Skip to main content
ClaudeWave
Skill1.5k repo starsupdated 5d ago

image-gen

|

Install in Claude Code
Copy
git clone --depth 1 https://github.com/0xsline/OpenChatCut /tmp/image-gen && cp -r /tmp/image-gen/src/agent/skills/image-gen ~/.claude/skills/image-gen
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Image Gen

Generate AI images via `submit_image` (configured provider keys only). Prefer one clear still per request unless the user asked for variants.

## Model Selection

| Model | Reference | Strengths | Max refs |
| --- | --- | --- | --- |
| `gpt-image-2` | [references/gpt-image-2.md](references/gpt-image-2.md) | Best text rendering, strongest prompt adherence | 16 |
| `nano-banana` | [references/nano-banana.md](references/nano-banana.md) | Strongest reference-image fidelity | 14 |
| `image-01` | [references/image-01.md](references/image-01.md) | MiniMax stills / live style; one subject reference via R2 | 1 |
| `grok-imagine` | [references/grok-imagine.md](references/grok-imagine.md) | xAI Grok Imagine; text-to-image, ≤4 outputs, 1K/2K | 0 |

- Default: `gpt-image-2` when that key is on.
- Reference-heavy → `nano-banana`.
- User named MiniMax / only MiniMax image key on → `image-01`.
- Respect capabilities: do not call a model whose vendor is not configured.

**IMPORTANT:** Before generating, READ the chosen model's reference.

## Tool Params

| Param               | Values                                                                  | Default |
| ------------------- | ----------------------------------------------------------------------- | ------- |
| `aspectRatio`       | `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `3:2`, `2:3`, `4:5`, `5:4`, `21:9` | `16:9`  |
| `imageSize`         | `512px`, `1K`, `2K`, `4K` (model-specific)                              | `1K`    |
| `width` / `height`  | GPT Image: 512–3840, /16; MiniMax: 512–2048, /8                         | —       |
| `quality`           | `low`, `medium`, `high`, `auto` (gpt-image-2 only)                      | `high`  |
| `referenceAssetIds` | Array of project asset ids — backend resolves bytes server-side         | —       |
| `name`              | Short descriptive asset name shown in the library                       | —       |
| `count`             | Number of images to generate (1–10; image-01 max 9)                     | `1`     |
| `promptOptimizer`   | MiniMax `image-01` only — `prompt_optimizer`                            | `false` |
| `seed`              | MiniMax `image-01` only                                                  | —       |
| `maskAssetId`, `background`, `moderation`, `inputFidelity` | GPT Image edit/output controls | — |
| `outputFormat`, `outputCompression` | GPT Image PNG/JPEG/WebP controls                         | PNG     |

## Defaults

- Aspect ratio: **16:9**. If the project composition is not 16:9, ASK the user which aspect ratio they want before generating.
- Size: **1K**.

## Ask Before Submit

- Never auto-upgrade size.
- Only pass `imageSize: "2K"` or `"4K"` when the user explicitly asks. Warn that 2K/4K are EXPERIMENTAL and may be slower.

## Reference Images

Use when the user provides source material to edit, blend, or use as visual guidance (e.g. "change the background", "combine these into a poster").

- Pass project asset ids via `referenceAssetIds`. The backend fetches and encodes them server-side — never pull the asset bytes yourself.
- When the user @-references an image asset, pass its id directly in `referenceAssetIds`.
- Formats accepted by backend: png, jpeg, webp, svg (auto-rasterized to png), heic, heif. Each ≤ 50MB.

## Run

```ts
// Basic generation
submit_image({
  model: "gpt-image-2",
  prompt: "a cute orange cat",
  name: "Cat",
});

// With quality (gpt-image-2 only)
submit_image({
  model: "gpt-image-2",
  prompt: "hero poster with bold title",
  quality: "high",
  name: "Hero Poster",
});

// With reference images — pass project asset ids; backend resolves bytes
submit_image({
  model: "gpt-image-2",
  prompt: "change background to beach",
  referenceAssetIds: ["<assetId>"],
  name: "Beach Edit",
});

// Reference-heavy with nano-banana
submit_image({
  model: "nano-banana",
  prompt: "composite poster",
  referenceAssetIds: ["<id1>", "<id2>"],
  name: "Composite",
});

// Multiple images
submit_image({
  model: "gpt-image-2",
  prompt: "product shots",
  count: 3,
  name: "Product",
});

// MiniMax (optional single subject reference; R2 must be configured for refs)
submit_image({
  model: "image-01",
  prompt: "matte product bottle on marble, soft studio light",
  name: "Bottle still",
  promptOptimizer: false,
});
```

OpenChatCut’s `submit_image` may return completed pool assets synchronously depending on the provider path. If a `jobId` is returned, use `track_progress`; otherwise treat the asset ids in the result as done.

## Rules

- Always provide `name` with a short descriptive asset name.
- Before submitting, briefly tell the user what you're about to generate — especially when generating multiple images.
- Only call models whose vendor key is configured (capabilities prompt).
openchatcutSkill

Connect an MCP-capable coding agent to OpenChatCut and edit local video projects. Use when the user asks to install, connect, or set up OpenChatCut; inspect or edit an OpenChatCut project; work with its timeline, transcript, captions, media, generation, motion graphics, audio, color, or export tools; or recover from an OpenChatCut MCP error.

ai-cinematic-short-filmSkill

Plan AI short films with story, shots, prompts, and continuity.

asset-importSkill

Use when acquiring or importing media into a OpenChatCut project asset library for video editing or creation, including local/attached videos, user-provided paths, public media URLs, web video/audio/image assets, upload fallback decisions, and deciding between import_media, download_media, or manual user action.

create-motion-graphicsSkill

Use whenever the agent needs to add, create, hand-author, patch, or place Motion Graphic JSX assets in a OpenChatCut project. This is the direct-authoring path: use create_motion_graphic_from_code / edit_asset / edit_item, not motion-graphic-gen or submit_motion_graphic. Covers project/timeline intake, project visual language, editable properties, asset binding, inline JSX authoring, existing asset updates, timeline placement, and verification.

explainer-videoSkill

Create finished explainer videos from a topic, script, outline, voiceover, product logic, data, technical concept, course material, or reference assets. Use when the user wants narration, motion graphics, stock footage, generated visuals, or mixed visuals to explain an idea.

exportSkill

Use when a OpenChatCut video editing or creation workflow needs export, render, download, share, final delivery, subtitle-file export, render choice, local-only asset handling, or export fallback explanation.

known-errorsSkill

Use when a OpenChatCut tool call fails or returns an unexpected shape.

livestream-to-clipsSkill

Cut an imported livestream recording of any genre into evidence-backed, platform-ready clips by combining transcript, visual, audio, interaction, and domain-specific signals. Use for commerce, gaming, talk, interview, education, entertainment, sports, music, IRL, creative, news, or mixed livestream recordings.