Skip to main content
ClaudeWave
Skill7.9k repo starsupdated 4d ago

yao-image

Image expert. ALWAYS invoke this skill when you need to read, analyze, describe, or generate images. Use for screenshots, photos, charts, diagrams, AI-generated images, or any visual content.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/YaoApp/yao /tmp/yao-image && cp -r /tmp/yao-image/tools/skills/yao-image ~/.claude/skills/yao-image
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Image Tools

Use these tools when you encounter images you cannot read natively, or when you need to generate new images.

## image_read

Send an image to a vision-capable model and get a text description.

### Local file (most common):
```bash
tai tool image_read --image_path /path/to/image.png --prompt "Describe this image"
```

### URL:
```bash
tai tool image_read --image_path https://example.com/photo.jpg --prompt "What is shown?"
```

### With a specific vision provider:
```bash
tai tool image_read --image_path /path/to/image.png --prompt "Describe" --provider llm.my-openai:gpt-4o
```

| Parameter  | Type    | Required | Description                                                     |
| ---------- | ------- | -------- | --------------------------------------------------------------- |
| image_path | string  | yes      | Image file path or URL                                          |
| prompt     | string  | no       | Analysis instruction (default: describe in detail)              |
| max_size   | integer | no       | Max dimension in pixels for longest edge (default: 1080)        |
| provider   | string  | no       | Vision provider connector ID. If omitted, uses default vision model |

Images are automatically resized (preserving aspect ratio) before sending to the vision model.
Supported formats: PNG, JPEG, GIF, WebP.

## image_generate

Generate a new image from a text prompt (text-to-image). For editing an existing image, use `image_edit` instead.

### Basic usage (always specify output):
```bash
tai tool image_generate --prompt "A serene mountain landscape at sunset" --output landscape.png
```

### With specific provider, model and size:
```bash
tai tool image_generate --prompt "A futuristic city skyline" --provider llm.my-openai --model gpt-image-1 --dimensions 1792x1024 --output output/city.png
```

| Parameter | Type   | Required | Description                                                       |
| --------- | ------ | -------- | ----------------------------------------------------------------- |
| prompt     | string | yes      | Text description of the image to generate                         |
| output     | string | yes      | Output file path for the generated image                          |
| provider   | string | no       | Provider connector ID (use `image_providers` to list). Auto-selects if omitted |
| dimensions | string | no       | Image dimensions (default: 1024x1024). Common: 1024x1024, 1024x1792, 1792x1024 |
| model      | string | no       | Model name to use. Overrides the provider's default model         |

If `output` is omitted, the image is saved to a default path in the working directory.

## image_edit

Edit or transform an existing image based on a text prompt (image-to-image). Use for style transfer, background replacement, adding/removing elements, or any modification that requires a reference image.

### Basic usage:
```bash
tai tool image_edit --image_path /path/to/photo.png --prompt "Change the background to a beach scene" --output edited.png
```

### With URL image:
```bash
tai tool image_edit --image_path https://example.com/photo.jpg --prompt "Make it look like a watercolor painting" --output watercolor.png
```

### With specific provider and model:
```bash
tai tool image_edit --image_path /path/to/original.png --prompt "Remove the person in the foreground" --provider llm.my-openai --model gpt-image-1 --dimensions 1024x1024 --output result.png
```

| Parameter  | Type   | Required | Description                                                       |
| ---------- | ------ | -------- | ----------------------------------------------------------------- |
| image_path | string | yes      | Reference image file path or URL                                  |
| prompt     | string | yes      | Text description of the desired edit or transformation            |
| output     | string | yes      | Output file path for the edited image                             |
| provider   | string | no       | Provider connector ID (use `image_providers` with `capability=image_editing`). Auto-selects if omitted |
| dimensions | string | no       | Output dimensions (default: 1024x1024). Common: 1024x1024, 1024x1792, 1792x1024 |
| model      | string | no       | Model name to use. Overrides the provider's default model         |

If `output` is omitted, the image is saved to a default path in the working directory.

## image_providers

List available image providers filtered by capability.

### List image generation providers (default):
```bash
tai tool image_providers
```

### List image editing providers:
```bash
tai tool image_providers --capability image_editing
```

### List vision (image reading) providers:
```bash
tai tool image_providers --capability vision
```

| Parameter  | Type   | Required | Description                                                 |
| ---------- | ------ | -------- | ----------------------------------------------------------- |
| capability | string | no       | `image_generation` (default), `image_editing`, or `vision`  |

Returns a list of providers with their available models and connector IDs that can be passed to `image_generate`, `image_edit`, or `image_read`.

## Constraints

Only use the parameters listed above for each tool. Do not pass unsupported parameters (such as `quality`, `style`, `n`, `response_format`, etc.) — they will be ignored or cause errors.
secretSkill
yao-agentSkill

Agent management expert. ALWAYS invoke this skill when you need to list available agents, download or reference agent source code, deploy agent code to the host, or query the LLM connector matrix. Do not guess agent structures — use this skill first.

yao-audioSkill

Audio expert. ALWAYS invoke this skill when the user asks to transcribe, recognize, or convert speech/audio to text.

yao-boardSkill

Kanban board and task query expert. ALWAYS invoke this skill when the user asks about boards, tasks, task status, or project progress. Do not guess task state — use this skill first.

yao-docSkill

Yao process documentation expert. ALWAYS invoke this skill when the user needs to discover available processes, read process signatures, or validate process names. Do not guess process APIs — use this skill first.

yao-processSkill

Yao process execution expert. ALWAYS invoke this skill when the user needs to call a Yao process, query data models, run scripts, or check process permissions. Do not call processes without checking this skill first.

yao-secretSkill

Secret management expert. ALWAYS invoke this skill when you need to read API keys, tokens, or other secrets configured by the user. Never hardcode credentials — use this skill to retrieve them securely.

yao-webSkill

Web information retrieval expert. ALWAYS invoke this skill when the user needs to search the web, fetch a URL, or access real-time information beyond training data. Do not guess or use stale knowledge — use this skill first.