Skip to main content
ClaudeWave
Skill977 repo starsupdated today

voice

The voice skill converts text to speech audio by executing the `mb voice` CLI command, generating MP3 files from provided text. Use it when users request audio generation, ask you to speak or read aloud, or want voice recordings of text content, with support for multiple TTS providers including Edge TTS, OpenAI, ElevenLabs, and Doubao with configurable voices and output options.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/xvirobotics/metabot /tmp/voice && cp -r /tmp/voice/src/skills/voice ~/.claude/skills/voice
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

## Text-to-Speech (Voice Output)

Generate MP3 audio from text using the `metabot voice tts` CLI.

### Quick Commands

```bash
# Generate MP3, prints file path to stdout
metabot voice tts "Hello, this is a test"

# Generate and play immediately
metabot voice tts "Hello" --play

# Save to specific file
metabot voice tts "Hello" -o greeting.mp3

# Override provider and voice
metabot voice tts "Hello" --provider doubao --voice zh_female_wanqudashu_moon_bigtts

# Pipe text (useful for long content)
echo "Long text here" | metabot voice tts
echo "Long text" | metabot voice tts -o output.mp3
```

### When to Use

- User asks you to "say", "speak", "read aloud", or "generate audio/voice"
- User wants a voice recording or audio version of text
- User requests TTS (text-to-speech) output

### Available Providers & Voices

**Edge TTS (default, free, no key needed):**
- `zh-CN-XiaoyiNeural` (default) — Female Chinese
- `zh-CN-YunxiNeural` — Male Chinese
- `zh-CN-XiaoxiaoNeural` — Female Chinese
- `en-US-JennyNeural` — Female English

**Doubao (default when Volcengine keys configured):**
- `zh_female_wanqudashu_moon_bigtts` (default) — Female Chinese
- Other Volcengine voice IDs from the TTS console

**OpenAI (when OPENAI_API_KEY set):**
- `alloy` (default), `echo`, `fable`, `onyx`, `nova`, `shimmer`

**ElevenLabs (when ELEVENLABS_API_KEY set):**
- Voice IDs from the ElevenLabs console

### Text Limits

- Doubao: ~300 Chinese characters (longer text is auto-truncated)
- OpenAI / ElevenLabs / Edge: ~4000 characters

### Guidelines

- For short text (greetings, alerts), use inline: `metabot voice tts "text"`
- For longer text, pipe through stdin: `echo "..." | metabot voice tts`
- The output file is MP3 format
- Use `--play` only when the user explicitly wants to hear the audio (it blocks until playback completes)
- When saving files for the user, use `-o` with a descriptive filename
- To send the audio to the user in Feishu, copy the file to the outputs directory:
  `cp /tmp/metabot-voice-xxx.mp3 /tmp/metabot-outputs/<chatId>/`