audio-cog
OpenSquilla-compatible audio generation adapter for webpage audio requests. Prefer OpenRouter config/API key in OpenSquilla; preserve the upstream CellCog workflow only as optional ClawHub provenance.
git clone --depth 1 https://github.com/opensquilla/opensquilla /tmp/audio-cog && cp -r /tmp/audio-cog/src/opensquilla/skills/bundled/audio-cog ~/.claude/skills/audio-cogSKILL.md
# Audio Cog - AI Audio Generation Powered by CellCog
Create professional audio with AI — voiceovers, music, sound effects, and personalized avatar voices.
## Meta-Skill Entrypoint
Meta-skills should run this skill as `skill_exec` when they need OpenRouter
audio. The entrypoint is a deterministic Python adapter. During MetaSkill
execution it receives a short-lived provider connection from ordinary Provider
Settings in the child process only; the credential and endpoint never enter
`with`, argv, the plan, or persisted run data. It calls the configured
OpenRouter audio model, writes a browser-playable WAV file under the supplied
output directory, and prints either `AUDIO_READY:` or a single failure label.
Do not spawn an LLM sub-agent just to generate audio.
Prefer JSON payload mode when the caller already has a narration script:
```json
{"script": "exact spoken narration text"}
```
In payload mode the adapter asks the audio model to speak exactly that
transcript and not add acknowledgements, titles, or setup text.
## OpenSquilla Compatibility Contract
When invoked from OpenSquilla, this skill is an adapter around the caller's
configured provider. Do not require `CELLCOG_API_KEY`, do not assume the
`cellcog` package is installed, and do not invent provider credentials.
For `AwesomeWebpageMetaSkill`:
- Use the code-owned OpenRouter capability candidate and the volatile provider
lease resolved from ordinary Provider Settings after explicit approval.
- Use only `config.awesome_webpage.openrouter.models.audio_generation` for
audio model selection.
- Save generated or processed files only under
`config.awesome_webpage.output_dir/project/assets/audio`.
- If the OpenRouter key, audio model, or output directory is missing, return a
concise `AUDIO_CONFIG_NEEDED` report listing the missing config keys.
- If the configured OpenRouter model cannot return a browser-playable audio
file, return `AUDIO_MODEL_UNSUPPORTED` with the narration/script, desired
duration, style, and target filename so the webpage can expose a clean
replacement slot instead of failing the whole project.
### On success: `AUDIO_READY` manifest line (required)
After every successful save, end your reply with one single-line JSON record
per file so `AwesomeWebpageMetaSkill` can collect and bind the assets:
```
AUDIO_READY: {"local_path": "project/assets/audio/<slug>.wav", "mime": "audio/wav", "duration_s": <int_or_null>, "voice": "<voice>", "script_preview": "<first 80 chars>"}
```
- One `AUDIO_READY:` line per audio file. No trailing prose on that line.
- `local_path` MUST be the relative path `project/assets/audio/...`. Do NOT
emit an absolute path here.
- On failure, emit one of `AUDIO_CONFIG_NEEDED`, `AUDIO_MODEL_UNSUPPORTED`, or
`AUDIO_GENERATION_FAILED` as a single-line label with the replacement-slot
path so the page can render a placeholder.
## OpenRouter Audio API Contract (hard rule for `openai/gpt-audio*`)
The default CellCog code-path is **wrong** for OpenSquilla and will fail.
OpenRouter routes `openai/gpt-audio` / `openai/gpt-audio-mini` through OpenAI's
audio-output mode, which has a strict request shape:
- `POST {base_url}/chat/completions` with body:
```
{
"model": "<audio_generation>",
"stream": true,
"modalities": ["text", "audio"],
"audio": {"voice": "alloy", "format": "pcm16"},
"messages": [...]
}
```
- `stream: true` is REQUIRED. Non-streaming requests are rejected with
HTTP 400 "Audio output requires stream: true".
- `audio.format` MUST be `pcm16` when streaming. `mp3`, `opus`, `flac`,
`wav` are all rejected as "unsupported_value" — there is no alternative
combo. Sending `format=mp3` (any stream setting) burns ~190 s of
per-attempt timeout for nothing; do not try it.
- Read the SSE response, base64-decode each `delta.audio.data` chunk,
concatenate the raw 24kHz mono signed-16-bit-little-endian PCM stream,
then save it as a browser-playable WAV file.
- Final on-disk asset is `.wav`. Set MIME to `audio/wav` in the manifest.
- If the required MetaSkill provider lease is missing or invalid, emit
`AUDIO_CONFIG_NEEDED` and exit 78 before any provider submission. Direct
standalone CLI use may still read `OPENROUTER_API_KEY`; never fall back to
`CELLCOG_API_KEY` or another provider.
- Provider/model failures after submission remain an exit-0 degradation with a
structured replacement slot. They are never automatically replayed.
Upstream CellCog instructions are intentionally omitted from the executable
prompt body. OpenSquilla meta-skills use the entrypoint above; provenance is
kept in frontmatter for registry/audit purposes.Submit audio or video for multilingual dubbing, poll status, and download dubbed audio. Use when the user asks for dubbing, 多语言配音, 视频翻译配音, 译制片, or wants a source clip dubbed into another language.
Generate a structured short-video shooting script from a topic. Emits a strict, machine-parseable shot list (3 shots by default) with image prompt + video prompt + voiceover + on-screen text per shot. Trigger when the user asks for a video script, 分镜, 短视频文案, AI视频, 短剧脚本, or wants visual prompts ready for image/video generation.
Use when the user asks to schedule recurring tasks, one-off reminders, timers, or cron-style jobs through the OpenSquilla cron tool.
Multi-round research with explicit methodology, evidence tracking, and citation-tagged synthesis. Trigger on 'deep dive', 'research report', 'literature review', 'investigate X across sources', 'multi-round investigation'. Distinct from the `summarize` skill, which is a single-pass condensation; this skill maintains a state file across iterations, tracks coverage, and produces a long-form report with per-claim citations. Three execution stages: plan (scope into sub-questions), iterate (record evidence per round), compile (synthesize report). The skill itself does not fetch the web — it tells the host agent which fetches to perform via OpenSquilla's existing web tools, and records what comes back.
Read, edit, or create Microsoft Word `.docx` files. Trigger this skill whenever the user mentions a Word document, .docx file, contract, report, brief, memo, or asks to extract text, modify an existing doc, generate one from a brief, or audit tracked changes. Three execution paths: text-and-structure extraction, in-place edit-by-run (preserves styles), and create-from-scratch with python-docx. Falls back to OOXML unzip-and-patch for layout work python-docx cannot reach.
Capture the current git diff (staged, working-tree, or staged file list) as text. Direct shell call for workflows that need repository diffs without an LLM agent loop.
GitHub operations via `gh` CLI: issues, PRs, CI runs, code review, API queries. Use when: (1) checking PR status or CI, (2) creating/commenting on issues, (3) listing/filtering PRs or issues, (4) viewing run logs. NOT for: complex web UI interactions requiring manual browser flows (use browser tooling when available), bulk operations across many repos (script with gh api), or when gh auth is not configured.
Query the per-turn DecisionEntry log for skill co-occurrence patterns, meta-skill usage stats, and the router fixture corpus. Returns a JSON summary suitable for downstream LLM consumption. Used by meta-skill-creator's harvest step but also useful standalone for 'which skills did I use most this week?'