Skip to main content
ClaudeWave
Skill6.9k repo starsupdated 2d ago

audio-cog

OpenSquilla-compatible audio generation adapter for webpage audio requests. Prefer OpenRouter config/API key in OpenSquilla; preserve the upstream CellCog workflow only as optional ClawHub provenance.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/opensquilla/opensquilla /tmp/audio-cog && cp -r /tmp/audio-cog/src/opensquilla/skills/bundled/audio-cog ~/.claude/skills/audio-cog
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Audio Cog - AI Audio Generation Powered by CellCog

Create professional audio with AI — voiceovers, music, sound effects, and personalized avatar voices.

## Meta-Skill Entrypoint

Meta-skills should run this skill as `skill_exec` when they need OpenRouter
audio. The entrypoint is a deterministic Python adapter. During MetaSkill
execution it receives a short-lived provider connection from ordinary Provider
Settings in the child process only; the credential and endpoint never enter
`with`, argv, the plan, or persisted run data. It calls the configured
OpenRouter audio model, writes a browser-playable WAV file under the supplied
output directory, and prints either `AUDIO_READY:` or a single failure label.
Do not spawn an LLM sub-agent just to generate audio.

Prefer JSON payload mode when the caller already has a narration script:

```json
{"script": "exact spoken narration text"}
```

In payload mode the adapter asks the audio model to speak exactly that
transcript and not add acknowledgements, titles, or setup text.

## OpenSquilla Compatibility Contract

When invoked from OpenSquilla, this skill is an adapter around the caller's
configured provider. Do not require `CELLCOG_API_KEY`, do not assume the
`cellcog` package is installed, and do not invent provider credentials.

For `AwesomeWebpageMetaSkill`:

- Use the code-owned OpenRouter capability candidate and the volatile provider
  lease resolved from ordinary Provider Settings after explicit approval.
- Use only `config.awesome_webpage.openrouter.models.audio_generation` for
  audio model selection.
- Save generated or processed files only under
  `config.awesome_webpage.output_dir/project/assets/audio`.
- If the OpenRouter key, audio model, or output directory is missing, return a
  concise `AUDIO_CONFIG_NEEDED` report listing the missing config keys.
- If the configured OpenRouter model cannot return a browser-playable audio
  file, return `AUDIO_MODEL_UNSUPPORTED` with the narration/script, desired
  duration, style, and target filename so the webpage can expose a clean
  replacement slot instead of failing the whole project.

### On success: `AUDIO_READY` manifest line (required)

After every successful save, end your reply with one single-line JSON record
per file so `AwesomeWebpageMetaSkill` can collect and bind the assets:

```
AUDIO_READY: {"local_path": "project/assets/audio/<slug>.wav", "mime": "audio/wav", "duration_s": <int_or_null>, "voice": "<voice>", "script_preview": "<first 80 chars>"}
```

- One `AUDIO_READY:` line per audio file. No trailing prose on that line.
- `local_path` MUST be the relative path `project/assets/audio/...`. Do NOT
  emit an absolute path here.
- On failure, emit one of `AUDIO_CONFIG_NEEDED`, `AUDIO_MODEL_UNSUPPORTED`, or
  `AUDIO_GENERATION_FAILED` as a single-line label with the replacement-slot
  path so the page can render a placeholder.

## OpenRouter Audio API Contract (hard rule for `openai/gpt-audio*`)

The default CellCog code-path is **wrong** for OpenSquilla and will fail.
OpenRouter routes `openai/gpt-audio` / `openai/gpt-audio-mini` through OpenAI's
audio-output mode, which has a strict request shape:

- `POST {base_url}/chat/completions` with body:
  ```
  {
    "model": "<audio_generation>",
    "stream": true,
    "modalities": ["text", "audio"],
    "audio": {"voice": "alloy", "format": "pcm16"},
    "messages": [...]
  }
  ```
- `stream: true` is REQUIRED. Non-streaming requests are rejected with
  HTTP 400 "Audio output requires stream: true".
- `audio.format` MUST be `pcm16` when streaming. `mp3`, `opus`, `flac`,
  `wav` are all rejected as "unsupported_value" — there is no alternative
  combo. Sending `format=mp3` (any stream setting) burns ~190 s of
  per-attempt timeout for nothing; do not try it.
- Read the SSE response, base64-decode each `delta.audio.data` chunk,
  concatenate the raw 24kHz mono signed-16-bit-little-endian PCM stream,
  then save it as a browser-playable WAV file.
- Final on-disk asset is `.wav`. Set MIME to `audio/wav` in the manifest.
- If the required MetaSkill provider lease is missing or invalid, emit
  `AUDIO_CONFIG_NEEDED` and exit 78 before any provider submission. Direct
  standalone CLI use may still read `OPENROUTER_API_KEY`; never fall back to
  `CELLCOG_API_KEY` or another provider.
- Provider/model failures after submission remain an exit-0 degradation with a
  structured replacement slot. They are never automatically replayed.

Upstream CellCog instructions are intentionally omitted from the executable
prompt body. OpenSquilla meta-skills use the entrypoint above; provenance is
kept in frontmatter for registry/audit purposes.
advanced-dubbing-studioSkill

Submit audio or video for multilingual dubbing, poll status, and download dubbed audio. Use when the user asks for dubbing, 多语言配音, 视频翻译配音, 译制片, or wants a source clip dubbed into another language.

ai-video-scriptSkill

Generate a structured short-video shooting script from a topic. Emits a strict, machine-parseable shot list (3 shots by default) with image prompt + video prompt + voiceover + on-screen text per shot. Trigger when the user asks for a video script, 分镜, 短视频文案, AI视频, 短剧脚本, or wants visual prompts ready for image/video generation.

cronSkill

Use when the user asks to schedule recurring tasks, one-off reminders, timers, or cron-style jobs through the OpenSquilla cron tool.

deep-researchSkill

Multi-round research with explicit methodology, evidence tracking, and citation-tagged synthesis. Trigger on 'deep dive', 'research report', 'literature review', 'investigate X across sources', 'multi-round investigation'. Distinct from the `summarize` skill, which is a single-pass condensation; this skill maintains a state file across iterations, tracks coverage, and produces a long-form report with per-claim citations. Three execution stages: plan (scope into sub-questions), iterate (record evidence per round), compile (synthesize report). The skill itself does not fetch the web — it tells the host agent which fetches to perform via OpenSquilla's existing web tools, and records what comes back.

docxSkill

Read, edit, or create Microsoft Word `.docx` files. Trigger this skill whenever the user mentions a Word document, .docx file, contract, report, brief, memo, or asks to extract text, modify an existing doc, generate one from a brief, or audit tracked changes. Three execution paths: text-and-structure extraction, in-place edit-by-run (preserves styles), and create-from-scratch with python-docx. Falls back to OOXML unzip-and-patch for layout work python-docx cannot reach.

git-diffSkill

Capture the current git diff (staged, working-tree, or staged file list) as text. Direct shell call for workflows that need repository diffs without an LLM agent loop.

githubSkill

GitHub operations via `gh` CLI: issues, PRs, CI runs, code review, API queries. Use when: (1) checking PR status or CI, (2) creating/commenting on issues, (3) listing/filtering PRs or issues, (4) viewing run logs. NOT for: complex web UI interactions requiring manual browser flows (use browser tooling when available), bulk operations across many repos (script with gh api), or when gh auth is not configured.

history-explorerSkill

Query the per-turn DecisionEntry log for skill co-occurrence patterns, meta-skill usage stats, and the router fixture corpus. Returns a JSON summary suitable for downstream LLM consumption. Used by meta-skill-creator's harvest step but also useful standalone for 'which skills did I use most this week?'