Install in Claude Code
Copygit clone --depth 1 https://github.com/Orkas-AI/Orkas /tmp/video-craft && cp -r /tmp/video-craft/resources/builtin/marketplace/agents/79df9cc89f5f/skills/video-craft ~/.claude/skills/video-craftThen start a new Claude Code session; the skill loads automatically.
Definition
SKILL.md
# video-craft
The difference between a video that merely renders and one that's actually good. Apply these standards while scripting, storyboarding, composing, generating, and editing — and run the self-check before publishing. These are general production-craft norms; the exact numbers are starting points, adjust to the brief.
## 1. The opening (hook)
- The first **1–3 seconds** decide whether anyone keeps watching. Frame 1 must already carry motion or a text hook — no blank intro, no logo sting, no slow build.
- Strong hook shapes: a sharp question, a counter-intuitive claim, the promised outcome ("by the end you'll…"), the stakes, or showing the finished result first ("here's what we'll build").
- On muted autoplay the **on-screen text** is the hook — assume no sound for the first beat.
- *Weak → strong:* ✗ "In this video we'll look at caching." (slow build, no stakes) → ✓ frame 1, bold on-screen text "Your API is slow. One line fixes it." (stakes + promised outcome, readable muted).
## 2. Story structure
- Arc: **hook → tension/gap → core idea(s) → proof/example → payoff/close (+ optional CTA)**. Land the first real payoff early; viewers drop off fast before they get value.
- **One new idea per ~30–45 s** of explainer. A 3-min video carries 4–6 ideas, no more. Cut "interesting but irrelevant" — it actively lowers comprehension.
- Connect beats with **"but" / "therefore"**, not "and then" — force logical (not just sequential) progression.
- For teaching: show the naive idea, let it half-work, then break it and introduce the one key insight — people remember what they feel they discovered. Surfacing a common misconception first, then correcting it, beats stating the right answer cold.
- Narration cadence: explainer **~150–160 wpm**, social/short **~180–200 wpm**, cinematic **~140–150 wpm**. Leave a **1–3 s** silence after a big reveal; avoid dead air > ~1.5 s between sentences.
## 3. Pacing & timing
- Shot/scene holds by format: explainer **~4–8 s**, short-form social **~1–3 s**, cinematic/contemplative **~10–20 s**. Don't hold the same length three times running — vary it.
- Cutting energy: rapid (15–30 cuts/min) = urgency; moderate (8–15) = standard teaching; slow (3–6) = documentary calm.
- A visual or audio **pattern interrupt** every ~20–30 s (short-form) / ~45–90 s (long-form) to re-grab attention.
- Completion drops with length (15 s clips finish far more often than 60 s) — keep it as short as the message allows; don't pad.
- Build animation timing **to the narration words**, not arbitrary beats. Hold a fully-built scene/chart **≥ 2–3 s** before moving on.
## 4. Visual design
- Composition: rule-of-thirds for key elements; center for stable establishing frames. One clear focal point per frame.
- Palette: **≤ 3–5 colors on screen** at once. Background is the least-saturated; foreground (largest, most saturated, central) reads first.
- Typography: **1–2 font families** total. At 1080p, titles ~60–90 px, body ~40–60 px, **never below ~40 px**; title at least ~50% larger than body. Max ~2 lines, ~32–42 chars/line.
- **Safe zones**: keep text inside ~80% of the frame (~192 px margin at 1080p). For vertical, keep essential content out of the bottom ~300 px (platform UI sits there) and away from the very top.
- Contrast ≥ **4.5:1** for any text (white-on-dark is safest). Use brightness/saturation, not red-green, to distinguish elements.
- Consistency is a feeling of quality: one color grade/LUT, one type system, 2–3 transition types — for the whole video.
## 5. Motion & animation
- **Never linear easing** — it reads robotic. Default ease-in-out for moves, ease-out for entrances (settle in), ease-in for exits (accelerate away). Vary easing; don't repeat one three times.
- Restraint: **one main element moves at a time**; stagger multi-element reveals by ~100–200 ms. A "static" shot means truly zero motion/zoom — if it moves, name the move.
- *Weak → strong:* ✗ four cards fly in together, linear, all 0.3 s → ✓ stagger them ~120 ms, ease-out, each settling before the next starts (one focal point at a time).
- Entrances ~0.3–0.5 s then hold for readable dwell; exits ~0.5–1 s. Optional small overshoot (~10–15%, settle in a few frames) adds life for playful pieces.
- Kinetic type: reveal text to a readable dwell (~3 s per ~60 chars); word-by-word reveal synced to narration boosts attention.
- Transitions carry meaning: **hard cut** = same topic/new angle (most invisible, most professional); **crossfade** = gentle topic change; **wipe/slide** = sequential steps; **zoom in/out** = into detail / out to context. Pick a small set and keep them consistent.
## 6. Captions / on-screen text
- Most social viewing is **muted** — captions are mandatory and are part of the pacing, not an afterthought.
- Bold sans-serif, **≥ ~42 px**, ≤ ~2 lines, ≤ ~32–42 chars/line, with a dark stroke or semi-opaque backing for legibility on any footage.
- Don't scroll text off before it can be read (~3 s per ~60 chars). Word-by-word highlighting in sync with the voice reads best.
- Lower-thirds: speaker name bold + role lighter; enter ~1–2 s, hold a few seconds, exit fast; never cover eyes/mouth.
- Don't make the viewer read on-screen text AND listen to different words at once.
## 7. Audio
- Pick music by energy: calm ~60–80 BPM, standard explainer ~90–110, upbeat ~110–130, high-energy ~120–140+. Use **instrumental** under narration (lyrics fight the voice); avoid big crescendos that bury speech.
- Levels: narration peaks loudest; **duck music ~18–20 dB below speech**; SFX between. Master around **−14 LUFS**, true-peak ≤ ~−1 dB; never clip 0 dB.
- SFX land **~10–20 ms before** the visual change they accent (ears lead eyes). Keep stacked SFX in different frequency bands.
- Silence is a tool — drop music for a few seconds at a major reveal; let the moment land.
## 8. Platform & format
- Aspect: **9:16** (TikTok/Reels/Shorts), **16:9** (YouTube/web), **1:1** when speaker+context both matter, cinematic lett