Skip to main content
ClaudeWave
Skill262 repo starsupdated today

runbook

The runbook skill creates single, production-grounded incident response procedures with explicit commands, expected outputs, and staleness tracking. Use it to document actual alert scenarios, recurring operational tasks, or known failure modes backed by concrete evidence like alert history or incident reports, rather than speculative or hypothetical situations.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/testdouble/han /tmp/runbook && cp -r /tmp/runbook/han-documentation/skills/runbook ~/.claude/skills/runbook
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Create or Update Runbook

## Operating Principles

- **YAGNI applies to runbooks themselves.** Apply the evidence-based YAGNI rule from
  [../../references/yagni-rule.md](../../references/yagni-rule.md). A runbook is worth writing only when the scenario is
  grounded in something real: an alert that has actually fired, a documented incident, a recurring task that exists, or
  a known failure mode on a service that receives production traffic. Runbooks for hypothetical alerts, "best practice
  says we should have one," or "we'll need this someday" are YAGNI candidates and the runbook should be deferred until
  the scenario actually occurs. The canonical anti-pattern from project history: Sentry runbooks for staging-only Sentry
  where data isn't reaching production — alerts that will never fire because no signal flows. The user always wins; the
  rule's job is to make the cost of speculative runbooks visible.
- **The companion evidence rule applies to the runbook's supporting evidence.** Apply the evidence rule from
  [../../references/evidence-rule.md](../../references/evidence-rule.md) to the citations that ground the scenario: name
  the trust class of each piece of evidence (alert history, incident report, on-call rotation pattern); cite the actual
  artifact (dashboard URL, ticket ID, log query) rather than paraphrased recollection; and surface single-source claims
  as such rather than presenting them as settled.
- **One runbook per invocation.** The skill produces a single runbook file. Multi-runbook batches conflate scope; rerun
  the skill per scenario.
- **Imperative commands with expected output.** The template requires every step to show the exact command and what
  success looks like. Prose paragraphs in place of commands are an authoring failure the skill prompts against.
- **Staleness is the failure mode.** The template requires owner, last-validated, last-edited, and a change-history
  entry so decay is visible rather than hidden. The skill does not enforce a review cadence — that is a team-level
  workflow concern — but the metadata fields make the cadence auditable.
- **Readability standard.** Source the standard by invoking `han-communication:readability-guidance` and apply it as you
  write the runbook. Hold its default audience frame: a capable reader who did not do this work and lacks the author's
  context — here, the operator following the runbook during an incident.

## Project Context

- Git user: !`git config user.name || echo unset` (!`git config user.email || echo unset`)
- OS username: !`whoami`
- Today's date: !`date +%Y-%m-%d`
- CLAUDE.md: !`find . -maxdepth 1 -name "CLAUDE.md" -type f`
- project-discovery.md: !`find . -maxdepth 3 -name "project-discovery.md" -type f`
- personal config directory: !`bash "${CLAUDE_PLUGIN_ROOT}/scripts/han-config-dir.sh" 2>/dev/null || echo "$HOME/.claude"`
- project .han/config.md: !`cat .han/config.md 2>/dev/null || echo ""`

As your first action, use the Read tool on `.han/config.md` inside the `personal config directory` path above. A read
that returns no file is no personal configuration: continue silently. When that file or the `project .han/config.md`
probe supplies content, apply it per [config-rule.md](../../references/config-rule.md), which governs precedence
between the two files, relative-path resolution, and what to do with a file that reads but cannot be used.

## Step 1: Determine Mode

Determine which mode to operate in based on the user's request:

| Mode                | When                                                                                                             | Then                                                                    |
| ------------------- | ---------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
| Creating new        | Drafting a runbook for a scenario the project does not yet have one for                                          | → Step 2                                                                |
| Updating existing   | Modifying an existing runbook (new step, validation date refresh, escalation change)                             | Read the existing runbook → Step 4                                      |
| Validating existing | User says they ran the procedure end-to-end and wants to refresh `Last validated` and add a change-history entry | Read the existing runbook → Step 4 (update mode, validation entry only) |

## Step 2: Apply the YAGNI Preflight

Before discovering structure or gathering context, gate the work. Ask the user (or confirm from their request) which of
the following describes the scenario:

1. **An alert that has actually fired** — name the alert, link the firing incident or alert manager record.
2. **A documented incident or post-mortem** — link it.
3. **A recurring scheduled task** that the team performs (weekly index rebuild, monthly cert rotation, etc.) — name the
   cadence and where the schedule lives.
4. **A live failure mode** on a service that receives production traffic, where the failure has occurred or is expected
   to occur with current measured pressure — name the service and the failure mode.
5. **Customer report or stakeholder commitment** requiring this procedure to be documented now — link it.

If none of these applies, recommend deferring the runbook. Surface the recommendation to the user with the trigger that
would justify revisiting:

> "I don't see a current trigger forcing this runbook. Per the project's YAGNI rule, runbooks for alerts that have never
> fired are an anti-pattern. Recommend deferring until {trigger — first alert fires, first occurrence of the failure
> mode, first run of the recurring task, customer commitment lands}. Override and proceed anyway?"

The user always wins. If they override, record the override in the runbook's Origin field as
`"override: written preventive