Skip to main content
ClaudeWave
Skill68.6k repo starsupdated 2d ago

senpi-qa

QA the omo Senpi adapter (packages/omo-senpi, packages/senpi-task) against the REAL senpi binary in strict isolation, and write every artifact to the one canonical evidence path .omo/evidence/omo-senpi-adapter/<slug>/. The live drivers under packages/omo-senpi/scripts/qa/ create their own isolated SENPI_CODING_AGENT_DIR and ignore the caller's, so the real ~/.senpi/agent is never written. Ships scripts/resolve-evidence-dir.mjs, which is the ONLY sanctioned way to pick an evidence directory: it rejects traversal, separators, absolute paths, and stray roots such as local-ignore/qa-evidence. Use whenever someone changes anything under packages/omo-senpi or packages/senpi-task, or wants to QA, smoke-test, verify, or debug the Senpi adapter, the task/team engine, the DAG, task RPC, or skill delivery. Triggers: senpi qa, qa senpi, senpi-qa, test senpi adapter, verify senpi task, senpi task e2e, senpi team e2e, task dag qa, live senpi driver, senpi evidence path.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent /tmp/senpi-qa && cp -r /tmp/senpi-qa/.agents/skills/senpi-qa ~/.claude/skills/senpi-qa
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Senpi QA

QA the omo Senpi adapter (`packages/omo-senpi/`) and the task engine
(`packages/senpi-task/`) by driving the REAL `senpi` binary. Unit tests never
count as live QA here: `bun run test:senpi` is the package gate, the drivers in
`packages/omo-senpi/scripts/qa/` are the harness proof.

## Golden rules

- **Evidence lives at exactly one path.** Every artifact goes under
  `.omo/evidence/omo-senpi-adapter/<slug>/`. Pick it with
  `scripts/resolve-evidence-dir.mjs` and nothing else — a hand-typed path is how
  runs end up somewhere like `local-ignore/qa-evidence/`, which no reviewer reads
  and the PR cannot cite.
- **The real agent dir stays untouched.** The live drivers build their own
  isolated `SENPI_CODING_AGENT_DIR` and deliberately IGNORE a caller-provided
  one, so `~/.senpi/agent` is never used as the sandbox. Report the driver's
  `realSenpiUntouched` / changed-path fields and the isolated agent-dir path;
  treat a whole-directory digest as supporting evidence, not proof by itself.
- **No binary means SKIP, not silence.** When `senpi` is absent the live drivers
  report `SKIP` or `FAIL` in their final JSON rather than degrading to the real
  home. A `SKIP` is not a pass — say so in the evidence README.
- **The captured JSON is the evidence.** No file on disk means the QA did not
  happen, which means no commit and no push.

## Resolve the evidence directory first

```bash
ev="$(node .agents/skills/senpi-qa/scripts/resolve-evidence-dir.mjs \
  --repo-root "$(git rev-parse --show-toplevel)" --slug <YYYYMMDD>-<short-slug>)"
mkdir -p "$ev"
```

The resolver returns an absolute path and creates nothing, so the caller decides
when the directory appears. A slug is ONE relative segment of lowercase letters,
digits, and hyphens (`20260820-senpi-qa-contract`). Separators, `.`/`..`,
traversal, absolute paths, and a non-git root are rejected with a non-zero exit
and a message naming the offending slug.

## Router: pick your case

| You changed… | Run | Proves |
|---|---|---|
| Any adapter code, as the fast precondition | `node packages/omo-senpi/scripts/qa/drive.mjs --self-test` | the driver + isolation harness itself works |
| Adapter wiring reaching a live session | `node packages/omo-senpi/scripts/qa/drive.mjs` | a real senpi run with the plugin loaded, isolated agent dir, and no attributed real-home changes |
| Task lifecycle (single + batch) | `SENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/task-e2e.mjs` | live task start/stream/terminal states |
| Team delivery, shutdown, reclaim, restart recovery | `SENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/team-e2e.mjs` | injection delivery and exactly-once recovery |
| Task RPC driver scripts | `node packages/omo-senpi/scripts/qa/task-rpc-e2e.mjs --self-test` | the RPC surface contract |
| Skill delivery into a task | `SENPI_BIN="$(command -v senpi)" node packages/omo-senpi/scripts/qa/task-load-skills-e2e.mjs` | skills reach the child |
| Continuation behavior | `node packages/omo-senpi/scripts/qa/probe-continuation.mjs` | turns continue as expected |
| DAG state machine / runners | `bun test packages/senpi-task` | unit + chaos invariants (NOT live proof) |

Point a driver's output at the resolved directory, e.g.:

```bash
TASK_E2E_OUT_DIR="$ev/live-task-dag" SENPI_BIN="$(command -v senpi)" \
  node packages/omo-senpi/scripts/qa/task-e2e.mjs
```

## Package gate

```bash
tsgo --noEmit -p packages/omo-senpi/tsconfig.json
bun run test:senpi
```

## Write the evidence README

Every run leaves `$ev/README.md` a reviewer can read without rerunning anything.
The required sections are the repo-wide evidence rules in the root
[`AGENTS.md`](../../../AGENTS.md) (what was tested / observed / why it is enough /
what was omitted). For Senpi, record the driver's changed-path/isolation fields
and sandbox agent-dir path. Some drivers report sandbox paths without removing
them; the caller must delete every task-owned sandbox and verify child PIDs are
terminal before writing the cleanup receipt.
get-unpublished-changesSkill

Compare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.

github-triageSkill

Read-only GitHub triage for issues AND PRs. 1 item = 1 background task (category: quick). Analyzes all open items and writes evidence-backed reports to /tmp/{datetime}/. Every claim requires a GitHub permalink as proof. NEVER takes any action on GitHub - no comments, no merges, no closes, no labels. Reports only. Triggers: 'triage', 'triage issues', 'triage PRs', 'github triage'.

hyperplanSkill

Adversarial multi-agent planning skill for omo-senpi. Self-orchestrates a 5-member hostile team (categories unspecified-low, unspecified-high, deep, ultrabrain, artistry) via the native lead team tools for ruthless cross-critique debate, distills only the insights that survive the attacks, then MANDATORILY hands the distilled bundle to a planner task (load_skills ulw-plan) for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', 'adversarial plan', 'hostile planning', 'cross-critique plan', '하이퍼플랜', '적대적 계획', '교차 비평'.

omomomoSkill

Easter egg command - about oh-my-opencode. Triggers: omomomo, about, easter egg.

opencode-qaSkill

QA opencode itself, per case: verify the CLI/terminal (opencode run, db, serve, export), prove a specific plugin hook/action/event fired via the SSE event stream, smoke-test the TUI under tmux, and investigate sessions in opencode's SQLite DB by id, title/name, or message text. Ships tested helper scripts (each with a --self-test) plus per-domain references. Use whenever someone wants to QA, smoke-test, verify, or debug opencode's CLI, HTTP server, plugin hooks/events, or TUI, or to find/inspect opencode sessions in the database. Triggers: opencode qa, qa opencode, test opencode, verify opencode hook, opencode session db, find opencode session by id/name/text, opencode tui test, opencode server health, opencode event stream.

pre-publish-reviewSkill

Nuclear-grade 16-agent pre-publish release gate. Runs /get-unpublished-changes to detect all changes since last npm release, spawns up to 10 ultrabrain agents for deep per-change analysis, invokes /review-work (5 agents) for holistic review, and 1 oracle for overall release synthesis. Runs ONLY when the user explicitly asks for a pre-publish review — a plain publish/release request MUST NOT trigger this; /publish ships directly. Triggers: 'pre-publish review', 'review before publish', 'release review', 'pre-release review', 'ready to publish?', 'can I publish?', 'pre-publish', 'safe to publish', 'publishing review', 'pre-publish check'.

publishSkill

Publish oh-my-opencode to npm by triggering the GitHub Actions publish workflow and verifying its artifacts. Ship-only: never runs pre-publish-review or re-reviews merged code unless the user explicitly asks. Argument: <patch|minor|major|explicit-semver>. Triggers: publish, release, deploy, npm publish.

remove-deadcodeSkill

Remove unused code from this project with ultrawork mode, LSP-verified safety, atomic commits. Triggers: remove dead code, dead code, cleanup, remove unused.