Skip to main content
ClaudeWave
Skill694 repo starsupdated today

flow-next-plan-review

The flow-next-plan-review skill conducts rigorous architectural reviews of Flow specification and design documents using a John Carmack-level standard. It integrates with backend review engines, RepoPrompt, Codex CLI, or GitHub Copilot CLI, selected via command-line flag, environment variable, or config file, and functions as a code review coordinator managing spec validation and anti-pattern detection across distributed review workflows.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/gmickel/flow-next /tmp/flow-next-plan-review && cp -r /tmp/flow-next-plan-review/plugins/flow-next/skills/flow-next-plan-review ~/.claude/skills/flow-next-plan-review
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Plan Review Mode

**Workflow is backend-split. Read [workflow.md](workflow.md) for common
orchestration and backend resolution, then read ONLY the file matching the
selected review backend:**

- `BACKEND=codex` → [workflow-codex.md](workflow-codex.md)
- `BACKEND=copilot` → [workflow-copilot.md](workflow-copilot.md)
- `BACKEND=cursor` → [workflow-cursor.md](workflow-cursor.md)
- `BACKEND=host` → [workflow-host.md](workflow-host.md)
- `BACKEND=rp` → [workflow-rp.md](workflow-rp.md)

Do not load the other backend files. `BACKEND=none` and explicit
`--review=export` terminate from the common workflow without loading any backend
file.

Conduct a John Carmack-level review of spec plans.

**Role**: Code Review Coordinator (NOT the reviewer)
**Backends** (branch on the common workflow's `RP_ELIGIBLE` probe):
- When `RP_ELIGIBLE=1`: RepoPrompt (rp), Codex CLI (codex), GitHub Copilot CLI
  (copilot), Cursor CLI (cursor), or host-native (`host`)
- When `RP_ELIGIBLE=0`: Codex CLI, GitHub Copilot CLI, Cursor CLI, or
  host-native — rp remains accepted explicitly but errors at runtime

## Preamble — execute common routing exactly once

Read and execute [workflow.md](workflow.md) Phase 0 once. It defines `$FLOWCTL`,
probes RepoPrompt eligibility, parses an explicit `--review` mode before
configured-backend resolution, resolves `SPEC_ID`, and handles `ASK`, `none`,
and `export`. Never invoke `flowctl review-backend` a second time.

When `RP_ELIGIBLE=0`, never steer the user toward rp. An explicit
`--review=rp`, `FLOW_REVIEW_BACKEND=rp`, or `review.backend=rp` remains valid
input and fails through the rp runtime check.

## Backend Selection

Priority (first match wins):

1. `--review=rp|codex|copilot|cursor|host|export|none`
2. Per-spec `default_review`
3. `FLOW_REVIEW_BACKEND`
4. `.flow/config.json` `review.backend`
5. Error — no auto-detection

Configured values accept `backend[:model[:effort]]`; `cursor` takes a model but
no effort, and `host`, `rp`, and `none` are bare-only. `export` is a one-off
mode, never a configured backend.

## Common Critical Rules

- The coordinator never self-declares a verdict.
- Stick to one backend for the full review/fix cycle.
- If `REVIEW_RECEIPT_PATH` is set, every review verdict writes a receipt.
- Any backend/transport failure outputs `<promise>RETRY</promise>` and stops;
  never silently fall back to a different backend. Autonomous/Ralph callers
  receive the same retry terminal and decide whether to re-enter. A no-verdict
  dispatch is refunded and recorded by flowctl; never manually reset the review
  counter for a transport failure. Exit 5 / `TRANSPORT_UNHEALTHY` means stop
  automatic retries and repair the backend.
- `none` skips only when selected explicitly or resolved from configuration.
- `export` emits the existing external-review artifact and terminal output,
  then returns; it never loads configured-backend guidance, writes a review
  receipt/status, or enters the fix loop.
- **Foreground rule:** run every `flowctl <backend> plan-review` call as one **blocking foreground** Bash call with a generous timeout (10 minutes; verdicts typically land in 1–7) — never `run_in_background` + monitor/poll (a background completion does not reliably resume a subagent context). Host-backend subagent dispatches are also blocking.

Backend-specific invocation, availability, model, session-continuity, receipt,
and anti-pattern rules live only in the selected backend file.

## Input

Arguments: $ARGUMENTS

Format: `<flow-spec-id> [focus areas] [--review=<mode>]`

## Workflow

1. Execute [workflow.md](workflow.md) Phase 0.
2. If it returns for `none` or `export`, stop. Do not read a backend file.
3. Read exactly the selected `workflow-<backend>.md`.
4. Execute one backend dispatch and carry its verdict directly into the shared
   Fix Loop below.
5. Continue in that loop until its terminal contract is satisfied.

## Fix Loop (INTERNAL - do not exit to Ralph)

**The fix loop never pauses for user confirmation.** Every valid finding is
fixed and re-reviewed automatically. A loop that stops to ask, or that exits
with a valid finding unfixed, has broken this. Never use AskUserQuestion in this
loop.

`MAJOR_RETHINK` is not a fix-loop input. Surface the reviewer's rationale and
stop with `BLOCKED: DESIGN_CONFLICT` (Ralph: `<promise>RETRY</promise>`). Only
`NEEDS_WORK` enters the loop.

Fix+re-review cycles are bounded at `${MAX_REVIEW_ITERATIONS:-8}`. The counter
is flowctl-owned; never keep an agent-side counter. On cap exhaustion, surface
surviving findings and stop (Ralph: `<promise>RETRY</promise>`).

**The cap is enforced deterministically by flowctl:** every dispatch reserves a
spec-scoped round before launch. SHIP / NEEDS_WORK / MAJOR_RETHINK / NEEDS_HUMAN consume it;
a no-verdict transport failure is durably recorded and refunded. At
`${MAX_REVIEW_ITERATIONS:-8}` verdict rounds, flowctl refuses with `ESCALATE:`
and exit 4. More than `${MAX_REVIEW_TRANSPORT_FAILURES:-2}` consecutive
no-verdict failures stop separately with `TRANSPORT_UNHEALTHY` + exit 5.
Callers invoke plan-review once and act on its terminal result. The verdict
counter resets only on SHIP or an explicit re-plan, never on an edit, fresh
invocation, or transport failure.**

**ANTI-PATTERN:** a delivered verdict is never a transport failure - never
re-dispatch or re-frame `NEEDS_WORK` as a backend/sandbox problem to claim a
refund. And never widen the reviewer sandbox: reviewers are read-only by
contract, so a sandbox-blocked reviewer means something asked it to mutate the
workspace. Fix that instead (Windows resolves via `auto`).

When the verdict is `NEEDS_WORK`:

1. Parse all valid issues from reviewer feedback.
2. Fix the user-edited current spec, never a checkpoint copy:

   ```bash
   $FLOWCTL spec set-plan <SPEC_ID> --file - --json <<'EOF'
   <updated current spec content>
   EOF
   ```

3. Sync affected task specs when requirements, acceptance, design decisions,
   interfaces, retry/error semantics
specsSkill
flow-next-captureSkill

Synthesize the current conversation context into a flow-next spec at `.flow/specs/<spec-id>.md` via `flowctl spec create + spec set-plan` — agent-native, source-tagged, with mandatory read-back before write. Triggers on /flow-next:capture, "capture spec", "lock down what we discussed", "make a spec from this conversation", "convert conversation to spec". Optional `mode:autofix` token runs without questions and requires `--yes` to commit. Optional `--rewrite <spec-id>` overwrites an existing spec; `--from-compacted-ok` overrides the incomplete-evidence refusal after compaction; `--override-strategy` proceeds despite a contradiction with an active STRATEGY.md track (and prompts to record the override as a decision); `--no-plan` sets the spec-level `no_plan` field after the write (explicit opt-in — never inferred).

flow-next-make-prSkill

Render a cognitive-aid PR body from flow-next state and open via gh. Triggers on /flow-next:make-pr with optional spec id and flags (--draft, --ready, --no-mermaid, --base <ref>, --memory, --dry-run). Auto-detects spec from current branch when no id given. NOT Ralph-blocked — autonomous loops can surface a draft PR for human review.

flow-next-auditSkill

Audit `.flow/memory/` entries against the current codebase and decide Keep / Update / Consolidate / Replace / Delete / Harden per entry. Triggers on /flow-next:audit, "audit memory", "review memory", "refresh learnings", "sweep stale memory", "consolidate overlapping memory entries", "graduate a recurring lesson into a gate". Optional `mode:autofix` token in arguments runs without questions and marks ambiguous as stale (Harden is never auto-applied). Optional scope hint after the mode token (concept, category, module, or path) narrows what gets audited.

flow-next-depsSkill

Show spec dependency graph and execution order. Use when asking 'what's blocking what', 'execution order', 'dependency graph', 'what order should specs run', 'critical path', 'which specs can run in parallel'.

flow-next-driveSkill

Drive any UI surface like a real user - a web app, a Chromium-backed desktop app (Electron / WebView2, reached over CDP), or a genuinely native app (macOS AppKit/SwiftUI, or a non-CDP webview) reached via the Cua Driver / Computer Use. Detects the surface, picks the best available driver, degrades gracefully. Use to navigate sites, verify deployed UI, test web or desktop apps, capture baseline screenshots, drive a sign-in flow, scrape data, fill forms, run an e2e check, or inspect current page state. Triggers on "check the page", "verify UI", "test the site", "test this app", "drive the app", "automate this desktop app", "read docs at", "look up API", "visit URL", "browse", "screenshot", "scrape", "e2e test", "login flow", "capture baseline", "see how it looks", "inspect current", "before redesign", "Electron app", "native app".

flow-next-epic-reviewSkill

[deprecated alias] Renamed to flow-next-spec-completion-review in flow-next 1.0 — invoke the new skill. Removed in 2.0.

flow-next-export-contextSkill

Export RepoPrompt context to a markdown file for review with an external LLM (ChatGPT, Claude web, etc.). Use when you want Carmack-level review but prefer an external model. Triggers on "export context", "export for external review", "export plan for ChatGPT", "export impl review context", "review with an external model", "export review context".