Skip to main content
ClaudeWave
Skill1.2k repo starsupdated 3d ago

moai-ref-cross-model-audit

>

Install in Claude Code
Copy
git clone --depth 1 https://github.com/modu-ai/moai-adk /tmp/moai-ref-cross-model-audit && cp -r /tmp/moai-ref-cross-model-audit/.claude/skills/moai-ref-cross-model-audit ~/.claude/skills/moai-ref-cross-model-audit
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Cross-Model Audit Convergence

This skill is the single load-point both plan-auditor and sync-auditor use when
the project opts into multi-model audit (`audit_model: multi`). It documents the
one MCP tool the auditor calls, the independence rule that tool enforces, and
how to fold the returned convergence result into the auditor's verdict.

## When to use convergence vs single-model

| Project setting | Path | Skill |
|---|---|---|
| `audit_model: claude` (default) | Claude reviews alone | (none — no second opinion needed) |
| `audit_model: codex` | Codex reviews alone | `moai-ref-owasp-checklist` etc., no convergence |
| `audit_model: glm` | GLM reviews alone | (same) |
| `audit_model: multi` | Claude + codex + GLM, converged | **this skill** |

Single-model paths do NOT load this skill. Convergence is only the multi-model
concern.

## The `audit_multi` MCP tool

The single tool surface is:

```
mcp__moai__audit_multi
```

It is exposed by the `moai mcp-server` stdio server (the self-hosted MCP server
shipped with the binary). The tool is a thin wrapper over the convergence
engine: it does NOT re-implement the codex or GLM backends — it fans out by
calling the existing single-backend handlers in parallel and synthesizes their
results.

### Input parameters

| Parameter | Type | Required | Notes |
|---|---|---|---|
| `claude_verdict` | object | YES | The in-session Claude review verdict. Object shape: `{verdict, summary, findings, next_steps}` — the same `review-output.schema.json` the single-backend tools return. |
| `target` | string | no | What the secondary backends review (`uncommittedChanges`, `baseBranch`). Passed through unchanged. |
| `focus` | string | no | Optional focus area forwarded to the secondary backends (e.g. `concurrency`, `auth`). |
| `gates` | object | no | Per-auditor gate map (`claude`/`codex`/`glm` ∈ `off`/`advisory`/`required`). When omitted, distributed defaults apply: claude required, codex required, glm advisory. |
| `session_id` | string | no | When set, the result is persisted to `.moai/state/audit-multi/<session>.json` so the multi-review-gate Stop hook reads the most recent result rather than re-invoking convergence. |
| `project_root` | string | no *(REQUIRED in a worktree)* | The tree the backends should read — this session's own `git rev-parse --show-toplevel`. Omitted from a worktree, the fan-out reads the PRIMARY checkout instead, so the backends review a diff that is not the one under audit and nothing in the result says so. Omit it only in the primary checkout. An unusable path is rejected with an error naming it, never silently replaced. |

### Output shape

The tool returns a `ConvergenceResult`:

```json
{
  "per_backend_verdicts": [
    {"backend": "claude", "gate": "required", "verdict": "pass", "summary": "...", "findings": [], "next_steps": []},
    {"backend": "codex",  "gate": "required", "verdict": "fail", "summary": "...", "findings": [...], "next_steps": [...]},
    {"backend": "glm",    "gate": "advisory", "verdict": "pass", "summary": "...", "findings": [], "next_steps": []}
  ],
  "overall_verdict": "fail",
  "disagreement_flag": true,
  "residual_risk_note": "cross-model disagreement (advisory, NOT a block): pass=[claude(required), glm(advisory)] fail=[codex(required)]",
  "fail_open_backends": []
}
```

- `overall_verdict` ∈ `{pass, fail}` — the existing review-output values. No
  new enum (disagreement is a flag, not a verdict value).
- `disagreement_flag` is `true` when either a required split OR an advisory-only
  conflict was detected.
- `residual_risk_note` describes the convergence outcome in prose (which
  backend(s) failed, or the shape of the split). Surface this in the audit
  report's residual-risk section.
- `fail_open_backends` lists the backends that returned `inconclusive` (missing,
  unauthenticated, or erroring) — surfaced so the report can name them.

### How to invoke

Call the tool with the in-session Claude analysis folded into the
`claude_verdict` object. Do NOT pass the full Claude analysis text as prompt
context for the secondary backends — see the Independence rule below.

```
result = mcp__moai__audit_multi({
  claude_verdict: { verdict: <your verdict>, summary: <one-line>, findings: [...], next_steps: [...] },
  target: "uncommittedChanges",
  focus: "concurrency",
  session_id: <current session id>
})
```

The orchestrator-side question channel is preserved: the tool returns a
structured result, never prompts the user. On a missing anchor or inconclusive
condition, surface the structured `overall_verdict: fail` + `residual_risk_note`
in the audit report and let the orchestrator translate.

## Independence rule (load-bearing)

> **Pass only the synthesized `claude_verdict` object to the MCP tool — NEVER
> the full Claude analysis text as prompt context for the secondary backends.**

The secondary backends (codex, GLM) are SUPER-REVIEWS: uncorrelated second
opinions. Their value collapses to a re-sample of Claude's reasoning the moment
they see Claude's analysis. The convergence engine enforces this structurally —
the `claude_verdict` is consumed ONLY by the synthesis step, and the secondary
backends receive `(target, focus, project_root)` — a scope, an area name, and a
directory, carrying no analysis between them — but the auditor must not
undermine the invariant by pasting Claude's reasoning into the `focus` field
either.

Concretely:

- `focus` carries a short AREA name (`concurrency`, `auth`, `secret handling`),
  not a paragraph of analysis.
- `claude_verdict.summary` is a one-line verdict rationale, not the full review.
- The findings you surface in the audit report come from
  `per_backend_verdicts[].findings` (each backend's own findings), NOT from
  echoing Claude's findings back.

## Convergence policy (how overall_verdict is derived)

The engine derives `overall_verdict` per a 4-case table:

| Case | Condition | overall_verdict | disagreement_flag |
|---|---|---|---|
| 1 | All required backen