Skip to main content
ClaudeWave
Skill5.8k estrellas del repoactualizado 4d ago

qa

General-purpose QA verdict for any artifact type

Instalar en Claude Code
Copiar
git clone --depth 1 https://github.com/Q00/ouroboros /tmp/qa && cp -r /tmp/qa/skills/qa ~/.claude/skills/qa
Después abre una sesión nueva de Claude Code; el skill carga automáticamente.

SKILL.md

# /ouroboros:qa

Standalone quality assessment for any artifact — code, documents, API responses, test output, or custom content. Unlike `ooo evaluate` (3-stage formal verification pipeline), `ooo qa` is a fast single-pass verdict with actionable suggestions.

## Usage

```
ooo qa [file_path | artifact_text]
ooo qa                                     # evaluate recent execution output
/ouroboros:qa [file_path | artifact_text]   # plugin mode
```

**Trigger keywords:** "ooo qa", "qa check", "quality check"

## How It Works

The QA Judge evaluates an artifact against a quality bar and returns a structured verdict:

1. **Parse the Quality Bar** — What EXACTLY must be true to pass?
2. **Assess Dimensions** — Correctness, Completeness, Quality, Intent Alignment, Domain-Specific
3. **Render Verdict** — Score (0.0-1.0) with PASS / REVISE / FAIL
4. **Determine Loop Action** — `done` (pass), `continue` (revise), `escalate` (fail)

### Verdict Thresholds

| Score Range  | Verdict | Loop Action |
|--------------|---------|-------------|
| >= 0.80      | PASS    | done        |
| 0.40 - 0.79  | REVISE  | continue    |
| < 0.40       | FAIL    | escalate    |

## Instructions

When the user invokes this skill:

### Step 0: Determine execution mode

This skill works in two modes. Determine which one **before** attempting any tool calls:

- **MCP mode** — If the QA MCP tool is available (already exposed, or loadable via discovery), use it:
  ```
  tool discovery query: "+ouroboros qa"
  ```
  If found (typically named `mcp__plugin_ouroboros_ouroboros__ouroboros_qa`), proceed with **QA Steps** below.

- **Fallback mode** — Only if the QA MCP tool is genuinely absent (no Ouroboros MCP server) skip to the **Fallback** section; an empty discovery result for an already-exposed tool is expected — call it directly rather than falling back. This skill is designed to work without MCP setup.

### QA Steps (MCP mode)

1. **Determine the artifact to evaluate:**
   - If user provides a file path: Read the file with Read tool
   - If user provides inline text: Use that directly
   - If no artifact specified: Look for the most recent execution output in conversation context
   - Ask user if unclear what to evaluate

2. **Determine the quality bar:**
   - If a seed YAML is available in context: Extract acceptance criteria from it
   - If user specifies a quality bar: Use that
   - If neither: Ask the user "What does 'good' mean for this artifact?"

3. **Determine artifact type:**
   - `code` — source code files
   - `test_output` — test results, CI output
   - `document` — specs, docs, READMEs
   - `api_response` — API responses, JSON payloads
   - `screenshot` — visual artifacts
   - `custom` — anything else

3.5. **Acting verification fan-out — probe in parallel, then judge (do not skip for behaviour-bearing artifacts):**
   A text judge can be fooled by a hopeful log line. When the artifact actually
   *does* something (code, an app, an API, a UI), fan out empirical probes
   using the host's native parallel sub-agent primitive — one probe sub-agent
   per acting modality the runtime actually exposes, all spawned **in the same
   message so they run concurrently**:
   - **process probe** (`Bash`/shell): run the command / start the app / run
     the declared smoke commands with bounded timeouts; capture exit codes and
     real output.
   - **browser probe** (browser-use tools, when the artifact serves HTTP or is
     a web UI): load it, click the primary flows, capture what actually
     renders and any console/network errors.
   - **computer-use probe** (desktop computer-use tools, when the artifact is
     a GUI/TUI): drive it like a user, screenshot the observed states.
   - **artifact probe** (file reads): verify declared files/paths exist with
     real content, not placeholders.
   Each probe returns structured evidence only — commands run, observed
   effects, screenshots/paths, pass/fail per probed behaviour. Every probe
   also hits the applicable adversarial classes (the QA tool lists them):
   `misleading_output` (claimed success vs. real effect), `hung_command`
   (bounded timeout?), `malformed_input`, `stale_state`, `dirty_worktree`.
   Skip a modality only when its tools are absent or the artifact type makes
   it meaningless — and say which modalities were skipped and why.

   Await all probes, then pass the merged evidence into the judge as
   `reference` (prefer observed behaviour over source text as the `artifact`
   when they disagree). **Empirical evidence outranks the judge**: if the
   judge scores PASS but any probe observed the behaviour failing, present
   the verdict as REVISE/FAIL on that evidence and say so explicitly — a
   score contradicted by observation is not a pass. If no acting tools are
   available at all, judge on the text alone but flag that behaviour was not
   observed.

4. **Call the `ouroboros_qa` MCP tool:**
   ```
   Tool: ouroboros_qa
   Arguments:
     artifact: <the content to evaluate>
     quality_bar: <what 'pass' means>
     artifact_type: "code"  (or other type)
     reference: <observed-behaviour evidence from step 3.5, plus any reference>
     pass_threshold: 0.80  (adjustable)
     seed_content: <seed YAML if available>
   ```

5. **Present results clearly:**
   - Show the score and verdict prominently
   - List dimension scores
   - Highlight specific differences found
   - Show actionable suggestions
   - End with next step guidance based on verdict:
     - **PASS (done)**: `Next: Your artifact meets the quality bar. Proceed with confidence.`
     - **REVISE (continue)**: `Next: Address the suggestions above, then run ooo qa again to re-check.`
     - **FAIL (escalate)**: `Next: Fundamental issues detected. Consider ooo interview to re-examine requirements, or ooo unstuck to challenge assumptions.`

### Iterative QA Loop

For iterative usage, track the `qa_session_id` and `iteration_history` from the response meta:

1. First call returns `qa_session_id` a