Skill5.8k repo starsupdated 4d ago
qa
General-purpose QA verdict for any artifact type
Install in Claude Code
Copygit clone --depth 1 https://github.com/Q00/ouroboros /tmp/qa && cp -r /tmp/qa/skills/qa ~/.claude/skills/qaThen start a new Claude Code session; the skill loads automatically.
Definition
SKILL.md
# /ouroboros:qa
Standalone quality assessment for any artifact — code, documents, API responses, test output, or custom content. Unlike `ooo evaluate` (3-stage formal verification pipeline), `ooo qa` is a fast single-pass verdict with actionable suggestions.
## Usage
```
ooo qa [file_path | artifact_text]
ooo qa # evaluate recent execution output
/ouroboros:qa [file_path | artifact_text] # plugin mode
```
**Trigger keywords:** "ooo qa", "qa check", "quality check"
## How It Works
The QA Judge evaluates an artifact against a quality bar and returns a structured verdict:
1. **Parse the Quality Bar** — What EXACTLY must be true to pass?
2. **Assess Dimensions** — Correctness, Completeness, Quality, Intent Alignment, Domain-Specific
3. **Render Verdict** — Score (0.0-1.0) with PASS / REVISE / FAIL
4. **Determine Loop Action** — `done` (pass), `continue` (revise), `escalate` (fail)
### Verdict Thresholds
| Score Range | Verdict | Loop Action |
|--------------|---------|-------------|
| >= 0.80 | PASS | done |
| 0.40 - 0.79 | REVISE | continue |
| < 0.40 | FAIL | escalate |
## Instructions
When the user invokes this skill:
### Step 0: Determine execution mode
This skill works in two modes. Determine which one **before** attempting any tool calls:
- **MCP mode** — If the QA MCP tool is available (already exposed, or loadable via discovery), use it:
```
tool discovery query: "+ouroboros qa"
```
If found (typically named `mcp__plugin_ouroboros_ouroboros__ouroboros_qa`), proceed with **QA Steps** below.
- **Fallback mode** — Only if the QA MCP tool is genuinely absent (no Ouroboros MCP server) skip to the **Fallback** section; an empty discovery result for an already-exposed tool is expected — call it directly rather than falling back. This skill is designed to work without MCP setup.
### QA Steps (MCP mode)
1. **Determine the artifact to evaluate:**
- If user provides a file path: Read the file with Read tool
- If user provides inline text: Use that directly
- If no artifact specified: Look for the most recent execution output in conversation context
- Ask user if unclear what to evaluate
2. **Determine the quality bar:**
- If a seed YAML is available in context: Extract acceptance criteria from it
- If user specifies a quality bar: Use that
- If neither: Ask the user "What does 'good' mean for this artifact?"
3. **Determine artifact type:**
- `code` — source code files
- `test_output` — test results, CI output
- `document` — specs, docs, READMEs
- `api_response` — API responses, JSON payloads
- `screenshot` — visual artifacts
- `custom` — anything else
3.5. **Acting verification fan-out — probe in parallel, then judge (do not skip for behaviour-bearing artifacts):**
A text judge can be fooled by a hopeful log line. When the artifact actually
*does* something (code, an app, an API, a UI), fan out empirical probes
using the host's native parallel sub-agent primitive — one probe sub-agent
per acting modality the runtime actually exposes, all spawned **in the same
message so they run concurrently**:
- **process probe** (`Bash`/shell): run the command / start the app / run
the declared smoke commands with bounded timeouts; capture exit codes and
real output.
- **browser probe** (browser-use tools, when the artifact serves HTTP or is
a web UI): load it, click the primary flows, capture what actually
renders and any console/network errors.
- **computer-use probe** (desktop computer-use tools, when the artifact is
a GUI/TUI): drive it like a user, screenshot the observed states.
- **artifact probe** (file reads): verify declared files/paths exist with
real content, not placeholders.
Each probe returns structured evidence only — commands run, observed
effects, screenshots/paths, pass/fail per probed behaviour. Every probe
also hits the applicable adversarial classes (the QA tool lists them):
`misleading_output` (claimed success vs. real effect), `hung_command`
(bounded timeout?), `malformed_input`, `stale_state`, `dirty_worktree`.
Skip a modality only when its tools are absent or the artifact type makes
it meaningless — and say which modalities were skipped and why.
Await all probes, then pass the merged evidence into the judge as
`reference` (prefer observed behaviour over source text as the `artifact`
when they disagree). **Empirical evidence outranks the judge**: if the
judge scores PASS but any probe observed the behaviour failing, present
the verdict as REVISE/FAIL on that evidence and say so explicitly — a
score contradicted by observation is not a pass. If no acting tools are
available at all, judge on the text alone but flag that behaviour was not
observed.
4. **Call the `ouroboros_qa` MCP tool:**
```
Tool: ouroboros_qa
Arguments:
artifact: <the content to evaluate>
quality_bar: <what 'pass' means>
artifact_type: "code" (or other type)
reference: <observed-behaviour evidence from step 3.5, plus any reference>
pass_threshold: 0.80 (adjustable)
seed_content: <seed YAML if available>
```
5. **Present results clearly:**
- Show the score and verdict prominently
- List dimension scores
- Highlight specific differences found
- Show actionable suggestions
- End with next step guidance based on verdict:
- **PASS (done)**: `Next: Your artifact meets the quality bar. Proceed with confidence.`
- **REVISE (continue)**: `Next: Address the suggestions above, then run ooo qa again to re-check.`
- **FAIL (escalate)**: `Next: Fundamental issues detected. Consider ooo interview to re-examine requirements, or ooo unstuck to challenge assumptions.`
### Iterative QA Loop
For iterative usage, track the `qa_session_id` and `iteration_history` from the response meta:
1. First call returns `qa_session_id` aMore from this repository
autoSkill
Automatically converge from goal to A-grade Seed and execute it
brownfieldSkill
Scan and manage brownfield repository/worktree defaults for interviews
cancelSkill
Cancel stuck or orphaned executions
ouroboros-configSkill
Open or drive the Ouroboros settings GUI (browser, TUI, or conversational fallback)
evaluateSkill
Evaluate execution with three-stage verification pipeline
evolveSkill
Start or monitor an evolutionary development loop
ouroboros-helpSkill
Full reference guide for Ouroboros commands and agents
interviewSkill
Socratic interview to crystallize vague requirements