Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
git clone https://github.com/cyanheads/evals-mcp-server{
"mcpServers": {
"evals": {
"command": "node",
"args": ["/path/to/evals-mcp-server/dist/index.js"]
}
}
}MCP Servers overview
<div align="center">
<h1>@cyanheads/evals-mcp-server</h1>
<p><b>Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.</b>
<div>9 Tools • 1 Resource</div>
</p>
</div>
<div align="center">
[](./CHANGELOG.md) [](./LICENSE) [](https://github.com/users/cyanheads/packages/container/package/evals-mcp-server) [](https://modelcontextprotocol.io/) [](https://www.npmjs.com/package/@cyanheads/evals-mcp-server) [](https://www.typescriptlang.org/) [](https://bun.sh/)
</div>
<div align="center">
[](https://github.com/cyanheads/evals-mcp-server/releases/latest/download/evals-mcp-server.mcpb) [](https://cursor.com/en/install-mcp?name=evals-mcp-server&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIkBjeWFuaGVhZHMvZXZhbHMtbWNwLXNlcnZlciJdfQ==) [](https://vscode.dev/redirect?url=vscode:mcp/install?%7B%22name%22%3A%22evals-mcp-server%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22%40cyanheads%2Fevals-mcp-server%22%5D%7D)
[](https://www.npmjs.com/package/@cyanheads/mcp-ts-core)
</div>
---
## Tools
Nine tools for authoring eval records — the draft loop (create, revise, discard, submit), the standalone deterministic checker, and read/list/export:
| Tool | Description |
|:---|:---|
| `evals_describe_schema` | Return the required and optional fields plus grader options for a task type. Call before drafting. |
| `evals_create_draft` | Create a draft eval record carrying its own grader; returns the parsed record, a review protocol, and a verification subagent prompt. |
| `evals_get_record` | Read a draft or submitted record by id; the id is stable across submit. |
| `evals_revise_draft` | Apply a surgical `set` / `append` / `unset` patch to a draft by dotted path; re-runs the self-consistency check. |
| `evals_discard_draft` | Delete a draft record by id. Draft-only. |
| `evals_run_check` | Run a grader spec against candidate answers and get PASS/REJECT per candidate, decoupled from any saved record. |
| `evals_submit_draft` | Finalize a draft through the committability gate, then freeze it. |
| `evals_list_records` | Browse and filter records by status, domain, task type, or tag. Returns a compact summary per record. |
| `evals_export_records` | Compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness and write the artifact under `exports/`. |
### `evals_describe_schema`
Return what a record of a given `task_type` needs before you draft it.
- Static — derived from the record and grader Zod schemas, no disk or runtime state
- Per-type gold shape, appropriate grader kind(s), required/optional fields, and authoring notes
- `task_type` is one of `numeric`, `exact_answer`, `set_answer`, `mcq`, `regex_answer`, `json_answer`, `free_response`
---
### `evals_create_draft`
Create and persist a draft eval record, then reflect it back as a review forcing function.
- Validates against the `task_type` discriminated union (per-type rules: `mcq` requires `choices`; `free_response` requires an `llm_rubric` grader)
- Runs a cheap self-consistency check — grader vs gold and each positive must PASS, vs each negative must REJECT
- Returns the normalized record parroted back behind a divider, a per-field review protocol, a ready-to-paste verification subagent prompt, and what's still required before submit
- Optional draft-time `verification` block and `captures` (EvalsIDs) when you already hold provenance
- Stays `draft` — passing self-consistency proves the grader discriminates, not that the gold is right
---
### `evals_revise_draft`
Surgically patch a draft so each change stays legible.
- Explicit `set` (dotted-path → value), `append` (dotted-path → array items), and `unset` (dotted paths) operations — not full-record rewrites
- Returns the updated record, an itemized list of what changed, and a re-run self-consistency verdict
- Re-validates the full shape and cross-field constraints after the patch
- Draft-only — submitted records are frozen; `task_type` cannot be patched (start a new draft to change the discriminant)
---
### `evals_run_check`
Run a grader against one or more candidates without touching a saved record.
- PASS/REJECT per candidate plus the resolved comparison value (e.g. the math.js-evaluated numeric target), so you see why each matched or missed
- `candidates` accepts strings, numbers, objects, or arrays — whatever the grader kind expects
- Supply `gold` for gold-relative kinds (`exact_match`); it is a no-op for target-embedding kinds like `numeric` and `mcq`
- `llm_rubric` cannot run on this server — submission relies on recorded independent verification
---
### `evals_submit_draft`
Finalize a draft through the committability gate, then freeze it.
- The gate runs the grader against the gold (must PASS), requires ≥1 declared negative case to be REJECTED, and requires a recorded, decorrelated independent verification that agrees with the gold
- Resolves and embeds any `captures` from `EVALS_CAPTURE_DIR`, cross-checking the gold against the authoritative captured value
- Rejects duplicates (same `content_hash` already submitted)
- On pass, flips the record to `submitted`, stamps `submitted_at` and a `checksum`, and freezes it; otherwise refuses with a typed error and the record stays a draft
- `free_response` `llm_rubric` is admitted on recorded independent verification and flagged `server_verified: false`
---
### `evals_export_records`
Compile submitted records to a downstream eval format.
- `jsonl` (lossless, the lingua franca), `csv` (a flattened, lossy spreadsheet summary), `inspect` (UK AISI Inspect AI), `lm-eval` (EleutherAI lm-evaluation-harness)
- Optional `domain` / `task_type` / `tag` filter
- Only submitted records are exported — drafts are skipped
- Writes the artifact under `exports/` and returns the file path, record count, byte size, and a short preview instead of dumping inline
## Resource
| Type | Name | Description |
|:---|:---|:---|
| Resource | `eval://record/{id}` | A single draft or submitted record by id — the same payload `evals_get_record` returns, for resource-capable clients. |
All record data is also reachable through the tool surface — `evals_get_record` for a single record, `evals_list_records` to browse. The resource is a convenience mirror for clients that support resources, not the access path.
## Features
Built on [`@cyanheads/mcp-ts-core`](https://www.npmjs.com/package/@cyanheads/mcp-ts-core):
- Declarative tool and resource definitions — single file per primitive, framework handles registration and validation
- Unified error handling — handlers throw, framework catches, classifies, and formats
- Typed error contracts — tools declare their domain failures (`reason` + recovery), surfaced to the agent
- Pluggable auth: `none`, `jwt`, `oauth`
- Structured logging with optional OpenTelemetry tracing
- STDIO and Streamable HTTP transports
Eval authoring:
- A `draft → review → surgical-revise → submit` loop, with the server acting as both scribe (normalize, persist, compile) and adversarial checker (run the record's own grader, reject what doesn't hold up)
- Records are a Zod `discriminatedUnion` keyed on `task_type` — `numeric`, `exact_answer`, `set_answer`, `mcq`, `regex_answer`, `json_answer`, `free_response`
- A typed grader DSL serialized with each record — deterministic kinds (`numeric` via math.js, `exact_match`, `set_match`, `regex`, `mcq`, `json_match`) run server-side; `llm_rubric` relies on recorded independent verification
- An enforced committability gate at submit: the gold must pass its own grader, ≥1 negative must be rejected, and a recorded decorrelated verification must agree with the gold
- Optional fleet grounding via the `captures` EvalsID field — link framework-written tool-call dumps, resolved from `EVALS_CAPTURE_DIR` and cross-checked against the gold (no server-to-server calls)
- Plain JSON files under `EVALS_DATA_DIR` — inspectable, diffable, version-controllable records
- Compile to JSONL, CSV, Inspect AI, and lm-evaluation-harness formats
Agent-friendly output:
- The two instructional tools (`evals_create_draft`, `evals_revise_draft`) carry the loop's review mechanism in their responses — the parsed record parroted back, a per-field review protocol, and a ready subagent prompt
- Self-consistency verdicts on every draft and revise — per-positive and per-negative results, not just a boolean
- `evals_list_records` discloses truncation when the limit is hit, so a partial set is never mistaken for the whole corpus
- The submit gate refuses with a typed `reason` + recovery hint, so a rejected record tells the agent exactly what to fix
## Getting started
Add the following to your MCP client configuration file. Set `EVALS_DATA_DIR` to a writable folder — the server manages `drafts/`, `submitted/`, and `eWhat people ask about evals-mcp-server
What is cyanheads/evals-mcp-server?
+
cyanheads/evals-mcp-server is mcp servers for the Claude AI ecosystem. Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP. It has 1 GitHub stars and its last recorded update is dated 2026-08-22.
How do I install evals-mcp-server?
+
You can install evals-mcp-server by cloning the repository (https://github.com/cyanheads/evals-mcp-server) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is cyanheads/evals-mcp-server safe to use?
+
Our security agent has analyzed cyanheads/evals-mcp-server and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains cyanheads/evals-mcp-server?
+
cyanheads/evals-mcp-server is maintained by cyanheads. The last recorded GitHub activity is dated 2026-08-22, with 5 open issues.
Are there alternatives to evals-mcp-server?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy evals-mcp-server to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/cyanheads-evals-mcp-server)<a href="https://claudewave.com/repo/cyanheads-evals-mcp-server"><img src="https://claudewave.com/api/badge/cyanheads-evals-mcp-server" alt="Featured on ClaudeWave: cyanheads/evals-mcp-server" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
The fastest path to AI-powered full stack observability, even for lean teams.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!