Local MCP server exposing the Cloudflare Clef-Flash decision model to AI agents — structured decisions with probability outputs, fully offline via llama.cpp.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add clefmcp -- npx -y clef-mcp{
"mcpServers": {
"clefmcp": {
"command": "npx",
"args": ["-y", "clef-mcp"]
}
}
}Resumen de MCP Servers
<div align="center">
<img src="./assets/demo.svg" alt="clef_decide — probability distributions for a production incident" width="880"/>
# clef-mcp
[](https://www.npmjs.com/package/clef-mcp)
[](https://www.npmjs.com/package/clef-mcp)
[](https://github.com/HighlyLoadedEgo/ClefMCP/actions/workflows/ci.yml)
[](https://registry.modelcontextprotocol.io)
[](LICENSE)
**A "reflex" for AI coding agents: structured decisions with probabilities — not prose.**
Local [MCP](https://modelcontextprotocol.io) server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the **Clef-Flash** decision model (9B, Apache-2.0 by Cloudflare) through a single tool: `clef_decide`. Pass a `state` and typed questions, get a **probability distribution over your options** in one forward pass. Fully local, offline, no tokens burned.
</div>
---
## Quick start
```bash
# 1. Detect hardware, download the model (~6 GB) + llama.cpp runtime, verify checksum + inference
npx clef-mcp install
# 2. Register the MCP server + agent skill in your clients (zcode, claude-code, codex, cursor)
npx clef-mcp setup
# 3. Run the MCP server (stdio)
npx clef-mcp
```
`clef-mcp` (step 3) never downloads anything. If the model is missing, tool calls return a structured `MODEL_NOT_INSTALLED` error with a hint. Steps 1 and 2 combine: `npx clef-mcp install --setup`.
## Why not just ask the LLM?
| | Chat LLM | `clef_decide` |
|---|---|---|
| Output | prose, you parse it | strict JSON: probability per option |
| Determinism | varies per run | single forward pass, no sampling |
| Latency (1 decision) | seconds of generation | ~0.5 s local |
| Context cost | grows with every decision | fixed, small schema |
| Privacy | depends on provider | 100% on-device, works offline |
| Calibration | vibes | softmax over trained option scores |
Sweet spot: **decision points inside agent loops** — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.
## The `clef_decide` tool
```json
{
"state": {
"task": "Fix failing tests",
"error": "TypeError: Cannot read properties of undefined"
},
"questions": {
"next_action": {
"type": "choice",
"instructions": "What should the coding agent do next?",
"criteria": {
"inspect": "Inspect the code and gather more information",
"modify": "Modify the code",
"test": "Run additional tests",
"ask_user": "Ask the user for clarification"
}
},
"confidence": {
"type": "score",
"instructions": "How confident are you in this decision?",
"criteria": ["very_low", "low", "medium", "high", "very_high"]
},
"is_outage": { "type": "noul", "instructions": "Is a service down?" }
}
}
```
| Type | Criteria | Answer |
|------|----------|--------|
| `choice` | map `option id → description`, or a plain list | probability per option |
| `score` | ordered list (index = score) | probability per level |
| `noul` | optional `{"true": "...", "false": "..."}` | `{"true": p, "false": 1-p}` |
Response — strictly structured, never prose:
```json
{
"model": "clef-flash",
"decisions": {
"next_action": { "answer": { "inspect": 0.72, "modify": 0.12, "test": 0.14, "ask_user": 0.02 } },
"confidence": { "answer": { "very_low": 0.01, "low": 0.04, "medium": 0.18, "high": 0.61, "very_high": 0.16 } },
"is_outage": { "answer": { "true": 0.9, "false": 0.1 } }
}
}
```
Batch up to **64 questions per call** — they are scored in one forward pass. `state` is treated strictly as **data**: never executed, never interpreted as instructions for the server.
## Prompts & resources
The server ships four MCP **prompts** (canned, decision-shaped asks — your client lists them via `prompts/list`):
| Prompt | Purpose |
|---|---|
| `incident-triage` | action + severity + user-impact questions for a production incident |
| `next-action` | what the coding agent should do next + confidence |
| `ticket-routing` | classify a message into a team + urgency |
| `security-review` | vulnerability yes/no, risk scale, first mitigation |
And three **resources** (read-only, no model needed):
| URI | Contents |
|---|---|
| `clef-mcp://capabilities` | live JSON: model, runtime, limits, error codes |
| `clef-mcp://evals/schema` | how to write eval cases |
| `clef-mcp://evals/dataset` | the bundled 30-case dataset |
## Measured, not marketed
Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:
| Scenario | Latency |
|---|---|
| Cold start (incl. model load, once per session) | ~4.4 s |
| 1 question | ~0.5 s |
| 10 questions, one call | ~2.9 s |
| 64 questions, one call | ~18.6 s |
Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — **86.7% pass** on the live model. Run it yourself: `clef-mcp evals`.
## Register with your MCP client
<details open>
<summary><b>Claude Code</b></summary>
```bash
claude mcp add clef-mcp -- clef-mcp
# or, without a global install:
claude mcp add clef-mcp -- npx -y clef-mcp
```
</details>
<details>
<summary><b>Codex</b> — <code>~/.codex/config.toml</code></summary>
```toml
[mcp_servers.clef-mcp]
command = "clef-mcp"
args = []
```
</details>
<details>
<summary><b>Cursor</b> — <code>.cursor/mcp.json</code></summary>
```json
{
"mcpServers": {
"clef-mcp": { "command": "clef-mcp", "args": [] }
}
}
```
</details>
<details>
<summary><b>ZCode</b> — <code>~/.zcode/cli/config.json</code> (user scope, auto-connect)</summary>
```json
{
"mcp": {
"servers": {
"clef-mcp": { "command": "clef-mcp", "args": [], "type": "stdio" }
}
}
}
```
</details>
Ready-made snippets: [`examples/`](examples/).
## Teach your agent (skill)
The schema tells the client *what* `clef_decide` accepts; agents also need to know *when* to reach for it and *how* to frame decisions. The bundled [`clef-decisions`](skills/clef-decisions/SKILL.md) skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:
```bash
clef-mcp setup # automatic
cp -r skills/clef-decisions ~/.agents/skills/ # manual, from repo
cp -r "$(npm root -g)/clef-mcp/skills/clef-decisions" ~/.agents/skills/ # from npm package
```
## Architecture
```mermaid
flowchart LR
subgraph clients [MCP clients]
CC[Claude Code]
CX[Codex]
CU[Cursor]
ZC[ZCode]
end
clients -- MCP stdio --> S[clef-mcp<br/>validation · limits · structured errors]
S -- SystemOne adapter --> R[ClefRuntime<br/>llama.cpp subprocess<br/>127.0.0.1]
R -- single forward pass --> M[("Clef-Flash<br/>9B · GGUF · local")]
M -. probabilities .-> S -. strict JSON .-> clients
```
The `ClefRuntime` interface (`load / decide / unload / health`) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is `POST /v1/systemone` — the same contract across llama.cpp and other Clef runtimes.
## CLI
```bash
clef-mcp # run the MCP server on stdio (default command)
clef-mcp install # detect hardware → download model + runtime → verify checksum → verify inference
clef-mcp setup # register the MCP server + agent skill in zcode / claude-code / codex / cursor
clef-mcp models # list models/quantizations and install status
clef-mcp status # runtime, model, memory summary
clef-mcp doctor # full diagnosis (platform, RAM, GPU, binary, model, checksum*, inference, MCP config)
clef-mcp uninstall # remove the model (and optionally the managed runtime)
clef-mcp evals # run the evaluation dataset against the installed model
```
Flags: `install --quant Q8_0 --yes --skip-probe`, `install --setup`, `setup --clients zcode,cursor --no-skill`, `doctor --deep` (re-hash the model file), `uninstall --runtime --yes`.
## Runtimes: llama.cpp and MLX
Two local runtimes behind the same `ClefRuntime` interface:
| | `llama-cpp` (default) | `mlx` |
|---|---|---|
| Platforms | macOS, Linux, Windows | macOS / Apple Silicon only |
| Model | GGUF from `ggml-org/Clef-Flash-GGUF` | MLX 4-bit from `mlx-community/clef-flash-4bit` |
| Extras | none | [uv](https://docs.astral.sh/uv/) on PATH (managed Python env) |
| Install | `clef-mcp install` | `clef-mcp install --runtime mlx` |
Switch at runtime with `CLEF_RUNTIME=mlx` (must be set for the MCP server process — e.g. in the client's `env` block). Both speak the same `POST /v1/systemone` contract. The MLX snapshot is fetched into `CLEF_HOME` via a uv-managed `huggingface_hub` (no global Python state) at a pinned revision.
## Configuration
| Variable | Default | Meaning |
|----------|---------|---------|
| `CLEF_MODEL` | `clef-flash` | Model id (per-call `model` also accepted) |
| `CLEF_HOME` | `~/.cache/clef-mcp` | Cache/model home |
| `CLEF_RUNTIME` | `llama-cpp` | `llama-cpp` \| `mlx` |
| `CLEF_LOG_LEVEL` | `error` | `error` \| `warn` \| `info` \| `debug` (stderr only) |
| `CLEF_LLAMA_BIN` | – | Explicit `llama-server` binary path (llama-cpp runtime) |
| `CLEF_LLAMA_RELEASE_TAG` | latest nightly | Pin the managed llama.cpp build |
| `CLEF_LLAMA_BATCH` | `8192` | llama.cpp physical batch (multi-question requests) |
| `CLEF_MLX_UV` | `uv` on PATH | Explicit uv binary (MLX runtime) |
| `CLEF_MAX_QUESTIONS` | `64` | Max questions per call |
| `CLEF_MAX_STATE_BYTES` | `1048576` | Max serialized `state` size |
| `CLEF_MAX_INSTRUCTION_CHARS` | `10000` | Max chars per question instructions |
Runtime resolution: `CLEF_LLAMA_BIN` → managed binary in `CLEF_HOME/runtLo que la gente pregunta sobre ClefMCP
¿Qué es HighlyLoadedEgo/ClefMCP?
+
HighlyLoadedEgo/ClefMCP es mcp servers para el ecosistema de Claude AI. Local MCP server exposing the Cloudflare Clef-Flash decision model to AI agents — structured decisions with probability outputs, fully offline via llama.cpp. Tiene 1 estrellas en GitHub y su última actualización registrada es del 2026-10-03.
¿Cómo se instala ClefMCP?
+
Puedes instalar ClefMCP clonando el repositorio (https://github.com/HighlyLoadedEgo/ClefMCP) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar HighlyLoadedEgo/ClefMCP?
+
Nuestro agente de seguridad ha analizado HighlyLoadedEgo/ClefMCP y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene HighlyLoadedEgo/ClefMCP?
+
HighlyLoadedEgo/ClefMCP es mantenido por HighlyLoadedEgo. La última actividad registrada en GitHub es del 2026-10-03, con 1 issues abiertos.
¿Hay alternativas a ClefMCP?
+
Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.
Despliega ClefMCP en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/highlyloadedego-clefmcp)<a href="https://claudewave.com/repo/highlyloadedego-clefmcp"><img src="https://claudewave.com/api/badge/highlyloadedego-clefmcp" alt="Featured on ClaudeWave: HighlyLoadedEgo/ClefMCP" width="320" height="64" /></a>Más MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.