Skip to main content
ClaudeWave

Local MCP server exposing the Cloudflare Clef-Flash decision model to AI agents — structured decisions with probability outputs, fully offline via llama.cpp.

MCP ServersRegistry oficial1 estrellas0 forks● TypeScriptApache-2.0Actualizado today
ClaudeWave Trust Score
95/100
✓ Verified
Passed
  • ✓Open-source license (Apache-2.0)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/4/2026
Install in Claude Code / Claude Desktop
Method: NPX · clef-mcp
Claude Code CLI
claude mcp add clefmcp -- npx -y clef-mcp
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "clefmcp": {
      "command": "npx",
      "args": ["-y", "clef-mcp"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
Casos de uso

Resumen de MCP Servers

<div align="center">

<img src="./assets/demo.svg" alt="clef_decide — probability distributions for a production incident" width="880"/>

# clef-mcp

[![npm version](https://img.shields.io/npm/v/clef-mcp.svg)](https://www.npmjs.com/package/clef-mcp)
[![npm downloads](https://img.shields.io/npm/dm/clef-mcp.svg)](https://www.npmjs.com/package/clef-mcp)
[![CI](https://github.com/HighlyLoadedEgo/ClefMCP/actions/workflows/ci.yml/badge.svg)](https://github.com/HighlyLoadedEgo/ClefMCP/actions/workflows/ci.yml)
[![Official MCP Registry](https://img.shields.io/badge/MCP_Registry-io.github.HighlyLoadedEgo-blue)](https://registry.modelcontextprotocol.io)
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](LICENSE)

**A "reflex" for AI coding agents: structured decisions with probabilities — not prose.**

Local [MCP](https://modelcontextprotocol.io) server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the **Clef-Flash** decision model (9B, Apache-2.0 by Cloudflare) through a single tool: `clef_decide`. Pass a `state` and typed questions, get a **probability distribution over your options** in one forward pass. Fully local, offline, no tokens burned.

</div>

---

## Quick start

```bash
# 1. Detect hardware, download the model (~6 GB) + llama.cpp runtime, verify checksum + inference
npx clef-mcp install

# 2. Register the MCP server + agent skill in your clients (zcode, claude-code, codex, cursor)
npx clef-mcp setup

# 3. Run the MCP server (stdio)
npx clef-mcp
```

`clef-mcp` (step 3) never downloads anything. If the model is missing, tool calls return a structured `MODEL_NOT_INSTALLED` error with a hint. Steps 1 and 2 combine: `npx clef-mcp install --setup`.

## Why not just ask the LLM?

| | Chat LLM | `clef_decide` |
|---|---|---|
| Output | prose, you parse it | strict JSON: probability per option |
| Determinism | varies per run | single forward pass, no sampling |
| Latency (1 decision) | seconds of generation | ~0.5 s local |
| Context cost | grows with every decision | fixed, small schema |
| Privacy | depends on provider | 100% on-device, works offline |
| Calibration | vibes | softmax over trained option scores |

Sweet spot: **decision points inside agent loops** — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.

## The `clef_decide` tool

```json
{
  "state": {
    "task": "Fix failing tests",
    "error": "TypeError: Cannot read properties of undefined"
  },
  "questions": {
    "next_action": {
      "type": "choice",
      "instructions": "What should the coding agent do next?",
      "criteria": {
        "inspect": "Inspect the code and gather more information",
        "modify": "Modify the code",
        "test": "Run additional tests",
        "ask_user": "Ask the user for clarification"
      }
    },
    "confidence": {
      "type": "score",
      "instructions": "How confident are you in this decision?",
      "criteria": ["very_low", "low", "medium", "high", "very_high"]
    },
    "is_outage": { "type": "noul", "instructions": "Is a service down?" }
  }
}
```

| Type | Criteria | Answer |
|------|----------|--------|
| `choice` | map `option id → description`, or a plain list | probability per option |
| `score` | ordered list (index = score) | probability per level |
| `noul` | optional `{"true": "...", "false": "..."}` | `{"true": p, "false": 1-p}` |

Response — strictly structured, never prose:

```json
{
  "model": "clef-flash",
  "decisions": {
    "next_action": { "answer": { "inspect": 0.72, "modify": 0.12, "test": 0.14, "ask_user": 0.02 } },
    "confidence":  { "answer": { "very_low": 0.01, "low": 0.04, "medium": 0.18, "high": 0.61, "very_high": 0.16 } },
    "is_outage":   { "answer": { "true": 0.9, "false": 0.1 } }
  }
}
```

Batch up to **64 questions per call** — they are scored in one forward pass. `state` is treated strictly as **data**: never executed, never interpreted as instructions for the server.

## Prompts & resources

The server ships four MCP **prompts** (canned, decision-shaped asks — your client lists them via `prompts/list`):

| Prompt | Purpose |
|---|---|
| `incident-triage` | action + severity + user-impact questions for a production incident |
| `next-action` | what the coding agent should do next + confidence |
| `ticket-routing` | classify a message into a team + urgency |
| `security-review` | vulnerability yes/no, risk scale, first mitigation |

And three **resources** (read-only, no model needed):

| URI | Contents |
|---|---|
| `clef-mcp://capabilities` | live JSON: model, runtime, limits, error codes |
| `clef-mcp://evals/schema` | how to write eval cases |
| `clef-mcp://evals/dataset` | the bundled 30-case dataset |

## Measured, not marketed

Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:

| Scenario | Latency |
|---|---|
| Cold start (incl. model load, once per session) | ~4.4 s |
| 1 question | ~0.5 s |
| 10 questions, one call | ~2.9 s |
| 64 questions, one call | ~18.6 s |

Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — **86.7% pass** on the live model. Run it yourself: `clef-mcp evals`.

## Register with your MCP client

<details open>
<summary><b>Claude Code</b></summary>

```bash
claude mcp add clef-mcp -- clef-mcp
# or, without a global install:
claude mcp add clef-mcp -- npx -y clef-mcp
```
</details>

<details>
<summary><b>Codex</b> — <code>~/.codex/config.toml</code></summary>

```toml
[mcp_servers.clef-mcp]
command = "clef-mcp"
args = []
```
</details>

<details>
<summary><b>Cursor</b> — <code>.cursor/mcp.json</code></summary>

```json
{
  "mcpServers": {
    "clef-mcp": { "command": "clef-mcp", "args": [] }
  }
}
```
</details>

<details>
<summary><b>ZCode</b> — <code>~/.zcode/cli/config.json</code> (user scope, auto-connect)</summary>

```json
{
  "mcp": {
    "servers": {
      "clef-mcp": { "command": "clef-mcp", "args": [], "type": "stdio" }
    }
  }
}
```
</details>

Ready-made snippets: [`examples/`](examples/).

## Teach your agent (skill)

The schema tells the client *what* `clef_decide` accepts; agents also need to know *when* to reach for it and *how* to frame decisions. The bundled [`clef-decisions`](skills/clef-decisions/SKILL.md) skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:

```bash
clef-mcp setup                                                    # automatic
cp -r skills/clef-decisions ~/.agents/skills/                     # manual, from repo
cp -r "$(npm root -g)/clef-mcp/skills/clef-decisions" ~/.agents/skills/  # from npm package
```

## Architecture

```mermaid
flowchart LR
    subgraph clients [MCP clients]
        CC[Claude Code]
        CX[Codex]
        CU[Cursor]
        ZC[ZCode]
    end
    clients -- MCP stdio --> S[clef-mcp<br/>validation · limits · structured errors]
    S -- SystemOne adapter --> R[ClefRuntime<br/>llama.cpp subprocess<br/>127.0.0.1]
    R -- single forward pass --> M[("Clef-Flash<br/>9B · GGUF · local")]
    M -. probabilities .-> S -. strict JSON .-> clients
```

The `ClefRuntime` interface (`load / decide / unload / health`) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is `POST /v1/systemone` — the same contract across llama.cpp and other Clef runtimes.

## CLI

```bash
clef-mcp              # run the MCP server on stdio (default command)
clef-mcp install      # detect hardware → download model + runtime → verify checksum → verify inference
clef-mcp setup        # register the MCP server + agent skill in zcode / claude-code / codex / cursor
clef-mcp models       # list models/quantizations and install status
clef-mcp status       # runtime, model, memory summary
clef-mcp doctor       # full diagnosis (platform, RAM, GPU, binary, model, checksum*, inference, MCP config)
clef-mcp uninstall    # remove the model (and optionally the managed runtime)
clef-mcp evals        # run the evaluation dataset against the installed model
```

Flags: `install --quant Q8_0 --yes --skip-probe`, `install --setup`, `setup --clients zcode,cursor --no-skill`, `doctor --deep` (re-hash the model file), `uninstall --runtime --yes`.

## Runtimes: llama.cpp and MLX

Two local runtimes behind the same `ClefRuntime` interface:

| | `llama-cpp` (default) | `mlx` |
|---|---|---|
| Platforms | macOS, Linux, Windows | macOS / Apple Silicon only |
| Model | GGUF from `ggml-org/Clef-Flash-GGUF` | MLX 4-bit from `mlx-community/clef-flash-4bit` |
| Extras | none | [uv](https://docs.astral.sh/uv/) on PATH (managed Python env) |
| Install | `clef-mcp install` | `clef-mcp install --runtime mlx` |

Switch at runtime with `CLEF_RUNTIME=mlx` (must be set for the MCP server process — e.g. in the client's `env` block). Both speak the same `POST /v1/systemone` contract. The MLX snapshot is fetched into `CLEF_HOME` via a uv-managed `huggingface_hub` (no global Python state) at a pinned revision.

## Configuration

| Variable | Default | Meaning |
|----------|---------|---------|
| `CLEF_MODEL` | `clef-flash` | Model id (per-call `model` also accepted) |
| `CLEF_HOME` | `~/.cache/clef-mcp` | Cache/model home |
| `CLEF_RUNTIME` | `llama-cpp` | `llama-cpp` \| `mlx` |
| `CLEF_LOG_LEVEL` | `error` | `error` \| `warn` \| `info` \| `debug` (stderr only) |
| `CLEF_LLAMA_BIN` | – | Explicit `llama-server` binary path (llama-cpp runtime) |
| `CLEF_LLAMA_RELEASE_TAG` | latest nightly | Pin the managed llama.cpp build |
| `CLEF_LLAMA_BATCH` | `8192` | llama.cpp physical batch (multi-question requests) |
| `CLEF_MLX_UV` | `uv` on PATH | Explicit uv binary (MLX runtime) |
| `CLEF_MAX_QUESTIONS` | `64` | Max questions per call |
| `CLEF_MAX_STATE_BYTES` | `1048576` | Max serialized `state` size |
| `CLEF_MAX_INSTRUCTION_CHARS` | `10000` | Max chars per question instructions |

Runtime resolution: `CLEF_LLAMA_BIN` → managed binary in `CLEF_HOME/runt
ai-agentsclefdecision-modelllama-cpplocal-llmmcpmodel-context-protocoloffline-inferencestructured-output

Lo que la gente pregunta sobre ClefMCP

¿Qué es HighlyLoadedEgo/ClefMCP?

+

HighlyLoadedEgo/ClefMCP es mcp servers para el ecosistema de Claude AI. Local MCP server exposing the Cloudflare Clef-Flash decision model to AI agents — structured decisions with probability outputs, fully offline via llama.cpp. Tiene 1 estrellas en GitHub y su última actualización registrada es del 2026-10-03.

¿Cómo se instala ClefMCP?

+

Puedes instalar ClefMCP clonando el repositorio (https://github.com/HighlyLoadedEgo/ClefMCP) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.

¿Es seguro usar HighlyLoadedEgo/ClefMCP?

+

Nuestro agente de seguridad ha analizado HighlyLoadedEgo/ClefMCP y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.

¿Quién mantiene HighlyLoadedEgo/ClefMCP?

+

HighlyLoadedEgo/ClefMCP es mantenido por HighlyLoadedEgo. La última actividad registrada en GitHub es del 2026-10-03, con 1 issues abiertos.

¿Hay alternativas a ClefMCP?

+

Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.

Despliega ClefMCP en tu cloud

Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.

¿Mantienes este repo? Añade un badge a tu README

Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.

Featured on ClaudeWave: HighlyLoadedEgo/ClefMCP
[![Featured on ClaudeWave](https://claudewave.com/api/badge/highlyloadedego-clefmcp)](https://claudewave.com/repo/highlyloadedego-clefmcp)
<a href="https://claudewave.com/repo/highlyloadedego-clefmcp"><img src="https://claudewave.com/api/badge/highlyloadedego-clefmcp" alt="Featured on ClaudeWave: HighlyLoadedEgo/ClefMCP" width="320" height="64" /></a>

Más MCP Servers

Alternativas a ClefMCP