MIMIR is a non-generative decision model: give it a context, a question and the options, and it answers with a typed, calibrated decision, the evidence behind it, and a certified signal for when to act — or to escalate instead.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Documented (README)
git clone https://github.com/abderahmane-ai/mimirResumen de Tools
# mimir-decisions
**Decisions your agents can act on.** MIMIR is a non-generative decision model: give it
a context, a question and the options, and get back a typed answer with calibrated
probabilities, the evidence behind it, and a certified verdict on whether to act or
escalate. No text generated. Nothing to parse. Nothing to hallucinate.
It beats Laya and GLiNER2.5-Decide head-to-head on six of ten tasks — by 54.7 points on
Banking77, 42.3 on MASSIVE, 36.3 on typed decisions — and where it cannot back an
answer, it abstains instead of guessing.
```bash
pip install "mimir-decisions[local]" # CPU engine
pip install "mimir-decisions[local-gpu]" # CUDA engine
pip install mimir-decisions # data models and HTTP client only
```
Python 3.11+. Documentation: <https://abderahmane-ai.github.io/mimir/>
---
## Why MIMIR
Most agents route, classify, and verify using a general-purpose language model: slow, expensive, and impossible to audit. MIMIR is built for structured decisions. It runs on ONNX Runtime in milliseconds, returns calibrated probabilities with every answer, and issues a mathematical certificate — a formal guarantee that its realised error rate stays at or below the risk level you ask for, measured on held-out data.
- **No generation.** Answers are drawn from the options you supply, not synthesised. The model cannot hallucinate an answer that wasn't on the list.
- **Calibrated confidence.** Probabilities are not softmax scores; they are calibrated to match realised accuracy on held-out data.
- **Certified deferral.** When confidence falls short of the certified threshold, the decision defers rather than guessing. The coverage and the error rate of taken decisions are proven.
- **One typed contract.** Seven decision types — choice, multi-choice, yes/no, verify, rank, rate, estimate — all returning the same result shape, over any context.
- **Portable.** The same Python interface works locally on CPU or GPU, over HTTP, and over MCP. Framework adapters exist for eight agent SDKs.
---
## Quickstart
```python
from mimir import Mimir
model = Mimir.from_pretrained("Mythologic/MIMIR-1")
result = model.choose(
"My card was charged twice for the same order.",
"Which team should handle this ticket?",
options={"billing": "Billing: payments, refunds", "security": "Security: account access"},
)
result.status # Status.DECIDED, Status.ABSTAINED or Status.DEFERRED
result.answer # an option id, or None when no option applies
result.probabilities # calibrated probability of each option id
result.certificate # the certified threshold the decision was checked against
```
`answer` is the model's prediction. `status` is the policy's verdict:
- `DECIDED` — the answer is an option and is certified at the requested risk level.
- `ABSTAINED` — no listed option applies, and that is certified.
- `DEFERRED` — not certified; `result.deferral.reason` is `below_threshold`, `out_of_distribution`, or `no_certified_threshold`.
The first call downloads the model from the Hugging Face Hub at the revision this package version pins, verifies its Sigstore signature, checks every file against the manifest's SHA-256, and loads it.
---
## Decision types
| Spec | Answer |
|---|---|
| `Choice(question, options)` | an option id, or `None` |
| `MultiChoice(question, options)` | the option ids that apply |
| `YesNo(question)` | `True` or `False` |
| `Verify(claim)` | `supported`, `contradicted`, or `not_enough_information` |
| `Rank(question, candidates)` | candidate ids, best first |
| `Rate(question, levels)` | a level id; levels given lowest first |
| `Estimate(question, low, high, unit)` | a number in `[low, high]`, with a confidence interval |
```python
from mimir import Context, Field, Passage, Rate, Table
context = Context(
passages=[Passage(title="Ticket #4412", text="The export has failed every night this week.")],
tables=[Table.from_rows([["2026-03-02", "failed"]], header=["date", "status"])],
fields=Field.from_json({"customer": {"plan": "enterprise", "seats": 240}}),
)
result = model.decide(context, Rate("How urgent is this?", ["low", "medium", "high"]), risk=0.01)
```
A context can be a string, a list of strings, a dict read as a JSON state, or a `Context` of typed passages, tables, and fields. `Table.from_dataframe(frame)` reads a pandas or polars DataFrame. `decide_many` batches multiple decisions, and every method has an async counterpart (`adecide`, `adecide_many`, …).
---
## Certification
`decide` takes a risk level certified by the loaded policy (`model.info().risk_levels`). A decision is taken only when its calibrated confidence clears a threshold certified on held-out data to keep the realised error rate at or below that risk with 95% confidence, and when the context passes the out-of-distribution gate. `decide_uncertified` returns the raw model answer with no policy applied.
A certificate covers one exact configuration: model files, variant, ONNX Runtime version, execution provider, and options. On hardware not listed in the certificate, the first load runs the release's equivalence set and requires every decision to match. To certify thresholds on your own labelled data:
```bash
mimir calibrate labelled.jsonl --risk 0.01 --confidence 0.95 --out policy.json
```
```python
model = Mimir.from_pretrained("Mythologic/MIMIR-1", policy="policy.json")
```
---
## Remote use
```python
from mimir import MimirClient
remote = MimirClient("https://mimir.internal", api_key="...")
remote.choose("...", "Which team?", options=["billing", "security"])
```
`MimirClient` has the same interface as `Mimir`, so all code, decision tools, and framework adapters accept either. It requires only the base install. Connection errors, timeouts, and 429/502/503/504/529 responses are retried with exponential backoff that honours `Retry-After`.
---
## Decision tools
```python
from mimir import Choice
route_ticket = model.tool(
"route_ticket",
Choice("Which team should handle this ticket?", options=["billing", "security"]),
description="Route a support ticket to the team that owns it.",
)
route_ticket("My card was charged twice")
route_ticket.input_schema, route_ticket.output_schema
```
Tools can also be declared in a YAML file, which the HTTP and MCP servers load:
```yaml
tools:
- name: route_ticket
description: Route a support ticket to the team that owns it.
decision:
type: choice
question: Which team should handle this ticket?
options: [billing, security]
```
---
## Tool-call checks
A tool-call check decides, against rules you write, whether an agent's pending tool call may run. A certified yes allows it, a certified no denies it, and anything else escalates to a person.
```python
check = model.tool_call_check(
["Refunds above 500 dollars need a manager's approval."], tools=["issue_refund"]
)
outcome = check("issue_refund", {"order": "4412", "amount": 900})
outcome.permission # Permission.ALLOW, Permission.DENY or Permission.ESCALATE
outcome.reason # one sentence for the agent or the approver
```
---
## Agent frameworks
Each adapter turns decision tools into the framework's native tool type and wires a tool-call check into that framework's own approval hook.
| Framework | Install | Tools | Tool-call check |
|---|---|---|---|
| OpenAI Agents SDK | `mimir-decisions[openai-agents]` | `as_function_tool` | `guard`: escalations pause the run for approval |
| LangChain / LangGraph | `mimir-decisions[langchain]` | `as_structured_tool` | `ToolCallCheckMiddleware`: escalations interrupt with the human-in-the-loop request |
| PydanticAI | `mimir-decisions[pydantic-ai]` | `as_toolset` | `guard`: escalations end the run with `DeferredToolRequests` |
| CrewAI | `mimir-decisions[crewai]` | `as_crewai_tool` | `tool_call_hook`: escalations go to your approver |
| Google ADK | `mimir-decisions[adk]` | `as_adk_tool` | `tool_call_callback`: escalations ask for ADK confirmation |
| Microsoft Agent Framework | `mimir-decisions[agent-framework]` | `as_function_tool` | `ToolCallCheckMiddleware`: only certified calls run |
| LlamaIndex | `mimir-decisions[llamaindex]` | `as_llamaindex_tool` | none |
| smolagents | `mimir-decisions[smolagents]` | `as_smolagents_tool` | none |
```python
from agents import Agent
from mimir.integrations.openai_agents import as_function_tool
agent = Agent(name="support", tools=[as_function_tool(route_ticket)])
```
Every framework also reaches MIMIR through its own MCP client. [`examples/`](examples) has a native, an MCP, and a checked agent for each framework, plus a Vercel AI SDK agent in TypeScript.
---
## HTTP server
```bash
pip install "mimir-decisions[local,server]"
MIMIR_API_KEYS=key-one,key-two mimir serve --host 0.0.0.0 --tools tools.yaml
```
| Route | Does |
|---|---|
| `POST /v1/decide` | one certified decision: `{context, decision, risk, alpha}` |
| `POST /v1/decide/uncertified` | the model's raw answer: `{context, decision}` |
| `POST /v1/decide/batch` | up to 64 decisions in one call |
| `POST /v1/tools/{name}` | a tool from `--tools`, given only `{context}` |
| `POST /v1/systemone` | Jev's request and response format |
| `GET /v1/models` | model, revision, runtime and certified risk levels |
| `GET /healthz`, `GET /readyz` | liveness, and readiness once the model is loaded |
| `GET /metrics` | Prometheus metrics |
Concurrent requests are batched. With keys in `MIMIR_API_KEYS`, every route except the probes requires `Authorization: Bearer <key>`. A server with no keys listens only on loopback unless started with `--allow-no-auth`. The OpenAPI 3.1 document is [`openapi.json`](openapi.json).
---
## MCP server
Each configured tool becomes an MCP tool that takes only a context; `--generic-tools` adds `mimir_choose`, `mimir_verify`, `mimir_rank`, and `mimir_rate`. A deferred decision is a normal result telling the agent to escalate.
```bash
uvx --from "mimir-decisions[locaLo que la gente pregunta sobre mimir
¿Qué es abderahmane-ai/mimir?
+
abderahmane-ai/mimir es tools para el ecosistema de Claude AI. MIMIR is a non-generative decision model: give it a context, a question and the options, and it answers with a typed, calibrated decision, the evidence behind it, and a certified signal for when to act — or to escalate instead. Tiene 1 estrellas en GitHub y su última actualización registrada es del 2026-09-28.
¿Cómo se instala mimir?
+
Puedes instalar mimir clonando el repositorio (https://github.com/abderahmane-ai/mimir) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar abderahmane-ai/mimir?
+
Nuestro agente de seguridad ha analizado abderahmane-ai/mimir y le ha asignado un Trust Score de 87/100 (tier: Trusted). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene abderahmane-ai/mimir?
+
abderahmane-ai/mimir es mantenido por abderahmane-ai. La última actividad registrada en GitHub es del 2026-09-28, con 0 issues abiertos.
¿Hay alternativas a mimir?
+
Sí. En ClaudeWave puedes explorar tools similares en /categories/tools, ordenados por popularidad o actividad reciente.
Despliega mimir en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/abderahmane-ai-mimir)<a href="https://claudewave.com/repo/abderahmane-ai-mimir"><img src="https://claudewave.com/api/badge/abderahmane-ai-mimir" alt="Featured on ClaudeWave: abderahmane-ai/mimir" width="320" height="64" /></a>Más Tools
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
An AI skill that provides design intelligence for building professional UI/UX across multiple platforms.
🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
Use Claude Code, Codex, VSCode, Pi, and OpenCode (and 6 other harnesses) for free (1.3B+ free tokens) from your terminal, app, IDE, or phone, and now from the browser with native browser sessions (multi-harness + multi-model) like OpenClaw (voice supported + ToS friendly)