Ekbasis: an open world model for agents. What an action will do, before it is done.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
/plugin marketplace add OpenInterpretability/ekbasis
/plugin install ekbasisResumen de Plugins
# Ekbasis  **Papers:** [Look When Unsure, Check When Sure](https://doi.org/10.5281/zenodo.23146970) · DOI [10.5281/zenodo.23146970](https://doi.org/10.5281/zenodo.23146970) · [When Does a Consequence Model Make AI Agents Safer?](https://doi.org/10.5281/zenodo.23197341) · DOI [10.5281/zenodo.23197341](https://doi.org/10.5281/zenodo.23197341) (data, scripts and the benchmark to rerun its studies: [paper/agents](https://github.com/OpenInterpretability/ekbasis/tree/main/paper/agents)) · **Code:** [github.com/OpenInterpretability/ekbasis](https://github.com/OpenInterpretability/ekbasis) · **Site:** [openinterp.org/ekbasis](https://openinterp.org/ekbasis) **Ekbasis** (ἔκβασις, *"how an action turns out"*) is an open **consequence model** for agents: given the current state and an action, it answers typed questions about what will happen — *will this command lose work? will it fail? what will X be afterwards?* — with a **calibrated probability**, in **one forward pass** (no generated text). It is the consequence layer of the Eikos family (Eikos decides; Ekbasis foresees), trained on synthetic worlds and on **real executions** of git commands in throwaway repositories, so its answers come from what actually happened, not from what an agent believes. > Every number on this page was measured on the exact release weights. [RELEASE_EVAL.md](https://github.com/OpenInterpretability/ekbasis/blob/main/RELEASE_EVAL.md) has the > results, every prediction, how this checkpoint was chosen, which measurements were pre-registered > ([PREREG_release_eval.md](https://github.com/OpenInterpretability/ekbasis/blob/main/PREREG_release_eval.md), [paper/prereg/](https://github.com/OpenInterpretability/ekbasis/tree/main/paper/prereg/)) and the deviations. ## Install The client needs a server: the hosted API (`EKBASIS_URL=https://openinterp.org/api/v1` and `EKBASIS_API_KEY=ekb_…`, key at [openinterp.org/console](https://openinterp.org/console)) or your own, with the open weights ([Quick start](#quick-start)). - **Command line and Python** (standard library only, Python ≥ 3.9): `ekbasis`, `ekbasis-claude-hook`, and `ekbasis-mcp` with the `mcp` extra. ```bash pip install ekbasis # or without installing: uvx ekbasis health pip install "ekbasis[mcp]" # with the MCP server (Python >= 3.10) pip install "git+https://github.com/OpenInterpretability/ekbasis" # from source ``` - **Claude Code plugin**: the git guard hook, the `ekbasis-guard` skill, the MCP server and the `ekbasis` command on the Bash tool's PATH, in one install: ```bash claude plugin marketplace add OpenInterpretability/ekbasis --sparse .claude-plugin plugins claude plugin install ekbasis@ekbasis ``` or, inside Claude Code, `/plugin marketplace add OpenInterpretability/ekbasis` then `/plugin install ekbasis@ekbasis` (without `--sparse` the marketplace step clones the whole repository, about 86 MB with the results and papers, once; with it, about 1 MB). The plugin (`plugins/ekbasis/`, about 0.5 MB) runs its own copy of the client, so it needs `python3` (≥ 3.9) and `sh` for the hooks and [`uv`](https://docs.astral.sh/uv/) for the MCP server, and nothing from PyPI. Set `EKBASIS_API_KEY` (a key alone means the hosted API) or `EKBASIS_URL`, in your shell or in the `"env"` block of `~/.claude/settings.json`, and restart Claude Code. Until one of them is set the guard is off: each session starts with one line saying how to set it up, and `ekbasis` says it is not set up (exit 3). Once set, the hook is the same fail-closed hook described in [Claude Code and MCP](#claude-code-and-mcp), and a hook that cannot run (`python3` missing or older than 3.9, an error in the hook, no answer within 27 s) asks you to confirm, saying why, instead of letting the command through. If you added `ekbasis-claude-hook` to `settings.json` by hand, remove it, or every line is checked twice. - **MCP server** for any MCP client (Python ≥ 3.10): ```bash uvx --python ">=3.10" --with "mcp>=1.2" ekbasis mcp # same as ekbasis-mcp from pip install "ekbasis[mcp]" ``` <!-- mcp-name: io.github.OpenInterpretability/ekbasis --> ## What kind of model is it? **Not a model that thinks. Not a model that judges. A model that foresees.** Ekbasis is a **world model for agents** in the precise sense: an action-conditioned model of how the state changes. It answers *what will be, if I do this*, as calibrated distributions over the next state's variables, in one forward pass. | | LLM, reasoning | System One (Eikos, Jev) | **Ekbasis** | |---|---|---|---| | It answers | anything, in text | what is: which option holds now | **what will be, if I do this** | | Kind of question | open | a judgment of the present | **an intervention: the outcome of acting** | | Output | generated text | calibrated distribution over the options | **calibrated distribution over the next state's variables** | | Learns from | human and model text | a teacher's judgments | **what actually happened when the action ran** | | Time | seconds to minutes of reasoning | one forward pass | **one forward pass per step; chains to long sequences** | | Role in an agent | plans and talks | judge | **simulator and guard: it foresees** | Two ways to picture it: - **The forward model the agents were missing.** Before you move your arm, the brain predicts the consequence of the motor command; that is what lets us act fast and correct before an error. One part plans, another predicts. The agent plans; Ekbasis predicts. - **The world-model module, made real.** LeCun's architecture for autonomous machine intelligence separates a world model (it predicts the next state given an action) from the actor and the critic. The LLM agent is the actor, AgentGuard the critic, Ekbasis the world model. It was trained with a JEPA-style loss that aligns its internal state with the true outcome. It is derived from an LLM (Qwen3.8-27B, through Eikos-27B) but does not work as one: its training objective and its readout make it answer with probabilities, not text. ## What it is for - **A check before acting**: ~0.1 s per question on one GPU, calibrated, so an agent can check *every* action and escalate only the risky or uncertain ones. - **A git guard** for coding agents: the state of your real repository is described in the format the model was trained on (including what decides a conflict), and the commands are checked before they run. - **Simulation and state tracking**: chain it one action at a time to follow long sequences; given a way to read the real state, it looks only when unsure (see [Look when unsure](#look-when-unsure-predict-observe-correct)). - **A fast first opinion before expensive reasoning**: answer from Ekbasis when it is confident, escalate to a large reasoning model when it is not. ## Quick start 1. **Serve the model** (one GPU with ~80 GB; vLLM ≥ 0.30). The model folder ships its serving code; run it in the Python environment where vLLM is installed: ```bash hf download caiovicentino1/Ekbasis-27B --local-dir Ekbasis-27B bash Ekbasis-27B/serve_vllm.sh Ekbasis-27B 8001 # vLLM on 127.0.0.1:8001 python Ekbasis-27B/serve.py --model Ekbasis-27B --vllm-url http://127.0.0.1:8001 --port 8000 ``` 2. **Install the client** (standard library only, Python ≥ 3.9), anywhere that can reach the server: ```bash pip install "git+https://github.com/OpenInterpretability/ekbasis" # ekbasis, ekbasis-claude-hook (MCP: below) export EKBASIS_URL=http://127.0.0.1:8000 ekbasis health ``` 3. **Check git commands before running them** (in any repository): ```bash ekbasis git-check -- "git checkout -- app.py" # Ekbasis: RISKY (lose uncommitted work: 99%) # 0% fails git checkout -- app.py # - may permanently lose uncommitted work (99%) ``` Exit code 0 when no risk is found, 2 when the commands may lose uncommitted work, 3 when the guard cannot foresee (the server cannot be reached or does not answer in time, the repository cannot be read, or a command points git at another repository or uses an alias): treat 3 as risky. 1 is a usage error. The guard fails closed; `--fail-open` turns "cannot foresee" into 0 with a warning. It is a warning layer that can be wrong, not a security boundary: keep confirmations, backups and least privilege ([docs/SECURITY.md](https://github.com/OpenInterpretability/ekbasis/blob/main/docs/SECURITY.md)). ## Builds Every build that works is released, so each machine runs the one that fits; each build's card shows its quality against bf16 on the same evaluation (the gate was pre-registered in `PREREG_quantized.md`). | Build | Size | Runs on | Speed (one RTX PRO 6000) | |---|---|---|---| | [Ekbasis-27B](https://huggingface.co/caiovicentino1/Ekbasis-27B) (bf16) | 55 GB | GPUs with 80 GB | the reference | | [Ekbasis-27B-FP8](https://huggingface.co/caiovicentino1/Ekbasis-27B-FP8) | 30.4 GB | GPUs with 48 GB; fastest with FP8 kernels (Ada, Hopper, Blackwell) | 1.6× the throughput of bf16 | | [Ekbasis-27B-INT4](https://huggingface.co/caiovicentino1/Ekbasis-27B-INT4) | 18.6 GB | GPUs with 32 GB; 24 GB with text only and a 4k context (command in its card) | about bf16's | | [Ekbasis-27B-MLX-4bit](https://huggingface.co/caiovicentino1/Ekbasis-27B-MLX-4bit) | 15 GB | Macs with Apple Silicon (32 GB or more recommended) | — | GPU sizes were tested by limiting vLLM to that much memory on one RTX PRO 6000 and checking that the answers match the build at full memory; the MLX build's quality gate ran with MLX on a Linux GPU, and it was then tested on an Apple-silicon Mac (its card, "On a Mac: tested"). ## Python ```python from ekbasis impo
Lo que la gente pregunta sobre ekbasis
¿Qué es OpenInterpretability/ekbasis?
+
OpenInterpretability/ekbasis es plugins para el ecosistema de Claude AI. Ekbasis: an open world model for agents. What an action will do, before it is done. Tiene 18 estrellas en GitHub y su última actualización registrada es del 2026-10-11.
¿Cómo se instala ekbasis?
+
Puedes instalar ekbasis clonando el repositorio (https://github.com/OpenInterpretability/ekbasis) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar OpenInterpretability/ekbasis?
+
Nuestro agente de seguridad ha analizado OpenInterpretability/ekbasis y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene OpenInterpretability/ekbasis?
+
OpenInterpretability/ekbasis es mantenido por OpenInterpretability. La última actividad registrada en GitHub es del 2026-10-11, con 3 issues abiertos.
¿Hay alternativas a ekbasis?
+
Sí. En ClaudeWave puedes explorar plugins similares en /categories/plugins, ordenados por popularidad o actividad reciente.
Despliega ekbasis en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/openinterpretability-ekbasis)<a href="https://claudewave.com/repo/openinterpretability-ekbasis"><img src="https://claudewave.com/api/badge/openinterpretability-ekbasis" alt="Featured on ClaudeWave: OpenInterpretability/ekbasis" width="320" height="64" /></a>Más Plugins
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
Write HTML. Render video. Built for agents.
Agent skill that removes signs of AI-generated writing from text
Academic Research Skills for Claude Code: research → write → review → revise → finalize
Editorial diagram design for Claude Code, Codex, GitHub Copilot, Cursor, Factory Droid, and Pi. 44 diagram types. Self-contained HTML + SVG. No shadows. No Mermaid slop.