Compression with a quality contract — cache-aware, causally-pruned LLM context compression for agentic runtimes, certified non-inferior across 7 domains. Works with any SDK.
git clone https://github.com/dshakes/distilResumen de Tools
<!-- mcp-name: io.github.dshakes/distil --> <p align="center"> <img src="docs/assets/banner.svg" alt="Distil — compression with a quality contract" width="100%"/> </p> <p align="center"> <a href="https://github.com/dshakes/distil/actions/workflows/ci.yml"><img src="https://github.com/dshakes/distil/actions/workflows/ci.yml/badge.svg" alt="CI"/></a> <a href="https://pypi.org/project/distil-llm/"><img src="https://img.shields.io/pypi/v/distil-llm?color=5ad1c9&label=pypi" alt="PyPI version"/></a> <a href="https://www.npmjs.com/package/distil-llm"><img src="https://img.shields.io/npm/v/distil-llm?color=5ad1c9&label=npm" alt="npm version"/></a> <a href="https://pypi.org/project/distil-llm/"><img src="https://img.shields.io/pypi/pyversions/distil-llm?color=5ad1c9" alt="Python versions"/></a> <a href="LICENSE"><img src="https://img.shields.io/pypi/l/distil-llm?color=8b7bff" alt="license"/></a> <a href="#-what-we-wont-pretend"><img src="https://img.shields.io/badge/runtime%20deps-0-5ad19a" alt="zero runtime deps"/></a> <a href="https://dshakes.github.io/distil/architecture.html"><img src="https://img.shields.io/badge/typed-py.typed%20%C2%B7%20mypy%20clean-8b7bff" alt="typed"/></a> <a href="https://dshakes.github.io/distil/adoption.html"><img src="https://img.shields.io/endpoint?url=https%3A%2F%2Fraw.githubusercontent.com%2Fdshakes%2Fdistil%2Fmetrics%2Fdata%2Fbadges%2Fdownloads-real.json" alt="PyPI installs/month, bot-filtered"/></a> </p> <h2 align="center">Compress your agent's context.<br/>Prove its decisions don't change.</h2> <p align="center"><b>Every other compressor asks you to <i>trust</i> it won't break your agent. Distil is the only one that proves it won't.</b><br/>On <b>500 real coding tasks</b>, compressed context <b>matched full context within statistical noise</b>: <b>42.0% vs 39.2%</b>. <sub>(SWE-bench Verified)</sub></p> <p align="center"> <img src="docs/assets/hero-terminal.svg" alt="Animated distil proof session: distil bench prints GATE: PASS (every trajectory certified non-inferior); distil wrap -- claude routes with zero config; a live line shows 53% smaller, equivalence 100%; then the proof ledger closes with 1,284,551 → 601,204 tokens (53.2% smaller), cost $18.41 → $8.72 calibrated to billed usage, 0 shadow decision changes across 63 A/B samples, 100% recoverable restore" width="84%"/> </p> <p align="center"><sub><code>uvx --from distil-llm distil bench</code> — runs the certificate gate in ~10s, no API key · <code>distil wrap -- claude</code> routes your agent, zero config.</sub></p> <!-- ═══ LIVE community counter — fed by the opt-in census, re-polls every 5 min ═══ --> <p align="center"><sub>◉ <b>LIVE</b> · measured from the opt-in census on a <a href="https://github.com/dshakes/distil/tree/metrics">public git branch</a>, never estimated</sub></p> <p align="center"> <a href="https://dshakes.github.io/distil/adoption.html"><img src="https://img.shields.io/endpoint?style=for-the-badge&url=https%3A%2F%2Fraw.githubusercontent.com%2Fdshakes%2Fdistil%2Fmetrics%2Fdata%2Fbadges%2Fsavings-tokens.json" alt="community tokens saved"/></a> <a href="https://dshakes.github.io/distil/adoption.html"><img src="https://img.shields.io/endpoint?style=for-the-badge&url=https%3A%2F%2Fraw.githubusercontent.com%2Fdshakes%2Fdistil%2Fmetrics%2Fdata%2Fbadges%2Fequivalence.json" alt="decision-equivalence"/></a> <a href="https://dshakes.github.io/distil/adoption.html"><img src="https://img.shields.io/endpoint?style=for-the-badge&url=https%3A%2F%2Fraw.githubusercontent.com%2Fdshakes%2Fdistil%2Fmetrics%2Fdata%2Fbadges%2Factive-installs.json" alt="active installs, 30d"/></a> </p> <p align="center"><b><a href="https://dshakes.github.io/distil/adoption.html">▶ Watch the counter tick live & audit every number →</a></b></p> <table align="center"><tr> <td align="center"><b>⚡ Get the savings</b><br/><sub>2 min, no config</sub><br/><br/><code>pipx install distil-llm</code><br/><code>distil onboard</code></td> <td align="center"><b>🔬 See the proof</b><br/><sub>real harness</sub><br/><br/><a href="#-the-proof"><b>benchmark ↓</b></a> · <a href="docs/PAPER.md">paper</a><br/><a href="https://dshakes.github.io/distil/compare.html">vs the others</a></td> </tr></table> <p align="center"><sub>Honest scope: +2.8pp is a point estimate (CI −0.6..+6.2pp — <b>non-inferiority certified, superiority not yet</b>). <a href="#-the-proof">Details, incl. what doesn't transfer →</a></sub></p> <p align="center"> <a href="#-use-it-now">Use it</a> · <a href="#-works-with-every-sdk">Integrations</a> · <a href="#-install-your-way">Install</a> · <a href="https://dshakes.github.io/distil/compare.html">vs the others</a> · <a href="https://dshakes.github.io/distil/getting-started.html"><b>Full Docs →</b></a> </p> --- <h3 align="center">Proof first — not a pitch 📊</h3> <p align="center"><img src="docs/assets/head-to-head.svg" alt="Distil vs LLMLingua-2 vs Headroom — token savings, decision-change rate, latency" width="100%"/></p> <table align="center"> <tr><th>On a real 500-instance long-horizon agent<br/><sub>(SWE-bench Verified, official harness)</sub></th><th>task success</th><th>tied with full context?</th><th>reversible + certified?</th></tr> <tr><td><b>Distil</b> (gated + surprise digest, v1.7)</td><td align="center"><b>42.0%</b></td><td align="center">✅ <b>tied</b> <sub>(+2.8pp point est., CI −0.6..+6.2 — n.s.)</sub></td><td align="center">✅</td></tr> <tr><td><b>Distil</b> (relevance-gated, E8)</td><td align="center"><b>36.8%</b></td><td align="center">✅</td><td align="center">✅</td></tr> <tr><td>Headroom <sub>(lossy)</sub></td><td align="center">32.6%</td><td align="center">❌ −6.6pp</td><td align="center">❌</td></tr> <tr><td>LLMLingua-2 <sub>(lossy — only 16/500 runs completed)</sub></td><td align="center">2.4%</td><td align="center">❌ −36.8pp</td><td align="center">❌</td></tr> <tr><td>no compression <sub>(full)</sub></td><td align="center">39.2%</td><td align="center">—</td><td align="center">—</td></tr> </table> <p align="center"><b>Distil is the only compressor statistically tied with full context — its v1.7 surprise-preserving digest reaches 42.0% vs 39.2% (paired non-inferiority certified; superiority not significant)</b> while every lossy tool craters. And on the live head-to-head above (graded by <code>claude-opus-4-8</code>), it certifies <b>83.2% savings at a 0% decision-change rate</b>, ~1,000× faster than the nearest tool <sub>(distil is pure-Python heuristics — no local ML model; competitors run transformer inference)</sub>. <a href="#-the-proof">Full breakdown ↓</a></p> --- ## 🚀 Use it now **One command sets you up and tells you what to do next:** ```bash pipx install distil-llm distil onboard # detects your agent + billing, wires the status line, prints a guided tour ``` It detects your environment (Claude Code · Codex · Gemini CLI; metered vs subscription) and hands you the exact commands. Or wrap your agent directly — **no config, no code change:** ```bash # Claude Code on a metered API key — saves real $$: distil wrap --expand -- claude # Claude Code on a Pro/Max subscription — flat-rate, ToS-safe (trims context, not $): distil wrap --lossless-only -- claude # Codex, Gemini CLI, aider — same pattern; env var auto-selected per agent: distil wrap --expand -- codex # → OPENAI_BASE_URL distil wrap --expand -- gemini # → GOOGLE_GEMINI_BASE_URL distil wrap --expand -- aider # → OPENAI_BASE_URL # Headless too — print mode, CI, and Agent SDK scripts route the same way: distil wrap -- claude -p "summarise this diff" distil wrap -- python my_agent_sdk_script.py ``` Each recognized agent (`claude` / `codex` / `gemini` / `aider`) auto-selects the right env var and upstream — no `--env-var` or `--upstream` flag needed. Prints `preset: <agent> detected → <VAR>` on start. Explicit flags always win. <details> <summary><b>Make it the default</b> — never type <code>distil wrap</code> again</summary> **Tired of typing `distil wrap` every time?** Make it the default — once: ```bash distil default # adds a managed shell alias so `claude` always routes through distil distil default --undo # remove it anytime (backed up before any change) ``` It detects your shell (zsh / bash / fish / PowerShell) and billing mode, writes the right line to the rc file your shell actually reads, and **tells you what it detected**. Want every SDK covered (not just the agent you type)? `distil default --always-on` runs a persistent proxy service — powerful, but it's a daemon you keep alive. </details> Then watch genuine savings from **your** traffic — measured, not estimated: ```bash distil leaderboard # cumulative tokens + $ saved, from the local ledger distil dashboard # live terminal TUI — token-trim + decision-equiv bars, Ctrl-C to exit distil dissect # per-session deep-dive: savings, digest inventory, anomalies (--html/--serve) ``` **Validate it on your traffic.** `--shadow` runs a fraction of requests twice (compressed **and** full) and compares the agent's chosen next action: ```bash distil wrap --shadow 0.1 -- claude # wrap + shadow 10% of requests distil shadow-stats # live decision-equivalence rate ``` Honest scope: that's next-action equivalence — a **proxy**, not task success ([E7](#-the-proof) shows it doesn't fully transfer under aggressive *lossy* compression). Distil fails safe to full context. > **Will it save money?** Only on **metered** billing (API key) — fewer tokens, fewer dollars. On a flat-rate **subscription** it trims context + latency, not the bill. Coding agents: short sessions ~7%, big wins on **long, many-turn** sessions the model never re-reads. --- ## 💡 Why Distil is different You don't need byte-equivalence — you need **decision-equivalence**: your agent taking the *same actions* with compressed context. That's measurable and certifiable. - **Certified, not estimat
Lo que la gente pregunta sobre distil
¿Qué es dshakes/distil?
+
dshakes/distil es tools para el ecosistema de Claude AI. Compression with a quality contract — cache-aware, causally-pruned LLM context compression for agentic runtimes, certified non-inferior across 7 domains. Works with any SDK. Tiene 7 estrellas en GitHub y se actualizó por última vez today.
¿Cómo se instala distil?
+
Puedes instalar distil clonando el repositorio (https://github.com/dshakes/distil) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar dshakes/distil?
+
dshakes/distil aún no ha sido auditado por nuestro agente de seguridad. Revisa el repositorio original en GitHub antes de usarlo en producción.
¿Quién mantiene dshakes/distil?
+
dshakes/distil es mantenido por dshakes. La última actividad registrada en GitHub es de today, con 0 issues abiertos.
¿Hay alternativas a distil?
+
Sí. En ClaudeWave puedes explorar tools similares en /categories/tools, ordenados por popularidad o actividad reciente.
Despliega distil en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/dshakes-distil)<a href="https://claudewave.com/repo/dshakes-distil"><img src="https://claudewave.com/api/badge/dshakes-distil" alt="Featured on ClaudeWave: dshakes/distil" width="320" height="64" /></a>Más Tools
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary