Self-hosted, OpenAI-compatible LLM gateway that measures the cost of every request and attributes it to the customer, workflow, or feature that drove it — without ever storing the prompt or response. Change one base_url; every call produces a payload-free economic receipt. Apache-2.0, zero dependency on any Inferrail-operated service.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add inferrail -- python -m inferrail{
"mcpServers": {
"inferrail": {
"command": "python",
"args": ["-m", "inferrail"],
"env": {
"OPENAI_API_KEY": "<openai_api_key>"
}
}
}
}OPENAI_API_KEYResumen de MCP Servers
# Inferrail
<!-- mcp-name: io.github.domondi1/inferrail -->
A self-hosted, OpenAI-compatible gateway that sits between your app and
your LLM provider, so every request produces a payload-free economic
receipt — measured cost and business attribution, with no prompt or
response ever stored.
[](https://github.com/domondi1/inferrail/actions/workflows/ci.yml)
[](LICENSE)
Point your existing OpenAI client at Inferrail instead of directly at
your provider. Nothing else changes — same request/response shape,
same streaming, same tool calls — except now every call gets logged
locally with its real token usage, its verified cost, and whatever
customer/workflow/feature you attribute it to.
**Developer preview (v0.1.1).** Works today; CLI flags, config shape, and
receipt fields may still change before 1.0 — see
[docs/PRODUCT.md](docs/PRODUCT.md).
## Why Inferrail
- **Know what each request cost, and who it was for** — without adding a
hosted observability vendor or logging prompts yourself.
- **One place to point an OpenAI-compatible client** instead of
provider-specific SDK code sprinkled through your app.
- **Runs entirely on your own machine.** No Inferrail-operated service
exists yet, and none of this depends on one.
- **Structurally can't store your prompts.** The receipt and telemetry
schemas have no field capable of holding message content — not a
setting, a guarantee.
## Install
```bash
pip install inferrail
```
Requires Python 3.11+.
## Quickstart
See the whole pipeline — request → receipt → cost report — with no API
key and no network call:
```bash
inferrail demo
```
Runs canned requests through Inferrail's real engine using a fake
in-memory provider. Every price and response is clearly labeled `DEMO`.
To try it for real, with your own key:
```bash
export OPENAI_API_KEY=<your-openai-api-key>
inferrail try "Reply with one word: ready" --customer acme
```
```
ready
Receipt ir_670b20135cfe4bcb8f6f
Provider openai
Model gpt-4o-mini
Input tokens 12
Output tokens 1
Cost $0.000002
Customer acme
Prompt stored no
Response stored no
Saved to ./inferrail-receipts.jsonl
Next:
inferrail report --by customer
```
`inferrail report --by customer` (or `provider`, `model`, `route`, or
any attribute you've attached) aggregates every receipt written so far
into a table of requests, tokens, and cost.
## Send a request
Run the gateway itself with the same zero-config defaults:
```bash
inferrail serve --quickstart
```
```bash
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-Inferrail-Attribute-Customer: acme" \
-d '{
"model": "default",
"messages": [{"role": "user", "content": "Say hello in five words."}]
}'
```
The response is standard OpenAI `choices`/`usage` plus a non-standard
`inferrail` block (route, provider, latency, retries) any OpenAI client
already ignores. `X-Inferrail-Attribute-*` headers are optional
attribution — never forwarded upstream. See
[examples/basic_chat_request.py](examples/basic_chat_request.py) for a
minimal Python client, or point any existing OpenAI-compatible SDK
(LangChain, LlamaIndex, CrewAI, ...) at `http://127.0.0.1:8000/v1`.
<details>
<summary>Framework examples (LangChain, LlamaIndex, CrewAI)</summary>
```python
# LangChain
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
base_url="http://127.0.0.1:8000/v1",
api_key="not-needed", # or your INFERRAIL_GATEWAY_TOKEN if auth is enabled
model="default",
)
```
```python
# LlamaIndex
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="default",
api_base="http://127.0.0.1:8000/v1",
api_key="not-needed",
is_chat_model=True,
context_window=8192,
)
```
```python
# CrewAI
from crewai import LLM
llm = LLM(
model="openai/default", # "openai/" prefix required by CrewAI
base_url="http://127.0.0.1:8000/v1",
api_key="not-needed",
)
```
</details>
## What Inferrail records
Every request writes one payload-free JSON receipt — no field on it can
hold a prompt or response:
```json
{
"receipt_id": "ir_1e6c916bac8940ca8a85",
"route": "default",
"provider": "openai",
"model": "gpt-4o-mini",
"status": "success",
"prompt_tokens": 842,
"completion_tokens": 191,
"pricing": {
"input_usd_per_million": "0.15",
"output_usd_per_million": "0.60",
"source": "https://developers.openai.com/api/docs/pricing",
"verified_date": "2026-08-16"
},
"estimated_cost_usd": "0.000241",
"attributes": { "customer": "acme", "workflow": "contract-review" },
"total_latency_ms": 15.96,
"retry_count": 0
}
```
If Inferrail can't verify a price for the (provider, model) pair, `pricing`
and `estimated_cost_usd` are `null` — never a guessed or fabricated cost.
All money math uses `Decimal`, never `float`. Details:
[docs/adr/0005](docs/adr/0005-privacy-preserving-economic-receipts.md).
## Model routing
`"model"` normally selects a named route from `inferrail.yaml` (e.g.
`"default"`), which maps to a provider + underlying model. If
`default_provider` is set in your config, a `model` that matches no route
is instead forwarded to that provider unchanged — so `"model":
"gpt-5.6-sol"` works with no route pre-registered for it. Named routes
always take priority. This passthrough is on by default for the
zero-config quickstart path, off by default otherwise. Full design:
[docs/adr/0007](docs/adr/0007-model-passthrough-routing.md).
## MCP
```bash
pip install "inferrail[mcp]"
```
An MCP server (`inferrail-mcp`), published on the MCP registry as
[`io.github.domondi1/inferrail`](https://registry.modelcontextprotocol.io),
exposes Inferrail's local receipt ledger to any MCP-aware agent (Claude
Code, Claude Desktop, Cursor, ...) as two **read-only** tools — neither
executes inference or spends provider budget:
| Tool | What it does |
|---|---|
| `get_spend` | Aggregates local receipts by provider/model/route/attribute, optional time window |
| `get_health` | Checks gateway reachability + most recent local receipt |
```json
{
"mcpServers": {
"inferrail": { "command": "inferrail-mcp" }
}
}
```
Claude Code: `claude mcp add inferrail -- inferrail-mcp`. Full contract:
[inferrail-mcp/README.md](inferrail-mcp/README.md).
## Configuration
For a real deployment instead of quickstart defaults:
```bash
cp inferrail.example.yaml inferrail.yaml
cp .env.example .env # then add a real OPENAI_API_KEY
inferrail config check # validate without starting a server
inferrail serve
```
`inferrail.yaml` only ever holds the *name* of an environment variable
for a secret, never the secret itself. Full shape (providers, routes,
telemetry, receipts, pricing overrides):
[inferrail.example.yaml](inferrail.example.yaml).
By default the gateway binds to `127.0.0.1:8000` with no auth. Set
`INFERRAIL_GATEWAY_TOKEN` to require callers to send `Authorization:
Bearer <token>` — see [SECURITY.md](SECURITY.md).
## What works today
- `POST /v1/chat/completions`: streaming (`stream: true`, real SSE
passthrough) and tool/function calling, single string message content,
no `n != 1`
- `GET /health`
- One provider adapter, generic over any OpenAI-compatible HTTP endpoint
- Named-route + optional passthrough model routing (above)
- Per-route retry with backoff on transient provider errors
- Local structured telemetry and payload-free cost receipts for every
request, plus `inferrail report --by <dimension>`
- CLI: `inferrail demo`, `try`, `serve` (`--quickstart`), `config check`,
`report`
**Not yet:** multi-provider intelligent routing, cost estimates for
models outside the built-in catalog or an explicit `pricing:` override,
budgets/spend limits, any non-OpenAI-compatible provider. Full scope:
[docs/PRODUCT.md](docs/PRODUCT.md).
## Documentation
- [docs/PRODUCT.md](docs/PRODUCT.md) — exact current scope
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — package layout, request
lifecycle
- [docs/adr/](docs/adr/) — why specific structural decisions were made
- [openapi.json](openapi.json) / [config.schema.json](config.schema.json)
/ [llms.txt](llms.txt) — machine-readable references for tooling and
agents
- [SECURITY.md](SECURITY.md)
## Development
```bash
git clone https://github.com/domondi1/inferrail.git && cd inferrail
pip install -e ".[dev,mcp]"
ruff check . && mypy && pytest
```
`pytest` needs no API key or network access — see
[CONTRIBUTING.md](CONTRIBUTING.md).
## License
Apache License 2.0 — see [LICENSE](LICENSE).
Lo que la gente pregunta sobre inferrail
¿Qué es domondi1/inferrail?
+
domondi1/inferrail es mcp servers para el ecosistema de Claude AI. Self-hosted, OpenAI-compatible LLM gateway that measures the cost of every request and attributes it to the customer, workflow, or feature that drove it — without ever storing the prompt or response. Change one base_url; every call produces a payload-free economic receipt. Apache-2.0, zero dependency on any Inferrail-operated service. Tiene 0 estrellas en GitHub y su última actualización registrada es del 2026-08-21.
¿Cómo se instala inferrail?
+
Puedes instalar inferrail clonando el repositorio (https://github.com/domondi1/inferrail) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar domondi1/inferrail?
+
Nuestro agente de seguridad ha analizado domondi1/inferrail y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene domondi1/inferrail?
+
domondi1/inferrail es mantenido por domondi1. La última actividad registrada en GitHub es del 2026-08-21, con 0 issues abiertos.
¿Hay alternativas a inferrail?
+
Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.
Despliega inferrail en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/domondi1-inferrail)<a href="https://claudewave.com/repo/domondi1-inferrail"><img src="https://claudewave.com/api/badge/domondi1-inferrail" alt="Featured on ClaudeWave: domondi1/inferrail" width="320" height="64" /></a>Más MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
The fastest path to AI-powered full stack observability, even for lean teams.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!