Skip to main content
ClaudeWave

PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.

SubagentsRegistry oficial2 estrellas1 forks● PythonMITActualizado today
ClaudeWave Trust Score
95/100
✓ Verified
Passed
  • ✓Open-source license (MIT)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/9/2026
Install as a Claude Code subagent
Method: Clone
Terminal
git clone https://github.com/Sahil170595/Chimeraforge && cp Chimeraforge/*.md ~/.claude/agents/
1. Clone the repository and copy the agent .md definitions into ~/.claude/agents (or .claude/agents inside a project).
2. Start a new Claude Code session to load the agents.
3. Delegate work to them with the Task/Agent tool or by name.
Casos de uso

Resumen de Subagents

# Chimeraforge

[![PyPI version](https://img.shields.io/pypi/v/chimeraforge.svg)](https://pypi.org/project/chimeraforge/)
[![Python](https://img.shields.io/pypi/pyversions/chimeraforge.svg)](https://pypi.org/project/chimeraforge/)
[![CI](https://github.com/Sahil170595/Chimeraforge/actions/workflows/ci.yml/badge.svg)](https://github.com/Sahil170595/Chimeraforge/actions)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)

<!-- mcp-name: io.github.Sahil170595/chimeraforge -->

**A local-first, model-agnostic LLM deployment planner.** It turns "which model, quantization, GPU, and backend -- how many, will it fit, will it hit my SLO, what will it cost" into a fast, honest, measured answer, from your shell, your Python, or your AI assistant.

```bash
uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
```

## The trust principle

**Every number is labeled `measured`, `extrapolated`, `derived`, `estimated`, or `unknown`, and the tool refuses to fake the ones it can't stand behind.** VRAM and KV-cache are `derived` -- exact arithmetic over the model's real architecture, not a measurement. Throughput is a measured lookup only on the rig the corpus was measured on; on any other GPU that row is scaled by memory bandwidth and reported as `extrapolated`, carrying the row it came from, the rig it was measured on and the ratio applied, because a 17.8x bandwidth extrapolation (RTX 4080 Laptop 432 GB/s -> B200 7700 GB/s) is not a measurement of your card. Failing that it is an explicit roofline `estimate` -- never presented as data it isn't. Quality below the bundled corpus reports `unknown`, not a made-up score. A 0-result plan names the exact gate that rejected every candidate instead of a generic "nothing found." No telemetry, no phone-home, works air-gapped.

Give it a model -- a size class, a Hugging Face repo, an Ollama tag, or manual overrides for an unreleased model -- and it searches the (model x quantization x backend x GPU count x tensor/pipeline parallelism) space against VRAM, quality, latency, cost, energy, and an opt-in safety gate, then hands back the cheapest config that meets your SLO.

**18 commands, one tool:** `plan` - `deploy` - `suggest` - `measure` - `workload` - `monitor` - `validate` - `doctor` - `contribute` - `catalog` - `safety` - `bench` - `eval` - `compare` - `refit` - `report` - `mcp` - `serve`.

The empirical corpus traces to Technical Reports TR108-TR137 (~204,000 real measurements on consumer GPUs). See the [CHANGELOG](CHANGELOG.md) for the full feature history.

<!-- corpus-shape:start (scripts/corpus_shape.py --write; do not edit by hand) -->
**What the planner itself reads.** The ~204,000 measurements are the research program's total across its reports. The tables `plan` looks numbers up in are far smaller:

| Table | Size | Shape |
|---|---|---|
| Decode throughput | 23 rows | FP16 only; 7 models, the largest llama3.2-3b at 3.21B; 9 rows on serving engines (ollama 3, tgi 3, vllm 3) and 14 on transformers research harnesses; every row measured on one GPU, the RTX 4080 Laptop 12GB (192-bit GDDR6, 432 GB/s) |
| Quantization speedups | 7 multipliers | applied to an FP16 row; a quantized throughput is never a measurement of that quant |
| Quality | 35 model x quant cells | n=20 items each (TR125) |
| Safety | 40 model x quant cells | refusal rate (TR134/TR142) |
| Latency service times | 9 | model x backend |
| Third-party audit | 42 scored cells on 17 GPUs | published benchmarks ([scorecard](corpora/SCORECARD.md)); 26 more published but too underspecified to score |

Anything outside those rows is `extrapolated` (scaled by memory bandwidth from the reference GPU), `derived` (exact arithmetic) or `estimated` (roofline), and every number in a plan says which. The data records its own limits:

- Every throughput row is FP16. Quantized throughput is the FP16 row times a quant multiplier, not a measurement of that quant.
- Every row was measured on one GPU. Other hardware is bandwidth-extrapolated and labelled 'extrapolated', not 'measured'.
- The largest model measured is 3.21B; predictions above that extrapolate the power law.
- No SGLang rows exist; that backend falls through to the fp16/power-law path.
<!-- corpus-shape:end -->

---

## Install

Try it with no install:

```bash
uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
pipx run chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
```

Install for real:

```bash
pip install chimeraforge            # planner + model resolution (HF/Ollama) + suggest/measure/safety/bench
pip install "chimeraforge[bench]"     # + GPU environment metadata for benchmarks (pynvml)
pip install "chimeraforge[mcp]"       # + MCP server so Claude/GPT/Cursor can call the planner
pip install "chimeraforge[eval]"      # + quality evaluation (ROUGE-L; BERTScore additionally needs `bert-score` + torch)
pip install "chimeraforge[refit]"     # + coefficient refitting (numpy, scipy)
pip install "chimeraforge[all]"       # everything
```

Python 3.10+. The core install covers the planner and network-facing commands (`httpx` is a core dep). `plan` / `suggest` / `catalog` run fully offline; `bench` / `measure` / `safety` need a running backend (Ollama, vLLM, TGI, or SGLang; `safety` supports Ollama only). Windows / macOS / Linux.

## Quickstart

```bash
# Plan a registry size class on your GPU
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2.0

# Plan ANY model -- a Hugging Face repo or an Ollama tag
chimeraforge plan --model Qwen/Qwen2.5-7B-Instruct --hardware "RTX 4090 24GB"
chimeraforge plan --model ollama:qwen3:14b --ollama-url http://localhost:11434

# Split a model too big for one GPU across several (tensor parallelism)
chimeraforge plan --model Qwen/Qwen2.5-72B-Instruct --hardware "H100 80GB" --tp 4

# Shrink the KV-cache, print the cost/latency/quality trade-off menu
chimeraforge plan --model-size 8b --hardware "RTX 4080 12GB" --kv-quant q8 --pareto

# Benchmark a live model and plan on the MEASURED numbers
chimeraforge plan --model qwen3:14b --measure

# Discover + rank what fits your GPU and budget
chimeraforge suggest --source ollama --hardware "RTX 4090 24GB" --budget 500
```

---

## Plan with your traffic, not your guesses

```bash
chimeraforge workload --from-log requests.jsonl --out workload.json
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --workload-profile workload.json
```

Derives the request rate, prompt and output lengths, traffic variance and prefix-cache hit rate from a request log or a live vLLM/SGLang `/metrics` endpoint. The variance one matters most: `plan` otherwise takes it as one of four presets, and it drives the whole queueing tail.

Metric names are per-engine and explicit -- vLLM has renamed two of these between versions, and a scraper that silently falls back to a stale name reports a fabricated measurement. An unknown engine is an error, and a field the source did not expose stays absent rather than acquiring a default.

## Decision briefs

```bash
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2 --report brief.md
```

Writes a markdown record of the decision: the recommendation, every assumption as an input rather than a finding, the alternatives table, the planner's warnings verbatim, and the exact command that regenerates it. Each number is tagged `measured` / `extrapolated` / `derived` / `estimated` / `unknown` in prose, not just with a symbol.

It refuses to render on a stale price snapshot and exits non-zero, rather than printing an old price in a nicer font -- a formatted document reads as more durable than a terminal line, and its reader will not re-derive the arithmetic.

## MCP server -- give Claude / GPT / Cursor the same numbers

GPU sizing is exactly where assistants fail: training-cutoff hardware prices and specs, plus error-prone KV-cache/batching arithmetic done from memory. `chimeraforge mcp` runs a stdio MCP server so an assistant calls the real planner against measured data instead of guessing.

```bash
pip install "chimeraforge[mcp]"
```

Claude Code:

```bash
claude mcp add --transport stdio chimeraforge -- uvx --from "chimeraforge[mcp]" chimeraforge mcp
```

Claude Desktop / Cursor (add to your MCP config file):

```json
{
  "mcpServers": {
    "chimeraforge": {
      "command": "uvx",
      "args": ["--from", "chimeraforge[mcp]", "chimeraforge", "mcp"]
    }
  }
}
```

The `--from "chimeraforge[mcp]"` pulls in the MCP SDK; `uvx` runs the server in a self-contained environment. If you have already `pip install "chimeraforge[mcp]"` into the environment your client launches, you can instead use `"command": "chimeraforge", "args": ["mcp"]`.

Exposes five tools: `chimeraforge_plan` (the full gate search), `chimeraforge_suggest` (the inverse -- rank what actually fits a given GPU), `chimeraforge_compare_api` (self-host vs hosted-API cost and the break-even volume), `chimeraforge_resolve_model` (grounds a model id in its real params/architecture), and `chimeraforge_list_hardware`. Every result carries the same `measured` / `extrapolated` / `estimated` / `unknown` provenance as the CLI, and the tool descriptions tell the model to prefer them over its own knowledge. `chimeraforge_plan` also returns a `launch` field -- the serve command for the recommended config -- so the assistant can answer "and how do I run it" without inventing flags. `chimeraforge_compare_api` prices against a *dated* snapshot and reports its age, so an assistant quotes a price with its capture date rather than presenting a stale figure as current.

---

## Commands

### `plan` -- predictive capacity planner

```bash
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2.0
chimeraforge plan --model Qwen/Qwen2.5-7B-Instruct --hardware "RTX 4090 24GB"   # any HF repo
chimeraforge plan --model ollama:qwen3:14b --ollama-url http://localhost:11434  # any Ollama tag
chimeraforge plan --model Qwen/Qwe
agentsbenchmarkingcapacity-planninggpullmllm-inferencemcpollamaperformancepythonquantizationrustvllmvram

Lo que la gente pregunta sobre Chimeraforge

¿Qué es Sahil170595/Chimeraforge?

+

Sahil170595/Chimeraforge es subagents para el ecosistema de Claude AI. PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge. Tiene 2 estrellas en GitHub y su última actualización registrada es del 2026-10-09.

¿Cómo se instala Chimeraforge?

+

Puedes instalar Chimeraforge clonando el repositorio (https://github.com/Sahil170595/Chimeraforge) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.

¿Es seguro usar Sahil170595/Chimeraforge?

+

Nuestro agente de seguridad ha analizado Sahil170595/Chimeraforge y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.

¿Quién mantiene Sahil170595/Chimeraforge?

+

Sahil170595/Chimeraforge es mantenido por Sahil170595. La última actividad registrada en GitHub es del 2026-10-09, con 5 issues abiertos.

¿Hay alternativas a Chimeraforge?

+

Sí. En ClaudeWave puedes explorar subagents similares en /categories/agents, ordenados por popularidad o actividad reciente.

Despliega Chimeraforge en tu cloud

Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.

¿Mantienes este repo? Añade un badge a tu README

Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.

Featured on ClaudeWave: Sahil170595/Chimeraforge
[![Featured on ClaudeWave](https://claudewave.com/api/badge/sahil170595-chimeraforge)](https://claudewave.com/repo/sahil170595-chimeraforge)
<a href="https://claudewave.com/repo/sahil170595-chimeraforge"><img src="https://claudewave.com/api/badge/sahil170595-chimeraforge" alt="Featured on ClaudeWave: Sahil170595/Chimeraforge" width="320" height="64" /></a>

Más Subagents

Alternativas a Chimeraforge