Skip to main content
ClaudeWave

PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.

SubagentsOfficial Registry2 stars1 forks● PythonMITUpdated today
ClaudeWave Trust Score
95/100
✓ Verified
Passed
  • ✓Open-source license (MIT)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/9/2026
Install as a Claude Code subagent
Method: Clone
Terminal
git clone https://github.com/Sahil170595/Chimeraforge && cp Chimeraforge/*.md ~/.claude/agents/
1. Clone the repository and copy the agent .md definitions into ~/.claude/agents (or .claude/agents inside a project).
2. Start a new Claude Code session to load the agents.
3. Delegate work to them with the Task/Agent tool or by name.
Use cases

Subagents overview

# Chimeraforge

[![PyPI version](https://img.shields.io/pypi/v/chimeraforge.svg)](https://pypi.org/project/chimeraforge/)
[![Python](https://img.shields.io/pypi/pyversions/chimeraforge.svg)](https://pypi.org/project/chimeraforge/)
[![CI](https://github.com/Sahil170595/Chimeraforge/actions/workflows/ci.yml/badge.svg)](https://github.com/Sahil170595/Chimeraforge/actions)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)

<!-- mcp-name: io.github.Sahil170595/chimeraforge -->

**A local-first, model-agnostic LLM deployment planner.** It turns "which model, quantization, GPU, and backend -- how many, will it fit, will it hit my SLO, what will it cost" into a fast, honest, measured answer, from your shell, your Python, or your AI assistant.

```bash
uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
```

## The trust principle

**Every number is labeled `measured`, `extrapolated`, `derived`, `estimated`, or `unknown`, and the tool refuses to fake the ones it can't stand behind.** VRAM and KV-cache are `derived` -- exact arithmetic over the model's real architecture, not a measurement. Throughput is a measured lookup only on the rig the corpus was measured on; on any other GPU that row is scaled by memory bandwidth and reported as `extrapolated`, carrying the row it came from, the rig it was measured on and the ratio applied, because a 17.8x bandwidth extrapolation (RTX 4080 Laptop 432 GB/s -> B200 7700 GB/s) is not a measurement of your card. Failing that it is an explicit roofline `estimate` -- never presented as data it isn't. Quality below the bundled corpus reports `unknown`, not a made-up score. A 0-result plan names the exact gate that rejected every candidate instead of a generic "nothing found." No telemetry, no phone-home, works air-gapped.

Give it a model -- a size class, a Hugging Face repo, an Ollama tag, or manual overrides for an unreleased model -- and it searches the (model x quantization x backend x GPU count x tensor/pipeline parallelism) space against VRAM, quality, latency, cost, energy, and an opt-in safety gate, then hands back the cheapest config that meets your SLO.

**18 commands, one tool:** `plan` - `deploy` - `suggest` - `measure` - `workload` - `monitor` - `validate` - `doctor` - `contribute` - `catalog` - `safety` - `bench` - `eval` - `compare` - `refit` - `report` - `mcp` - `serve`.

The empirical corpus traces to Technical Reports TR108-TR137 (~204,000 real measurements on consumer GPUs). See the [CHANGELOG](CHANGELOG.md) for the full feature history.

<!-- corpus-shape:start (scripts/corpus_shape.py --write; do not edit by hand) -->
**What the planner itself reads.** The ~204,000 measurements are the research program's total across its reports. The tables `plan` looks numbers up in are far smaller:

| Table | Size | Shape |
|---|---|---|
| Decode throughput | 23 rows | FP16 only; 7 models, the largest llama3.2-3b at 3.21B; 9 rows on serving engines (ollama 3, tgi 3, vllm 3) and 14 on transformers research harnesses; every row measured on one GPU, the RTX 4080 Laptop 12GB (192-bit GDDR6, 432 GB/s) |
| Quantization speedups | 7 multipliers | applied to an FP16 row; a quantized throughput is never a measurement of that quant |
| Quality | 35 model x quant cells | n=20 items each (TR125) |
| Safety | 40 model x quant cells | refusal rate (TR134/TR142) |
| Latency service times | 9 | model x backend |
| Third-party audit | 42 scored cells on 17 GPUs | published benchmarks ([scorecard](corpora/SCORECARD.md)); 26 more published but too underspecified to score |

Anything outside those rows is `extrapolated` (scaled by memory bandwidth from the reference GPU), `derived` (exact arithmetic) or `estimated` (roofline), and every number in a plan says which. The data records its own limits:

- Every throughput row is FP16. Quantized throughput is the FP16 row times a quant multiplier, not a measurement of that quant.
- Every row was measured on one GPU. Other hardware is bandwidth-extrapolated and labelled 'extrapolated', not 'measured'.
- The largest model measured is 3.21B; predictions above that extrapolate the power law.
- No SGLang rows exist; that backend falls through to the fp16/power-law path.
<!-- corpus-shape:end -->

---

## Install

Try it with no install:

```bash
uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
pipx run chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
```

Install for real:

```bash
pip install chimeraforge            # planner + model resolution (HF/Ollama) + suggest/measure/safety/bench
pip install "chimeraforge[bench]"     # + GPU environment metadata for benchmarks (pynvml)
pip install "chimeraforge[mcp]"       # + MCP server so Claude/GPT/Cursor can call the planner
pip install "chimeraforge[eval]"      # + quality evaluation (ROUGE-L; BERTScore additionally needs `bert-score` + torch)
pip install "chimeraforge[refit]"     # + coefficient refitting (numpy, scipy)
pip install "chimeraforge[all]"       # everything
```

Python 3.10+. The core install covers the planner and network-facing commands (`httpx` is a core dep). `plan` / `suggest` / `catalog` run fully offline; `bench` / `measure` / `safety` need a running backend (Ollama, vLLM, TGI, or SGLang; `safety` supports Ollama only). Windows / macOS / Linux.

## Quickstart

```bash
# Plan a registry size class on your GPU
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2.0

# Plan ANY model -- a Hugging Face repo or an Ollama tag
chimeraforge plan --model Qwen/Qwen2.5-7B-Instruct --hardware "RTX 4090 24GB"
chimeraforge plan --model ollama:qwen3:14b --ollama-url http://localhost:11434

# Split a model too big for one GPU across several (tensor parallelism)
chimeraforge plan --model Qwen/Qwen2.5-72B-Instruct --hardware "H100 80GB" --tp 4

# Shrink the KV-cache, print the cost/latency/quality trade-off menu
chimeraforge plan --model-size 8b --hardware "RTX 4080 12GB" --kv-quant q8 --pareto

# Benchmark a live model and plan on the MEASURED numbers
chimeraforge plan --model qwen3:14b --measure

# Discover + rank what fits your GPU and budget
chimeraforge suggest --source ollama --hardware "RTX 4090 24GB" --budget 500
```

---

## Plan with your traffic, not your guesses

```bash
chimeraforge workload --from-log requests.jsonl --out workload.json
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --workload-profile workload.json
```

Derives the request rate, prompt and output lengths, traffic variance and prefix-cache hit rate from a request log or a live vLLM/SGLang `/metrics` endpoint. The variance one matters most: `plan` otherwise takes it as one of four presets, and it drives the whole queueing tail.

Metric names are per-engine and explicit -- vLLM has renamed two of these between versions, and a scraper that silently falls back to a stale name reports a fabricated measurement. An unknown engine is an error, and a field the source did not expose stays absent rather than acquiring a default.

## Decision briefs

```bash
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2 --report brief.md
```

Writes a markdown record of the decision: the recommendation, every assumption as an input rather than a finding, the alternatives table, the planner's warnings verbatim, and the exact command that regenerates it. Each number is tagged `measured` / `extrapolated` / `derived` / `estimated` / `unknown` in prose, not just with a symbol.

It refuses to render on a stale price snapshot and exits non-zero, rather than printing an old price in a nicer font -- a formatted document reads as more durable than a terminal line, and its reader will not re-derive the arithmetic.

## MCP server -- give Claude / GPT / Cursor the same numbers

GPU sizing is exactly where assistants fail: training-cutoff hardware prices and specs, plus error-prone KV-cache/batching arithmetic done from memory. `chimeraforge mcp` runs a stdio MCP server so an assistant calls the real planner against measured data instead of guessing.

```bash
pip install "chimeraforge[mcp]"
```

Claude Code:

```bash
claude mcp add --transport stdio chimeraforge -- uvx --from "chimeraforge[mcp]" chimeraforge mcp
```

Claude Desktop / Cursor (add to your MCP config file):

```json
{
  "mcpServers": {
    "chimeraforge": {
      "command": "uvx",
      "args": ["--from", "chimeraforge[mcp]", "chimeraforge", "mcp"]
    }
  }
}
```

The `--from "chimeraforge[mcp]"` pulls in the MCP SDK; `uvx` runs the server in a self-contained environment. If you have already `pip install "chimeraforge[mcp]"` into the environment your client launches, you can instead use `"command": "chimeraforge", "args": ["mcp"]`.

Exposes five tools: `chimeraforge_plan` (the full gate search), `chimeraforge_suggest` (the inverse -- rank what actually fits a given GPU), `chimeraforge_compare_api` (self-host vs hosted-API cost and the break-even volume), `chimeraforge_resolve_model` (grounds a model id in its real params/architecture), and `chimeraforge_list_hardware`. Every result carries the same `measured` / `extrapolated` / `estimated` / `unknown` provenance as the CLI, and the tool descriptions tell the model to prefer them over its own knowledge. `chimeraforge_plan` also returns a `launch` field -- the serve command for the recommended config -- so the assistant can answer "and how do I run it" without inventing flags. `chimeraforge_compare_api` prices against a *dated* snapshot and reports its age, so an assistant quotes a price with its capture date rather than presenting a stale figure as current.

---

## Commands

### `plan` -- predictive capacity planner

```bash
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2.0
chimeraforge plan --model Qwen/Qwen2.5-7B-Instruct --hardware "RTX 4090 24GB"   # any HF repo
chimeraforge plan --model ollama:qwen3:14b --ollama-url http://localhost:11434  # any Ollama tag
chimeraforge plan --model Qwen/Qwe
agentsbenchmarkingcapacity-planninggpullmllm-inferencemcpollamaperformancepythonquantizationrustvllmvram

What people ask about Chimeraforge

What is Sahil170595/Chimeraforge?

+

Sahil170595/Chimeraforge is subagents for the Claude AI ecosystem. PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge. It has 2 GitHub stars and its last recorded update is dated 2026-10-09.

How do I install Chimeraforge?

+

You can install Chimeraforge by cloning the repository (https://github.com/Sahil170595/Chimeraforge) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is Sahil170595/Chimeraforge safe to use?

+

Our security agent has analyzed Sahil170595/Chimeraforge and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.

Who maintains Sahil170595/Chimeraforge?

+

Sahil170595/Chimeraforge is maintained by Sahil170595. The last recorded GitHub activity is dated 2026-10-09, with 5 open issues.

Are there alternatives to Chimeraforge?

+

Yes. On ClaudeWave you can browse similar subagents at /categories/agents, sorted by popularity or recent activity.

Deploy Chimeraforge to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: Sahil170595/Chimeraforge
[![Featured on ClaudeWave](https://claudewave.com/api/badge/sahil170595-chimeraforge)](https://claudewave.com/repo/sahil170595-chimeraforge)
<a href="https://claudewave.com/repo/sahil170595-chimeraforge"><img src="https://claudewave.com/api/badge/sahil170595-chimeraforge" alt="Featured on ClaudeWave: Sahil170595/Chimeraforge" width="320" height="64" /></a>

More Subagents

Chimeraforge alternatives