Skip to main content
ClaudeWave

Compare text LLMs across providers via OpenRouter — bring your own API key.

PluginsOfficial Registry0 stars0 forks● PythonMITUpdated today
ClaudeWave Trust Score
87/100
✓ Trusted
Passed
  • ✓Open-source license (MIT)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Documented (README)
Last scanned: 10/2/2026
Install as a Claude Code plugin
Method: Clone
Claude Code
/plugin marketplace add thejaredchapman/evalforge-lite
/plugin install evalforge-lite
1. Inside Claude Code, add the marketplace and install the plugin with the commands above.
2. Follow any post-install configuration from the README.
3. Restart the session if commands or hooks do not show up immediately.
Use cases

Plugins overview

# EvalForge Lite

<!-- mcp-name: io.github.thejaredchapman/evalforge-lite -->

Compare text LLMs side by side. Write a few test prompts, pick up to four models
across OpenRouter, Amazon Bedrock, Google Vertex AI and Microsoft Foundry, and
EvalForge Lite sends the same prompts to all of them, scores the answers
automatically, and shows a leaderboard with letter grades, response time, speed
and estimated cost. Use it as a web app or as an MCP server for Claude and other
assistants. You bring your own credentials, or the person hosting it keeps them
on the server for you.

**Documentation site:** https://thejaredchapman.github.io/evalforge-lite/

## Why use it

- **Test on your own prompts.** Choose models from evidence about your tasks, not a generic benchmark.
- **Up to 4 models per run, across 4 backends.** `X` (OpenRouter) and `X@bedrock` are separate targets, so you can check one model on two platforms in a single run.
- **Automatic grading.** A judge model scores each answer against your rubric, and rule checks (`contains`, `regex`, `json_valid`, `max_length`, available through the API and MCP) add a pass or fail. You get a 0-100 score and a letter grade.
- **A second opinion on every response.** Each answer is also evaluated on six criteria (answered, quality, instruction following, completeness, helpfulness, safety) with strengths and weaknesses written out.
- **Speed and cost beside quality.** Latency, tokens per second and estimated cost for every model, and a "What matters most?" selector that moves the "Best for ..." badge without a new run.
- **Policy gate.** Upload a company policy and prompts that violate it are blocked before any model is called. If the check itself fails, the prompt is blocked.
- **Reports.** Download a PDF or a CSV for any of your last five runs.
- **No accounts, no database.** Credentials are used for one request and not stored. Nothing is written to disk.

## Quick start

### 1. Run the web app on your computer

Requires Python 3.10 or newer (3.12 recommended; download from
https://www.python.org/downloads/) and git.

    git clone https://github.com/thejaredchapman/evalforge-lite.git
    cd evalforge-lite
    python3.12 -m venv venv
    source venv/bin/activate
    pip install -r requirements.txt
    python app.py

Open http://localhost:8000, paste a key for at least one backend (an OpenRouter
API key is the quickest: https://openrouter.ai/workspaces/default/keys), add a
test case, pick two to four models, and click **Run comparison**. Runs are
limited to 3 per 8 hours per browser session.
Full walkthrough: [Getting started](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/getting-started.md).

### 2. Use it from Claude (MCP server)

With [`uv`](https://docs.astral.sh/uv/) installed:

    uvx evalforge-lite

Add it to Claude Code in one line:

    claude mcp add evalforge-lite -- uvx evalforge-lite

Or install the Claude Code plugin, which bundles the same server:

    claude plugin marketplace add thejaredchapman/evalforge-lite
    claude plugin install evalforge-lite@evalforge

Then ask your assistant to compare models. It gets 9 tools: `list_models`,
`suggest_models`, `list_availability`, `set_policy`, `evaluate_prompt`,
`run_comparison`, `list_runs`, `get_report`, `get_report_csv`.
Details, Claude Desktop config and credential shapes: [MCP server](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/mcp-server.md).

### 3. Host it for other people

Deploy with the included `render.yaml` (gunicorn, one worker) or any host that
can run `gunicorn --workers 1 --threads 4 --bind 0.0.0.0:$PORT app:app`. By
default every visitor supplies their own key. Optionally keep provider keys on
the server with environment variables and a shared daily cap (50 per 24 hours by
default). Keep it at one worker: all state is in memory per process.
Full guide: [Hosting and server-side keys](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md).

## Good to know

- The app does not read `.env` by itself. To use values from it, run `set -a; source .env; set +a` before `python app.py`.
- Each run is limited to 4 models, and each browser session gets 3 runs per 8 hours.
- All state lives in memory and is cleared when the server restarts. See [Privacy and limits](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/privacy-and-limits.md).
- Upgrading from an older version and reading the CSV or API fields? See the notes in [Troubleshooting and FAQ](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/troubleshooting-and-faq.md#changes-to-exports-and-api-fields).

## Backends and credentials

| Backend | What you provide |
|---|---|
| OpenRouter | One API key |
| Amazon Bedrock | A region, plus a Bedrock API key or AWS access keys (optional session token) |
| Google Vertex AI | A project id and region, plus an access token or service-account JSON |
| Microsoft Foundry | A resource name and region, plus an API key or Entra ID access token |

A separate **judge backend** setting chooses where the judge and policy gate
run. Bedrock, Vertex and Foundry costs are estimates from catalog prices, not
your cloud bill. See [Backends and credentials](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/backends-and-credentials.md).

## Documentation

| Page | What is in it |
|---|---|
| [Overview](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/index.md) | What it is, who it is for, the three ways to use it |
| [Getting started](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/getting-started.md) | Install, run, and your first comparison |
| [Web app guide](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/web-app.md) | Every part of the screen, in order |
| [Comparing models](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/comparing-models.md) | Reading metrics, grades, evaluation, cost and their limits |
| [Backends and credentials](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/backends-and-credentials.md) | Keys, regions, `X@backend` targets |
| [MCP server](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/mcp-server.md) | Install paths, all 9 tools, example prompts |
| [Hosting and server-side keys](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md) | Deploying for others, operator-held keys, daily cap |
| [Troubleshooting and FAQ](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/troubleshooting-and-faq.md) | Common messages, fixes, and notes on CSV/API field changes |
| [Privacy and limits](https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/privacy-and-limits.md) | What data goes where, what is stored, every limit |

## Test

    pytest tests/ -v

Every model and HTTP call is mocked, so the suite needs no API key and makes no
network calls.

## Contributing

Contributions are welcome: bug reports, model-catalog updates, new checks, docs,
and new backends. See [CONTRIBUTING.md](https://github.com/thejaredchapman/evalforge-lite/blob/main/CONTRIBUTING.md) for setup, tests and the
pull request process. When the app shows an error, the popup's **Report an issue
on GitHub** button opens a pre-filled bug report.

## License

[MIT](https://github.com/thejaredchapman/evalforge-lite/blob/main/LICENSE)

What people ask about evalforge-lite

What is thejaredchapman/evalforge-lite?

+

thejaredchapman/evalforge-lite is plugins for the Claude AI ecosystem. Compare text LLMs across providers via OpenRouter — bring your own API key. It has 0 GitHub stars and its last recorded update is dated 2026-10-01.

How do I install evalforge-lite?

+

You can install evalforge-lite by cloning the repository (https://github.com/thejaredchapman/evalforge-lite) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is thejaredchapman/evalforge-lite safe to use?

+

Our security agent has analyzed thejaredchapman/evalforge-lite and assigned a Trust Score of 87/100 (tier: Trusted). See the full breakdown of passed checks and flags on this page.

Who maintains thejaredchapman/evalforge-lite?

+

thejaredchapman/evalforge-lite is maintained by thejaredchapman. The last recorded GitHub activity is dated 2026-10-01, with 0 open issues.

Are there alternatives to evalforge-lite?

+

Yes. On ClaudeWave you can browse similar plugins at /categories/plugins, sorted by popularity or recent activity.

Deploy evalforge-lite to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: thejaredchapman/evalforge-lite
[![Featured on ClaudeWave](https://claudewave.com/api/badge/thejaredchapman-evalforge-lite)](https://claudewave.com/repo/thejaredchapman-evalforge-lite)
<a href="https://claudewave.com/repo/thejaredchapman-evalforge-lite"><img src="https://claudewave.com/api/badge/thejaredchapman-evalforge-lite" alt="Featured on ClaudeWave: thejaredchapman/evalforge-lite" width="320" height="64" /></a>

More Plugins

evalforge-lite alternatives