OrangePro local-first CLI + MCP server for behavior mapping, grounded test generation, and dynamic proof.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add orangepro-mcp -- npx -y @orangepro/mcp-server{
"mcpServers": {
"orangepro-mcp": {
"command": "npx",
"args": ["-y", "@orangepro/mcp-server"]
}
}
}MCP Servers overview
<p align="center">
<img src="https://github.com/OrangeproAI/orangepro-mcp/raw/main/docs/logo-horizontal.svg" alt="OrangePro" width="320" />
</p>
<p align="center">
<strong>See which code your tests really protect. Prove it by breaking the code on purpose.</strong>
</p>
<p align="center">
<a href="https://www.npmjs.com/package/@orangepro/mcp-server"><img src="https://badge.fury.io/js/@orangepro%2Fmcp-server.svg" alt="npm version" /></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-green.svg" alt="MIT License" /></a>
<a href="https://www.npmjs.com/package/@orangepro/mcp-server"><img src="https://img.shields.io/npm/dw/@orangepro/mcp-server.svg" alt="npm downloads" /></a>
<a href="https://glama.ai/mcp/servers/OrangeproAI/orangepro-mcp"><img src="https://glama.ai/mcp/servers/OrangeproAI/orangepro-mcp/badges/score.svg" alt="Glama score" /></a>
<a href="https://registry.modelcontextprotocol.io/?q=orangepro"><img src="https://img.shields.io/badge/MCP_Registry-orangepro-orange.svg" alt="MCP Registry" /></a>
</p>
---
OrangePro maps every public function in your repository, links each one to the tests that actually call it, and ranks what's left by what it can break. For a test you care about, it breaks the function in an isolated copy and checks that the test fails. It runs on your machine, uses no AI model for scoring, and sends nothing to OrangePro. Model calls happen only if you add your own key, for optional test generation and suggestions.
```bash
npx -y @orangepro/mcp-server@latest start .
```
---
## Contents
- [What you get](#what-you-get)
- [Why developers use it](#why-developers-use-it)
- [Quick start](#quick-start)
- [Evidence tiers](#evidence-tiers)
- [How the ranking works](#how-the-ranking-works)
- [Prove a test](#prove-a-test)
- [Use with your coding agent](#use-with-your-coding-agent)
- [Configuration](#configuration)
- [Language support](#language-support)
- [Privacy and network use](#privacy-and-network-use)
- [Feedback](#feedback)
- [Reference](#reference)
---
## What you get
Every run writes two reports to `.orangepro/`:
| File | For | What's in it |
|---|---|---|
| `short_behavior-coverage.html` | Leads, reviewers, anyone in a hurry | One page: the headline, where to start, the code paths that delete data and their test evidence, the top-ranked items, the tests behind each proof, and how to reproduce the run. Each name links to its line at the analysed commit (GitHub and GitLab). |
| `behavior-coverage.html` | Developers | The full interactive map: every behavior and its evidence, flows from entry points through services, the ranked list with suggested tests, and the settings used. |
### The one-page summary
<img width="880" alt="OrangePro one-page summary of its own repository: 60 functions mapped, 45 linked to a test, 1 proven; the three highest-ranked items without test evidence; and the proof for OrangeProClient.get" src="https://github.com/OrangeproAI/orangepro-mcp/raw/main/docs/images/short-report.png" />
*OrangePro run on its own repository at commit `3cc5c19`, with no AI model. The proof at the bottom replaced `OrangeProClient.get` with a fixed return value in an isolated copy, and the project's own test, unchanged, failed at its assertion.*
### The detailed report
**<a href="https://orangeproai.github.io/orangepro-mcp/twenty-crm-behavior-coverage.html" target="_blank">→ Live example: Twenty CRM (5,237 behaviors mapped)</a>**
<img width="895" alt="OrangePro system map: entry lanes, services, evidence tiers" src="https://github.com/user-attachments/assets/1ceba779-e0ec-4ec1-99ce-001bc3589b42" />
*System map: entry lanes (GraphQL, HTTP, jobs) flowing into services, sized by traffic, colored by evidence tier, red-ringed by risk.*
---
## Why developers use it
- **It tells you what a test actually checks, not just what it runs.** A test that calls a function isn't necessarily checking it. `opro prove` replaces the function body with a fixed return value in an isolated copy and reruns your own test, unchanged. If the test still passes, it wasn't protecting that function.
- **It puts consequences first.** Two lists sit above the ranking:
- **Can destroy data:** code paths that reach a delete or purge with no proven test. Cache evictions don't count.
- **Changing fast:** code that changes often with nothing proving it works.
- **It's honest about evidence.** A function is "linked" only when a test calls that exact function. A test with a similar name is a lead, never coverage. Tests that only check mocks don't count.
- **It's repeatable.** Same commit, full git history, same config and same version give the same ranking. Each report records fingerprints, so you can tell when a change in results came from the code and when it came from the tool.
- **Your coding agent can drive it.** As an MCP server it gives Claude Code, Cursor, Copilot, Codex and others the ranked gaps, the suggested test location and the proof step in one loop.
- **It's free, local and open source (MIT).** No account, no API key for analysis, and no upload.
---
## Quick start
```bash
cd /path/to/your/repo
# install the repo's own dependencies first (npm ci, uv sync, go mod download, ...)
npx -y @orangepro/mcp-server@latest start .
open .orangepro/short_behavior-coverage.html
```
Analysis, ranking and proof need no model key. Test generation is optional and uses your own key (see [Model setup](#reference)).
**Tips for the most accurate run:**
- **Use a full clone, not a shallow one.** Change history drives part of the ranking, and the report says when history was partial.
- **Exclude what isn't product code** with `rank_exclude_paths` (see [Configuration](#configuration)).
- **Rerun after a change.** The detailed report shows what entered, moved up or got resolved since the last run.
**What's written:**
```
.orangepro/
├── short_behavior-coverage.html ← one-page summary
├── behavior-coverage.html ← full interactive report
├── graph.json ← the evidence graph (deterministic)
├── ledger.json ← proof certificates
├── rtm.md ← traceability matrix
└── config.json ← optional per-repo settings
orangepro_generated/ ← generated tests (only with a model key); your files are never edited
```
---
## Evidence tiers
Every behavior gets exactly one tier.
| Tier | What it means |
|---|---|
| **Dynamically Proven** | A test passed on the original code and failed at its own assertion when this function was broken in an isolated copy. |
| **Runtime-covered** | A coverage tool you ran executed this code. |
| **Statically Linked** | A test calls this exact function. That's a structural link, not proof. |
| **Name match only** | A test with a similar name exists. That's a lead, not evidence. |
| **No test found** | Nothing links a test to it. |
**What OrangePro doesn't count, so linked numbers are a floor, not a coverage percentage:**
- tests that only assert on mocks;
- calls made over HTTP or from end-to-end suites;
- in Python, objects that reach a test only through a fixture.
Check the tests before writing new ones for a flagged path.
> **"Dynamically Proven 0" is normal on a first run.** Proof runs your tests, so it happens only for the functions you choose, or within the attempt budget of `opro start`.
---
## How the ranking works
Each unproven function gets an OrangePro Risk Score, **ORS = P × I × D**:
- **P**: how likely it is to change. Change history, fan-out, new code and size.
- **I**: what it can break. Incoming references, entry-point position, data sensitivity, and a floor for paths that reach a delete or purge within two calls.
- **D**: how hard a break would be to notice. Evidence tier, and whether it runs unattended (jobs and schedulers).
The score sets the order of work. It is not a defect probability. The weights are fixed and no AI model is involved. The report shows the inputs behind every row.
---
## Prove a test
```bash
# Python: the mutation value is derived from the function (return annotation or its observed result)
opro prove-loop --target-symbol 'sym:app/billing/invoices.py#void_invoice' \
--test 'tests/test_invoices.py::test_void_marks_invoice_void' --replacement sentinel
# TypeScript / JavaScript: give the inert body to substitute
opro prove-loop --target-symbol 'sym:src/orders.service.ts#OrdersService.cancel' \
--test src/orders.service.spec.ts --replacement 'return null;'
```
- **Proven:** the test passes on the original code and fails at its own assertion on the broken copy.
- **Not proven:** the test still passes on the broken copy, so it doesn't protect that function. That's a finding too.
- **Unrunnable:** setup failed. This is never counted either way.
**Find weak tests without a model key:**
```bash
opro roast . # passing tests whose targeted mutant still survives
```
---
## Use with your coding agent
OrangePro runs as an MCP server. Add it to your client's config:
```json
{
"mcpServers": {
"orangepro-local": {
"command": "npx",
"args": ["-y", "@orangepro/mcp-server@latest", "mcp"]
}
}
}
```
| Client | Where to put it |
| --- | --- |
| Claude Code | `.mcp.json` or `~/.claude.json` |
| Cursor | `~/.cursor/mcp.json` or Settings → MCP |
| VS Code / Copilot | MCP settings |
| Codex / OpenCode / Windsurf | `npx -y @orangepro/mcp-server@latest agent --client codex` prints the setup |
**A prompt that runs the whole loop:**
> "Use `orangepro_start`, then `orangepro_find_test_gaps`. For the top gap, write a test at the suggested path, run it, then call `orangepro_prove_loop` and tell me whether it was proven."
---
## Configuration
Optional. Put it in `.orangepro/config.json` in the repository, or in `~/.orangepro/config.json` for defaults across repositories. Every setting that changes the ranking is shown in the report.
```json
{
"classification": {
"rank_exclude_paths": ["ui/**", "docs/**", "scripts/What people ask about orangepro-mcp
What is OrangeproAI/orangepro-mcp?
+
OrangeproAI/orangepro-mcp is mcp servers for the Claude AI ecosystem. OrangePro local-first CLI + MCP server for behavior mapping, grounded test generation, and dynamic proof. It has 18 GitHub stars and its last recorded update is dated 2026-10-09.
How do I install orangepro-mcp?
+
You can install orangepro-mcp by cloning the repository (https://github.com/OrangeproAI/orangepro-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is OrangeproAI/orangepro-mcp safe to use?
+
Our security agent has analyzed OrangeproAI/orangepro-mcp and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains OrangeproAI/orangepro-mcp?
+
OrangeproAI/orangepro-mcp is maintained by OrangeproAI. The last recorded GitHub activity is dated 2026-10-09, with 0 open issues.
Are there alternatives to orangepro-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy orangepro-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/orangeproai-orangepro-mcp)<a href="https://claudewave.com/repo/orangeproai-orangepro-mcp"><img src="https://claudewave.com/api/badge/orangeproai-orangepro-mcp" alt="Featured on ClaudeWave: OrangeproAI/orangepro-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.