PluginBench
MCP Server
Active
MIT

OrangePro MCP Server

io.github.OrangeproAI/orangepro

Find test gaps, generate grounded tests, and prove behavior with mutation testing.

What is the OrangePro MCP server?

OrangePro is an MCP server that maps every public behavior in your codebase, scores each by real test evidence, and identifies structural blind spots before users find them. It generates grounded tests for top gaps and uses mutation testing to dynamically prove that tests actually catch behavior changes. Runs locally with no external dependencies.

OrangePro analyzes your code to create an interactive behavior-coverage report showing which functions and flows are tested, which are only partially covered, and which have no test evidence. It ranks untested behaviors by risk, generates candidate tests for the highest-impact gaps, and confirms test quality through mutation killing. Use it to close test gaps before they reach production.

How to install OrangePro

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • OPENAI_API_KEY
    secret

    Optional OpenAI BYOK key for AI candidate links and test generation.

  • ANTHROPIC_API_KEY
    secret

    Optional Anthropic BYOK key for AI candidate links and test generation.

  • OLLAMA_BASE_URL

    Optional Ollama endpoint for local AI candidate links and test generation.

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "orangepro": {
      "command": "npx",
      "args": [
        "-y",
        "@orangepro/mcp-server",
        "mcp"
      ],
      "env": {
        "OPENAI_API_KEY": "<YOUR_OPENAI_API_KEY>",
        "ANTHROPIC_API_KEY": "<YOUR_ANTHROPIC_API_KEY>",
        "OLLAMA_BASE_URL": "<YOUR_OLLAMA_BASE_URL>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • orangepro_start — One-command setup: analyze codebase, build evidence graph, generate report, and suggest next actions
  • orangepro_analyze_sources — Build or refresh the evidence graph from source code
  • orangepro_generate_tests — Generate grounded tests for identified gaps, optionally scoped to a branch diff
  • orangepro_prove — Run mutation-kill oracle to confirm a test actually catches behavior changes
  • orangepro_prove_loop — Setup, dynamic proof, and report refresh for a single behavior
  • orangepro_find_test_gaps — List behaviors with weak or missing tests, ranked by risk
  • orangepro_graph_score — Calculate graph readiness score (0–100)
  • orangepro_status — Check workspace state without generating artifacts
  • orangepro_doctor — Recommend next evidence to improve test quality
  • orangepro_rtm — Generate requirements traceability matrix
  • orangepro_stats — Aggregate statistics across the codebase
  • orangepro_changed_impact — Show what a diff touches and its test coverage impact
  • orangepro_record_run — Record a test run result for evidence tracking
  • orangepro_explain_test — Explain why a test was generated and what it targets
  • orangepro_export_evidence_pack — Export metadata-only evidence pack for external use
  • orangepro_update_graph — Incrementally update the evidence graph
  • orangepro_ai_links — Discover weak behavior-to-symbol suggestions using AI
  • orangepro_ai_flows — Find candidate service-boundary flows using AI

Use cases

  • Identify untested public functions and flows ranked by blast radius before shipping to production
  • Generate and run candidate tests for the highest-risk gaps, then prove they catch real behavior changes
  • Track test coverage deltas across PRs to prevent regressions from entering the codebase
  • Analyze multi-service integration flows to find untested cross-boundary behaviors
  • Score code readiness (0–100) and get recommendations for what evidence to add next

OrangePro MCP server FAQ

What is OrangePro and how does it work?

OrangePro maps every public behavior in your codebase, assigns evidence tiers (Dynamically Proven, Runtime-covered, Statically Linked, Unconfirmed Candidate, or No Signal), and generates an interactive HTML report showing test gaps ranked by risk. It uses mutation testing to confirm that tests actually catch behavior changes.

Is OrangePro free?

Yes. The local tool is free and open-source (MIT license). Analysis, scoring, and mutation proof require no API key. Test generation is optional and uses your own API key (BYOK) if you configure one.

How do I install OrangePro in Cursor or Claude?

Add to your MCP config (Cursor: ~/.cursor/mcp.json, Claude Code: .mcp.json): {"mcpServers": {"orangepro-local": {"command": "npx", "args": ["-y", "@orangepro/mcp-server@latest", "mcp"]}}}. Then tell your agent to use orangepro_start and orangepro_generate_tests.

What authentication or API keys are required?

None for analysis, scoring, or mutation proof. Test generation is optional and requires one of: ANTHROPIC_API_KEY, OPENAI_API_KEY, or OLLAMA_BASE_URL (local, no key needed). Keys are read from environment at runtime and never persisted.

What languages does OrangePro support?

Static behavior mapping works for TypeScript, JavaScript, Python, Go, Java, Kotlin, Rust, PHP, C#, Ruby, Swift, C, and C++. Test generation and dynamic proof are currently supported for TypeScript/JavaScript (Jest/Vitest/Mocha), Python (pytest), Go (*_test.go), and Java (JUnit 4/5); other languages are planned.

Does OrangePro upload my code anywhere?

No. It runs entirely locally, reads code in-process, and never uploads to any OrangePro server. Your source files are never modified. If you configure a model key, code context goes directly to your chosen provider (OpenAI, Anthropic, or local Ollama), not through OrangePro.

README (reference)

Source of truth, from the repository.

<p align="center"> <img src="https://github.com/OrangeproAI/orangepro-mcp/raw/main/docs/logo-horizontal.svg" alt="OrangePro" width="320" /> </p> <p align="center"> <strong>Find the behaviors your tests miss. Generate grounded tests that actually run.</strong> </p> <p align="center"> <a href="https://www.npmjs.com/package/@orangepro/mcp-server"><img src="https://badge.fury.io/js/@orangepro%2Fmcp-server.svg" alt="npm version" /></a> <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-green.svg" alt="MIT License" /></a> <a href="https://www.npmjs.com/package/@orangepro/mcp-server"><img src="https://img.shields.io/npm/dw/@orangepro/mcp-server.svg" alt="npm downloads" /></a> <a href="https://glama.ai/mcp/servers/OrangeproAI/orangepro-mcp"><img src="https://glama.ai/mcp/servers/OrangeproAI/orangepro-mcp/badges/score.svg" alt="Glama score" /></a> <a href="https://registry.modelcontextprotocol.io/?q=orangepro"><img src="https://img.shields.io/badge/MCP_Registry-orangepro-orange.svg" alt="MCP Registry" /></a> </p>

OrangePro maps every public behavior in your codebase, scores each one by real test evidence, and shows you the structural blind spots before your users find them. Runs locally. Your code never leaves your machine.

npx -y @orangepro/mcp-server@latest start .
<!-- TODO: Replace with a terminal GIF showing the command running and report opening -->

Table of Contents


What you get

One command produces an interactive HTML report:

npx -y @orangepro/mcp-server@latest start .
open .orangepro/behavior-coverage.html

The report has two modes: Simple (integration-level blind spots, plain English) and Expert (full behavior list, evidence tiers, flows, system map). Toggle with the pill switch at the top.

<a href="https://orangeproai.github.io/orangepro-mcp/twenty-crm-behavior-coverage.html" target="_blank">→ Live example: Twenty CRM (5,237 behaviors mapped)</a>

<img width="895" alt="OrangePro system map — entry lanes, services, evidence tiers" src="https://github.com/user-attachments/assets/1ceba779-e0ec-4ec1-99ce-001bc3589b42](https://github.com/user-attachments/assets/a4d85b98-4f19-4647-8dd9-db5911574f49" />

System map — entry lanes (GraphQL, HTTP, Jobs) flowing into services, sized by traffic, colored by evidence tier, red-ringed by risk.

<img width="818" alt="Priority gaps" src="https://github.com/user-attachments/assets/30a512b6-7830-48db-a00f-a616e7176ea8" />

Priority gaps of another open source Project HONO — top 20 unproven behaviors ranked by blast radius, with generated test drafts.


Evidence tiers

Every behavior gets exactly one tier. Nothing is labeled "tested" on faith.

TierColorWhat it means
Dynamically Proven🟢A real test kills a targeted mutation of this behavior
Runtime-covered🟢Coverage tool executed this code
Statically Linked🟡A test imports and calls this code — structural link, not proof
Unconfirmed Candidate⚪A similar test file exists — a lead, not evidence
No Signal🔴Nothing tests this behavior

"Dynamically Proven 0" is normal on first run. Proof requires running tests against targeted mutations. That's the trust model.


Quick start

cd /path/to/your/repo
npm install          # install the repo's own dependencies first

npx -y @orangepro/mcp-server@latest start .
open .orangepro/behavior-coverage.html

No API key needed. The report shows your system map, evidence tiers, priority gaps, and delta since last run.

Want test generation? Add a model key (BYOK):

export ANTHROPIC_API_KEY="..."   # or OPENAI_API_KEY / OLLAMA_BASE_URL
npx -y @orangepro/mcp-server@latest start .

AI output never changes evidence tiers. Only the mutation-kill oracle can mint Dynamically Proven.

Output:

.orangepro/
├── behavior-coverage.html   ← open this
├── graph.json               ← deterministic evidence graph
├── COVERAGE_REPORT.md       ← coverage and gap summary
└── ai/                      ← candidate flows (when a key is configured)

orangepro_generated/         ← generated tests; your source files are never touched

Each rerun shows a delta banner: what entered the codebase, what moved up in risk, what got resolved.


Use with your coding agent

OrangePro runs as an MCP server. Add to your client's config:

{
  "mcpServers": {
    "orangepro-local": {
      "command": "npx",
      "args": ["-y", "@orangepro/mcp-server@latest", "mcp"]
    }
  }
}
ClientWhere to put it
Claude Code.mcp.json or ~/.claude.json
Cursor~/.cursor/mcp.json or Settings → MCP
VS Code / CopilotMCP settings
Codex / OpenCodeRun npx -y @orangepro/mcp-server@latest agent --client codex

The workflow: Tell your agent:

"Use orangepro_start, then orangepro_generate_tests with base_ref=main. Write each test to its suggested_path, run it, and report pass/fail."

The agent writes the test, runs it, calls orangepro_prove, and the behavior turns Dynamically Proven. One prompt, full loop.


Works with

<p> <strong>Claude Code</strong> · <strong>Cursor</strong> · <strong>GitHub Copilot</strong> · <strong>Codex</strong> · <strong>Windsurf</strong> · <strong>OpenCode</strong> · <strong>VS Code</strong> </p>

Any MCP-compatible agent can drive OrangePro. No vendor lock-in.


How it works

┌─────────────┐     ┌──────────────┐     ┌─────────────┐
│  Your Code  │ ──► │  Knowledge   │ ──► │  Evidence   │
│  (any lang) │     │    Graph     │     │   Tiers     │
└─────────────┘     └──────────────┘     └─────────────┘
                           │
                    ┌──────┴──────┐
                    ▼             ▼
             ┌───────────┐  ┌──────────┐
             │ Gap Report│  │ Generate │
             │ + Risks   │  │  Tests   │
             └───────────┘  └──────────┘
PhaseWhat happensNeeds a model key?
AnalyzeAST walk → behaviors, flows, evidence tiersNo
ScoreGraph readiness score (0–100)No
GenerateGrounded tests for top gapsYes (BYOK)
ProveMutation-kill oracle confirms test breaks if behavior changesNo

Same code = same score. Deterministic. Always.


Language support

LanguageStatic mappingGenerated testsDynamic proof
TypeScript / JavaScript✓✓ Jest / Vitest / Mocha✓
Python✓✓ pytest✓
Go✓✓ *_test.go✓
Java✓✓ JUnit 4/5✓
Kotlin, Rust, PHP, C#, Ruby, Swift, C, C++✓plannedplanned

Static mapping works across many languages via tree-sitter. Dynamic proof is deliberately narrower — each language needs a runner, mutation locator, and sandbox profile.


Highest-value local run

Use the repository's own setup and test commands first, and keep unit and integration coverage in separate artifacts. Then run opro start; it performs analysis, ingests the artifacts, attempts targeted proof, generates report-visible drafts, and writes the final report. A separate opro analyze is unnecessary when opro start follows it.

# 1. Install/build exactly as the repository documents.
# 2. Run the repository's unit and integration coverage commands separately.
# 3. Record artifact provenance (example paths and commands):
mkdir -p .orangepro
# create .orangepro/coverage-suites.json using the schema below

opro coverage .                    # optional preflight: discover/generate artifacts
opro start . --proof-limit 5 --generate-limit 20
{
  "artifacts": {
    ".orangepro/coverage/unit.coverprofile": {
      "suite": "unit",
      "command": "make unit-test-coverage"
    },
    ".orangepro/coverage/integration.coverprofile": {
      "suite": "integration",
      "command": "make integration-test-coverage"
    }
  }
}

Without this manifest, OrangePro conservatively infers clear unit/integration names and labels everything else unclassified; it never guesses that an aggregate profile is unit-only. The report shows unit, integration, their overlap, unclassified coverage, and the combined union separately. --proof-limit controls dynamic proof attempts (which may draft a test for proof); --generate-limit independently controls the additional report-visible risk-gap drafting lane. A generation run also records its terminal status and exact reason, so a compiler/import failure is not misreported as a generic dependency problem.


Privacy

  • No stored source. Reads code in-process. Never uploads to an OrangePro server.
  • No existing-source mutation. Never edits your source or test files.
  • Your keys stay yours. Read from env at call time, never persisted.
  • BYOK is direct. Code context goes to the model provider you configure. OrangePro is not in that path.

<details> <summary><strong>CLI reference</strong></summary>
opro                          # analyze + report + agent next actions
opro start --base main        # same, scoped to a branch diff
opro analyze                  # build the evidence graph
opro score                    # graph readiness (0–100)
opro gaps --limit 10          # top 10 untested behaviors
opro generate --base main     # tests for PR diff
opro generate --single        # top gap, whole repo
opro prove                    # mutation-kill oracle
opro rtm                      # traceability matrix
opro export                   # metadata-only evidence pack
opro mcp                      # run as MCP server (stdio)
opro doctor                   # what evidence to add next
opro coverage                 # discover/generate artifacts; analyze or start ingests them

Add --json to any read command for machine output. Run opro help for the full reference.

</details> <details> <summary><strong>MCP tools (18 total)</strong></summary>
ToolWhat it does
orangepro_startOne-command setup: analyze + report + next actions
orangepro_analyze_sourcesBuild/refresh the evidence graph
orangepro_generate_testsGenerate grounded tests for gaps
orangepro_proveRun mutation-kill oracle on a behavior
orangepro_prove_loopSetup + dynamic proof + report refresh for one behavior
orangepro_find_test_gapsList behaviors with weak/missing tests, ranked by risk
orangepro_graph_scoreGraph readiness score (0–100)
orangepro_statusWorkspace state without generating anything
orangepro_doctorRecommend next evidence to improve quality
orangepro_rtmRequirements traceability matrix
orangepro_statsAggregate statistics
orangepro_changed_impactWhat a diff touches (requires git + base ref)
orangepro_record_runRecord a test run result
orangepro_explain_testExplain why a test was generated
orangepro_export_evidence_packExport metadata-only evidence pack
orangepro_update_graphIncremental graph update
orangepro_ai_linksWeak behavior→symbol suggestions (optional AI)
orangepro_ai_flowsCandidate flow discovery (optional AI)
</details> <details> <summary><strong>PR workflow</strong></summary>
opro generate --base main              # tests for what this branch changed
opro generate --pr 1234                # checks out PR #1234
opro generate --changed                # current branch diff vs main

Each generated test includes:

  • Grounding — the real files, symbols, and existing tests it cites
  • Run hints — where to write it, how to run it
  • Scenario bucket — what failure mode it targets

If dependencies aren't installed, tests are kept as Manual tests (Given/When/Then steps with the blocker named). Install dependencies and re-run to convert them to runnable tests.

</details> <details> <summary><strong>Test categories</strong></summary>

Generation is evidence-gated. A category is produced only when the graph has supporting evidence.

CategoryWhat it targets
Happy pathPrimary expected behavior
Validation errorBad/invalid input handling
Edge caseBoundaries, empty/null, concurrency, retries
Integration flowMulti-step behavior across services
Security / privacyAuth, injection, data leakage
RegressionPinning a previously-broken behavior
</details> <details> <summary><strong>Model setup (BYOK)</strong></summary>

Analysis, scoring, and proof need no model key. Generation does.

ProviderEnvironment variable
OpenAI-compatibleOPENAI_API_KEY (optional: OPENAI_BASE_URL, OPENAI_MODEL)
AnthropicANTHROPIC_API_KEY (optional: ANTHROPIC_MODEL)
Ollama (local, no key)OLLAMA_BASE_URL (optional: OLLAMA_MODEL)

Auto-detect order: OpenAI → Ollama → Anthropic. Override with --provider and --model. The defaults are gpt-5.3-codex for OpenAI and claude-sonnet-5 for Anthropic.

Run opro setup to configure interactively. Keys stay in your environment — never written to graph, config, or artifacts.

</details> <details> <summary><strong>AI candidate lanes</strong></summary>

With a provider key, OrangePro stages weak AI behavior→symbol links and AI-suggested candidate flows. These are review/generation worklists, not evidence:

  • AI links appear as AI-linked suggestions.
  • AI flows are stored separately from deterministic flows.
  • Neither lane changes evidence tiers or denominator counts.

Use them when you want the agent to find likely service-boundary flows faster; ignore them for a deterministic-only report.

</details>

What's on the hosted platform

This repo is the free local tool. The OrangePro platform adds:

  • Persistent knowledge graph across PRs and repos
  • PR/CI policy gates over evidence tiers and risk deltas
  • Jira / Confluence / TestRail / OpenAPI enrichment
  • Cross-repo intelligence and recurring-flow memory
  • Production incident correlation and regression targeting
  • Team dashboards and test lifecycle management

Contributing

git clone https://github.com/OrangeproAI/orangepro-mcp.git
cd orangepro-mcp && npm ci && npm run build
npm test

PRs welcome. Please open an issue first for large changes.


<p align="center"> MIT License · <a href="https://orangepro.ai">orangepro.ai</a> </p>

Related MCP servers

MCP server for controlling Discord servers via bot token

0
TypeScript
MIT
View repository →

Automate Google Ad Manager: campaigns, line items, creatives, inventory, reporting — 51 tools.

2
Python
MIT
View repository →

Browse, subscribe to, and call 3,000+ APIs autonomously. x402 USDC pay-per-call on Base.

View repository →

8,000+ pay-per-call APIs via x402 USDC micropayments on Base. No API keys needed.

1
JavaScript
MIT
View repository →

Regulatory intelligence (FDA, ICH, EMA) and sanctions screening (OFAC, EU, UK) for Claude

0
View repository →

MCP server for Orderly Network documentation, SDK patterns, and perpetual futures trading infrastructure

8
JavaScript
MIT
View repository →