ControlKeel MCP Server
io.github.aryaminus/controlkeel
Agent control plane for governed engineering: policy validation, findings tracking, and review gates for AI workflows.
What is the ControlKeel MCP server?
ControlKeel is an agent control plane MCP server that enforces policy, validates agent output, and tracks findings before risky work reaches production. It captures intent and policy as typed memory, runs deterministic checks and optional human reviews, and persists evidence across sessions to prevent rediscovery of domain knowledge.
ControlKeel sits between your AI coding agents and production, learning your team's intent rules, review preferences, and delivery habits. It transforms domain knowledge from documentation into enforceable policy checks, deterministic validation, and human-gated approval workflows. Use it to catch policy violations before they reach main, maintain consistency across agent sessions, and build regression evidence for recurring behaviors.
How to install ControlKeel
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
ck_attach— Attach ControlKeel to a supported agent host (OpenCode, Copilot, Claude, Codex, etc.)controlkeel setup— Initialize ControlKeel for a repository with project or user scopecontrolkeel attach— Bind ControlKeel to a specific agent host or runtimecontrolkeel attach doctor— Diagnose and verify ControlKeel attachment healthcontrolkeel provider doctor— Validate provider configuration and connectivitycontrolkeel status— Display current governance state and configurationcontrolkeel findings— List and review policy violations and validation findingscontrolkeel validate— Run deterministic policy checks on agent outputBenchmark engine— Persisted benchmark suite for regression evidence and eval scoring
Use cases
- Validate AI agent code changes against team policy before merging to main
- Track and review high-risk actions with human approval gates when policy requires it
- Persist domain knowledge and governance rules across agent sessions to avoid re-explaining context
- Catch policy violations deterministically with zero provider tokens using local scanners
- Build regression evidence from recurring behaviors to automate known-safe patterns
ControlKeel MCP server FAQ
ControlKeel is an MCP server that acts as a control plane for AI agents, enforcing policy validation, tracking findings, and gating risky actions. It learns your team's intent rules and delivery habits, turning them into typed memory and deterministic checks that persist across sessions.
Yes, ControlKeel is open-source and available via npm (@aryaminus/controlkeel), Homebrew, or direct download. Local governance, CLI, and MCP tools are included; optional cloud telemetry and team features are available but dormant by default.
Install the CLI via `npm i -g @aryaminus/controlkeel`, `brew install controlkeel`, or the platform-specific installer. Run `controlkeel setup` in your repository, then `controlkeel attach <host>` (e.g., `controlkeel attach opencode` for Cursor). Alternatively, paste the one-line setup prompt into your agent to automate the process.
Local project-scoped governance requires no external auth. Optional cloud features (team operations, workspace sync, cloud evidence) use OAuth-scoped credentials and are opt-in; they remain dormant until explicitly configured.
ControlKeel supports OpenCode (Cursor), GitHub Copilot, Claude, Codex, and other MCP-compatible hosts. See docs/support-matrix.md and docs/agent-integrations.md for the canonical host inventory and integration tiers.
ControlKeel runs deterministic validation checks on agent output before risky work reaches production. It catches policy violations with zero provider tokens, optionally gates high-impact actions for human review, and persists findings and proof bundles for audit and regression evidence.
README (reference)
Source of truth, from the repository.
ControlKeel
Turn the way your team works into enforceable memory for AI agents. - @arya_minus
ControlKeel is an agent control plane for day-to-day governed engineering. Through observation, findings and evaluation, it learns your intent rules, review taste and delivery habits, turning them into typed memory, policy checks and proof bundles. CK sits between your coding agents and production as a portable "company brain": comparing intended delivery against actual delivery and turning raw agent intent into policy-validated tasks.
If you're using an AI agent today, you probably have an *.md telling it how to behave. But a rules/specs file is just a promise made to the model. ControlKeel enforces the output. Beyond just catching bugs, CK solves the "Unknown Unknowns" problem: having to re-explain your domain knowledge in every single session.
Product loop
- Capture intent and policy — scope, risk, budget, domain pack, and human taste become CK state.
- Validate agent output — deterministic checks and optional advisory review produce findings before risky work reaches main.
- Gate only when needed — humans approve high-impact actions when intent, risk, or policy requires it.
- Persist evidence — findings, reviews, proofs, memory, cost, and task outcomes survive host switches.
- Improve with evals — traces and recurring failures become bounded regression evidence for specific suites and subjects.
The operating rule is simple: spend tokens on discovery, not rediscovery. Recurring behavior moves through a human-gated deterministic promotion path; once its regression evidence passes, checks, APIs, CLIs, or workflows handle the known case and agents handle only exceptions and genuinely new uncertainty.
ControlKeel transforms your domain knowledge from "raw" intent and "shelfware" documentation into a living system that remembers, enforces, and evolves.
Quick start
One-line setup via your agent
Copy/paste this into your agent (OpenCode, Codex, Claude, or another supported host):
Set up ControlKeel for this repository. Read and follow https://raw.githubusercontent.com/aryaminus/controlkeel/main/README.md, https://raw.githubusercontent.com/aryaminus/controlkeel/main/docs/getting-started.md, https://raw.githubusercontent.com/aryaminus/controlkeel/main/docs/support-matrix.md, and https://raw.githubusercontent.com/aryaminus/controlkeel/main/docs/agent-integrations.md. Install ControlKeel if missing, run `controlkeel setup`, detect this agent host, attach the strongest supported path with `controlkeel attach <host>`, then run `controlkeel attach doctor`, `controlkeel provider doctor`, `controlkeel status`, `controlkeel findings`, and the host-native MCP check. If CK is available only as MCP, call `ck_attach` for this host. Apply only safe local fixes and redact secrets from logs. Pause and ask before continuing if the host needs workspace trust, manual provider configuration, a restart after attach/plugin changes, or a plan-review approval that cannot auto-wait. Ensure the project is trusted and restart the host after attach/plugin changes.
CLI install
Install the CLI:
brew tap aryaminus/controlkeel && brew install controlkeel
# or
npm i -g @aryaminus/controlkeel
# or
curl -fsSL https://github.com/aryaminus/controlkeel/releases/latest/download/install.sh | sh
Windows PowerShell:
irm https://github.com/aryaminus/controlkeel/releases/latest/download/install.ps1 | iex
First governed run:
controlkeel
controlkeel setup
controlkeel attach opencode # project scope by default; use another supported host as needed
controlkeel attach doctor
controlkeel provider doctor
controlkeel status
controlkeel findings
Run setup and attach from the repository you want to govern. Project scope writes host files only inside that repository. --scope user writes host-level files under your user configuration only for targets that explicitly support user scope; it does not turn project binding or proof state into global state.
For the complete first-run path, use docs/getting-started.md. For host truth, use docs/support-matrix.md and docs/agent-integrations.md.
Benchmark-backed evidence
ControlKeel includes a persisted benchmark engine. Current user-facing evidence is bounded to the named suite, subject, and scoring definition below; docs/benchmarks.md is the canonical reference for full tables, caveats, JSON exports, and agent-host protocols.
Verified with-vs-without-CK baseline (host_comparison_v1, 12 risky scenarios)
Verified with ControlKeel 0.3.45:
- Risky suite
host_comparison_v1:null_policy_baselinecaught 0/12;controlkeel_validatecaught 12/12, blocked 9/12, and hit expected rules 9/12 with median deterministic validation time 52 ms, 0 provider tokens. - Paired benign suite
benign_baseline_v1:controlkeel_validateproduced 0/10 catches, 0/10 blocks, FPR 0.000, median deterministic validation time 42 ms, 0 provider tokens.
Read the numbers precisely: deterministic scanner evidence is not the same as model-backed agent-host evidence. Reproduction commands and the OpenCode/Copilot/Claude/Codex comparison protocol live in docs/benchmarks.md.
What ships today
- Local governance: CLI, the full local stdio MCP tool set, project binding, host attach/export bundles, scanner validation, findings, reviews, proof bundles, budgets, and typed memory.
- Host and runtime support: native attach for supported hosts, runtime exports for headless/outer-loop systems, a narrower OAuth-scoped hosted MCP set, minimal A2A, and fallback validation/proxy paths.
- Team/project operations: org membership, invitations, workspace GitHub repo bindings, service accounts, webhooks, workspace tool policy, and policy-set APIs.
- Cloud evidence paths: opt-in cloud telemetry, workspace keys, cloud run packages, runtime callbacks, and dormant-until-configured bidirectional sync for findings, reviews, digests, and memory records.
- Observability loop: timelines, memory quality, costs, trends, problem clusters, eval candidates, benchmark drafts/history, and promotion advisories.
Docs map
- docs/README.md — documentation map by job
- docs/getting-started.md — install to first finding
- docs/support-matrix.md — canonical host/protocol inventory
- docs/agent-integrations.md — integration mechanisms and support tiers
- docs/benchmarks.md — benchmark scoring, metadata, and claim discipline
- docs/observability-feedback-loop.md — local evidence-to-regression loop
- docs/control-plane-claim-matrix.md — README claim-to-test matrix for governance, memory, cloud sync, and human gates
- docs/api-reference.md and docs/cli-reference.md — code-aligned surfaces
- docs/packages.md — package and distribution catalog
- docs/self-hosting.md — self-host deployment guidance
Development
mix setup
mix phx.server
mix test
mix precommit
Phoenix + Ecto on SQLite. Uses Req for HTTP. Single-binary builds ship through Burrito and GitHub Releases.
Related MCP servers

Multi-tradition astrology. 19 modes, 18 tools. Deterministic ephemeris engine.
Deterministic local-first context and impact maps for coding agents from tasks, issues, and diffs.

Database-driven FP enforcement and project management for AI-maintained codebases

PWA Debug Layer
Debug PWAs in your real browser via MCP: service-worker, cache, installability & framework state.

Infimium
Private AI context layer for semantic code search, dependency graphs, memory, and workspaces—100% local, zero token bloat.
Tells you what leaked in a .har capture, who's on the page, and what to strip. No network calls.

