PluginBench
MCP Server
Active
Apache-2.0

io.github.sipyourdrink-ltd/bernstein MCP Server

io.github.sipyourdrink-ltd/bernstein

Deterministic orchestrator for CLI coding agents with byte-identical run receipts, 40+ adapters, and air-gap support.

What is the io.github.sipyourdrink-ltd/bernstein MCP server?

Bernstein is a deterministic orchestrator for CLI coding agents (Claude Code, Codex, Gemini CLI, and 40+ others) that runs them in parallel with reproducible task scheduling and verifiable run receipts. It isolates each coding task in its own git worktree, gates task completion on concrete signals like passing tests, and records runs with optional HMAC-chained audit logs for offline verification.

Bernstein orchestrates multi-agent coding workflows with deterministic scheduling, isolation by construction, and cryptographic proof of execution. Unlike systems with LLMs in the coordination loop, Bernstein uses plain Python scheduling so runs are reproducible end-to-end. Each agent task gets its own git worktree, artifacts are verified against declared contracts, and runs can be replayed and audited offline with signed receipts.

How to install io.github.sipyourdrink-ltd/bernstein

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "bernstein": {
      "command": "uvx",
      "args": [
        "bernstein"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • Task decomposition and planning — Breaks goals into tasks with roles, owned files, and completion signals via a single LLM call, then executes deterministically in plain Python
  • Parallel agent execution — Spawns 40+ supported CLI agents (Claude Code, Codex, Gemini CLI, Cursor, Aider, etc.) in isolated git worktrees running concurrently
  • Deterministic replay and verification — Records complete run journals and lineage spines; replays runs to detect non-determinism at the exact step via hash mismatch
  • Signed run receipts — Generates Ed25519-signed receipts binding journal state, lineage spine, and optional audit chains for offline verification without live state
  • HMAC-chained audit log — Optional audit mode (BERNSTEIN_AUDIT=1) records cryptographically-chained receipts proving no tampering of recorded actions
  • Reliability floor evaluation — Runs tasks k times under fixed coordination to compute pass^k reliability floor (all k attempts must pass) with signed sealed results
  • Terminal dashboard (bernstein live) — Two-column TUI showing live agent logs and task board with independent scrolling panes
  • Browser dashboard (bernstein gui serve) — Web UI listing tasks with working-tree diffs and full-width activity feed
  • Workflow manifests — Declarative YAML DAGs of agent, command, and loop nodes for multi-stage plans
  • Artifact mode tasks — Non-code deliverables (reports, datasets, action logs) with working directories and completion via signed lineage receipts
  • Sandbox backends — Optional stricter filesystem enforcement for task isolation
  • Agent catalogs — Point roles at agent definitions outside built-in templates via YAML/SKILL.md directories or Claude Code plugin layouts
  • Datasources — Read-only query receipts with schema snapshot binding for each result
  • Repository hygiene gates — Verify and sync translated README files (readme-l10n) to prevent drift from English source

Use cases

  • Run multiple coding agents in parallel on the same goal, with deterministic scheduling and reproducible results
  • Verify agent work offline by replaying runs and checking cryptographic receipts without live state or SaaS
  • Isolate coding tasks in separate git worktrees to prevent agents from interfering with each other's changes
  • Audit compliance by recording HMAC-chained action logs and generating signed run receipts for regulatory proof
  • Evaluate agent reliability by running tasks k times and computing a pass^k floor (all k attempts must pass)
  • Mix cheap local models and expensive cloud models in the same workflow for cost-optimized multi-agent orchestration

io.github.sipyourdrink-ltd/bernstein MCP server FAQ

What is Bernstein?

Bernstein is a deterministic orchestrator for CLI coding agents that runs them in parallel with reproducible scheduling, isolated git worktrees, and cryptographically-signed run receipts you can verify offline. It supports 40+ agents including Claude Code, Codex, Gemini CLI, Cursor, and Aider.

Is Bernstein free?

Yes, Bernstein is open-source under the Apache 2.0 license. It runs locally on your machine with no SaaS hop or third-party data plane required.

How do I install Bernstein?

Install via `uv tool install bernstein`, `pipx install bernstein`, or `pip install bernstein`. Docker images are available at `ghcr.io/sipyourdrink-ltd/bernstein`. An air-gap wheelhouse is provided for offline installation.

What agents does Bernstein support?

Bernstein has 40+ built-in adapters for Claude Code, Codex CLI, Gemini CLI, GitHub Copilot CLI, Cursor, Aider, Goose, Muse Code, OpenAI Agents SDK, Amp, Cody, Continue, Devin Terminal, Junie, Kilo, Kiro, AWS Q Developer, Ollama, OpenCode, OpenHands, Open Interpreter, gptme, Plandex, AIChat, Letta Code, Qwen, and more. Any agent with a `--prompt` flag works through the generic wrapper.

Does Bernstein require authentication?

Bernstein itself is free and open-source. Authentication depends on the agents you use: Claude Code requires a Claude API key, Codex requires GitHub Copilot, etc. Run `bernstein doctor` to check if a CLI agent is installed and authenticated.

Can I verify a run offline?

Yes. Bernstein generates signed run receipts that bind the journal state, lineage spine, and optional audit chains under a single Ed25519 signature. You can verify the receipt offline with just the file and the operator's public key: `bernstein verify receipt run-receipt.json --public-key key.pem`.

README (reference)

Source of truth, from the repository.

<div align="center"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/sipyourdrink-ltd/bernstein/main/docs/assets/logo-dark.svg"> <source media="(prefers-color-scheme: light)" srcset="https://raw.githubusercontent.com/sipyourdrink-ltd/bernstein/main/docs/assets/logo-light.svg"> <img alt="Bernstein" src="https://raw.githubusercontent.com/sipyourdrink-ltd/bernstein/main/docs/assets/logo-light.svg" width="340"> </picture> <br> <img alt="Bernstein - deterministic multi-agent CLI orchestration" src="https://raw.githubusercontent.com/sipyourdrink-ltd/bernstein/main/docs/assets/banner-readme.webp" width="820"> <br>

"To achieve great things, two things are needed: a plan and not quite enough time." - attributed to Leonard Bernstein

deterministic multi-agent CLI orchestration

CI PyPI GHCR Python 3.12+ License OpenSSF Scorecard CodeQL Open in Codespaces MCP Toplist <a href="https://deepwiki.com/sipyourdrink-ltd/bernstein"><img src="https://deepwiki.com/badge.svg" alt="Ask DeepWiki"></a>

website · docs · install · first run · glossary · limitations · name policy · sponsor

简体中文 · 繁體中文 · 日本語 · 한국어 · हिन्दी · বাংলা

</div>

Status: beta. Solo-maintained, under active development. The version number counts releases, not maturity - minor versions may change interfaces. Pin the version for anything you depend on; regressions get fixed fast, file them.

Bernstein is a deterministic orchestrator for CLI coding agents (Claude Code, Codex, Gemini CLI, and 40+ more). It runs them in parallel, gates what they produce, and records enough of the run that you can check it afterwards. Air-gap install profile included. Apache-2.0.

at a glance

Four things set it apart; everything after is detail.

  • No LLM in the coordination loop. Scheduling is plain Python, so a run is reproducible end to end. Replay yesterday's plan and get yesterday's task graph.
  • Checkable after the fact. The replay journal records every run, and the always-on lineage spine records every lineage-bearing step; the opt-in HMAC-chained audit log (BERNSTEIN_AUDIT=1) adds receipts you verify offline. Non-determinism surfaces as a hash mismatch at the exact step, not a flaky re-run. Non-code deliverables get the same treatment: a task can declare an artifact contract (report, dataset, action log, ops result) and completes on a signed lineage receipt rather than a git commit.
  • Isolated by construction. Each coding task gets its own git worktree behind merge gates; artifact-mode tasks get a working directory under .sdd/workspaces/. Agents share no mutable workspace by default; the only shared state is the task backlog, which is claimed atomically. Stricter filesystem enforcement is opt-in, from the sandbox backends. Disable worktrees and every task runs in the shared checkout.
  • Broad and local. 40+ CLI agent adapters plus a generic --prompt wrapper, file-based state, no SaaS hop, no third-party data plane.

The full list is on the capabilities page; the feature matrix is the exhaustive index.

install in 30 seconds

uv tool install bernstein    # or: pipx install bernstein
bernstein init
bernstein doctor             # checks a CLI agent is installed and authenticated
bernstein -g "fix the failing test in tests/test_foo.py"

pipx, pip, brew, dnf, npm, and Docker are covered in the install guide; the air-gapped wheelhouse has its own air-gap guide.

<img alt="A real bernstein demo run: mock agents fix four seeded bugs, ending on the run's signed receipt verifying offline" src="https://raw.githubusercontent.com/sipyourdrink-ltd/bernstein/main/docs/assets/demo-run/demo.gif" width="820">

The recording above is a real run, and it ships with its own proof. The cast, the signed run receipt derived from that run's journal, and the public key that pins it all live in docs/assets/demo-run/. Verify the run you just watched, offline:

bernstein verify receipt docs/assets/demo-run/run-receipt.json \
    --public-key docs/assets/demo-run/run-receipt.pub.pem

CI re-verifies the committed receipt on every push to main — and proves a tampered copy fails — so the published evidence cannot rot into a decorative file. scripts/record_demo.sh regenerates the recording, receipt, and key from a fresh real run; nothing inside the terminal is synthesised.

A run in flight is watchable from either operator surface. Both read the same task API, so neither is a lagging mirror of the other. In bernstein live, the left and right columns scroll independently as whole panes, so widgets below the fold remain reachable in shorter terminals.

A two-column terminal dashboard - agents with their live logs on the left, the task board on the right - with a full-width activity feed and a cost line underneathA browser dashboard listing sixty-two tasks with eleven running, one of them opened to its working-tree diff
bernstein live — the terminal dashboardbernstein gui serve — the browser dashboard

prove a run

Determinism here is something you check, not something you take on faith. Run once with audit enabled, then verify what was recorded:

BERNSTEIN_AUDIT=1 bernstein -g "fix the failing test in tests/test_foo.py"
bernstein replay list                 # run ids recorded on disk
bernstein replay latest --verify      # recompute the journal head, name the first divergent step
bernstein lineage verify <run_id>     # recompute the always-on lineage spine
bernstein audit verify                # HMAC chain + Merkle seal (written because audit was enabled)
bernstein audit diagnose <run_id> --signal gate --sign-key KEY
                                      # name the exact step a failure entered the run, as a signed receipt
bernstein verify run <run_id> --signing-key-path key.pem   # sign one portable run receipt
bernstein verify receipt .sdd/runs/<run_id>/run-receipt.json  # verify it offline: file only

The journal is written on every run; the lineage spine is always on and gains an entry for each lineage-bearing step, so a short run can finish with a valid, empty spine. bernstein audit verify only has a chain to check when the run was started with BERNSTEIN_AUDIT=1, a compliance preset, or bernstein run --audit. The --audit flag belongs to bernstein run; on the bernstein -g form above, set the environment variable.

One run receipt binds the journal head, the lineage-spine head when the run wrote spine entries, and, opt-in, an audit-chain range, under a single Ed25519-signed subject with the public key embedded. A reviewer holding that file and the operator's public key can confirm the embedded actions and chains were not changed: no HMAC key, no live .sdd/, and exit 2 naming the first divergent step on tamper. That receipt identifies the journal state it embeds; proving that state is the complete finished journal additionally requires an independent head/count seal. With the file alone and no --public-key pin, the check is integrity-only — it proves the receipt is internally consistent, not who signed it, and the verdict says so. Details in deterministic replay.

The same checkability applies to evaluation numbers. bernstein bench run <suite> --reliability k (also spelled bernstein eval --reliability k) runs every task k times under fixed coordination, then reports a pass^k floor (all k attempts must pass) alongside the pass@1 ceiling. That result is sealed in a signed receipt which bernstein bench reliability-verify recomputes offline, so a fabricated floor fails verification. Details: pass^k reliability floor.

how it works

Each goal moves through four stages:

  1. Decompose. The manager breaks your goal into tasks with roles, owned files, and completion signals. One LLM call, then plain Python from there.
  2. Spawn. Agents start in isolated git worktrees, one per coding task; an artifact-mode task gets a plain working directory instead. Main branch stays clean.
  3. Verify. The janitor checks concrete signals: tests pass, files exist, lint clean, types correct.
  4. Merge. Verified work lands in main. Failed tasks get retried or routed to a different model.

Why the scheduler is plain Python, and what that trades away: why deterministic.

everyday commands

cd your-project
bernstein init                    # creates .sdd/ workspace, bernstein.yaml + templates/
bernstein -g "Add rate limiting"  # agents spawn, work in parallel, verify, exit
bernstein live                    # watch progress in the TUI dashboard
bernstein run plan.yaml           # multi-stage plan: skip LLM planning, execute directly
bernstein stop                    # graceful shutdown with drain

The full operator surface (PR automation, schedules, chat bridges, the autofix daemon) is in operator commands.

Repository hygiene gates: bernstein readme-l10n verify fails a PR whose translated READMEs drifted from the English source (naming the stale section), bernstein readme-l10n sync rebinds them after an English edit. See readme-l10n.

supported agents

Claude Code, Codex CLI, Gemini CLI, GitHub Copilot CLI, Cursor, Aider, Goose, Muse Code, OpenAI Agents SDK, Amp, Cody, Continue, Devin Terminal, Junie, Kilo, Kiro, AWS Q Developer, Ollama, OpenCode, OpenHands, Open Interpreter, gptme, Plandex, AIChat, Letta Code, Qwen, and more. The adapter index carries install commands for 30 of them. bernstein integrations list enumerates all 51 wired-in integrations from src/bernstein/adapters/registry.py, the single source of truth for what resolves. 49 of them are selectable agent adapters; the other two rows are the mock test stub and the self-hosted-endpoints endpoint profile. Anything else with a --prompt flag works through the generic wrapper.

Mix agents in the same run: cheap local models for boilerplate, heavier cloud models for architecture. bernstein integrations list --installed shows what is available on your machine.

beyond the front page

Everything deep lives on the docs site:

pagewhat it covers
capabilitiesthe full capability list: MCP server mode, signed agent cards, sandbox backends, artifact sinks, regulatory mappings
who this is forwhere the value lands, and where Bernstein is the wrong tool
workflowsdeclarative YAML DAGs of agent / command / loop nodes
web UIbrowser dashboard on the same API the TUI uses
cloud executionexperimental: run agents on Cloudflare Workers with R2 workspace sync against your own account. The hosted api.bernstein.run service is not yet available
datasourcesread-only query receipts, plus a query driver that binds each result to the schema snapshot it was derived against
agent catalogspoint roles at agent definitions outside the built-in templates - a generic YAML/SKILL.md directory, or a Claude Code plugin-layout tree
securityscorecard, fuzzing, hardening
architecturehow it works under the hood

why the name?

Bernstein is named after Leonard Bernstein, the American conductor and composer. The project orchestrates a crew of CLI coding agents the way Bernstein conducted the New York Philharmonic: every player on cue, the score deterministic, the conductor accountable for the result.

i wrote bernstein because i was paying $400/month in claude bills running three coding agents in parallel and getting nondeterministic merges. Apache 2.0, solo maintained. Live stats: bernstein.run.

mentioned in

Listed in vinta/awesome-python, covered in Augment Code's open-source agent orchestrators roundup, and listed in Python Weekly #742. We also wrote up the approach as the deterministic zero-LLM orchestration pattern in awesome-agentic-patterns.

<details> <summary>All coverage: 20+ awesome lists, directories, newsletters, and peer citations</summary> <br>

The full tracked list, including every awesome-list entry, catalog listing, prior-art citation, and newsletter mention, lives in docs/mentions.md. Entries are added as they appear; corrections welcome by issue or PR.

</details>

contributing, support, license

PRs welcome; CONTRIBUTING.md has setup and code style. Security reports go through SECURITY.md. If Bernstein saves you time: GitHub Sponsors. Contact: forte@bernstein.run.

Citation metadata lives in CITATION.cff. License: Apache-2.0; the project name is covered separately in TRADEMARKS.md.


Alex Chernysh · GitHub · X · bernstein.run

<!-- mcp-name: io.github.sipyourdrink-ltd/bernstein -->

Related MCP servers

Verifies Bernstein run receipts and hash chains; lists the shipped presets and adapters. Read-only.

0
TypeScript
Apache-2.0
View repository →

Deploy full-stack web apps with database, file storage, auth, and RBAC via a single API call.

View repository →

Multi-chain EVM blockchain data with companion smart contracts

0
TypeScript
MIT
View repository →

AI-powered management for UniFi Access doors, credentials, policies, visitors, and events.

716
Python
MIT
View repository →

AI-powered management for UniFi Network, Protect, and Access controllers via MCP.

716
Python
MIT
View repository →

Manage UniFi Protect cameras, events, recordings, and smart detections via AI agents.

716
Python
MIT
View repository →