io.github.shitianfang/jev-use MCP Server
io.github.shitianfang/jev-use
Batch yes/no, pick-one, and score judgments from Jev for faster LLM decisions with typed escalation.
What is the io.github.shitianfang/jev-use MCP server?
The jev-use MCP server integrates Jev, a fast decision model, with Claude Code, Codex, and pi to offload non-text tasks like yes/no checks, single-choice picks, and scoring. It batches judgments in one call and returns typed escalation when Jev cannot decide, letting the LLM focus on writing while Jev handles rapid decisions.
jev-use lets AI agents delegate fast, structured decisions to Jev instead of calling the main LLM, reducing latency and token use. It's designed for tasks like gating shell commands, context compaction, routing, and real-time decision loops where speed matters more than prose output.
How to install io.github.shitianfang/jev-use
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
judge— Batch typed yes/no (check), pick-one (pick), and score (rate) judgments in one call with confidence scores and escalation flags.gate— Pre-call routing to deny dangerous actions (e.g., shell commands) at ~230ms with a reason, zero LLM tokens.
Use cases
- Gate shell commands by safety risk before execution
- Compact conversation context by having Jev mark messages keep/drop, then LLM summarizes dropped blocks
- Route decisions in real-time loops (e.g., Pong ball control) where latency is critical
- Validate LLM outputs (e.g., reject wrong geocodes, rerun flaky tests) with fast confidence scores
- Triage incidents or support tickets by scoring urgency and routing to appropriate handlers
io.github.shitianfang/jev-use MCP server FAQ
An MCP server that lets Claude, Codex, and pi call Jev—a fast decision model—for structured yes/no, pick-one, and scoring tasks, with typed escalation when Jev is unsure.
The MCP server itself is MIT open-source. Jev judgments require an API key from TypeSafe, OpenRouter, or Vercel AI Gateway; a mock backend is available for local testing.
Run `npx -y jev-use install` to auto-wire Claude Code, Codex, or pi. Set one env var (TYPESAFE_API_KEY, OPENROUTER_API_KEY, or AI_GATEWAY_API_KEY) in your agent's environment.
One API key for your chosen Jev provider (TypeSafe direct, OpenRouter, or Vercel AI Gateway). Use JEV_BACKEND=mock for keyless dry runs.
Use Jev for fast, non-prose decisions: gating actions, routing, scoring, validation, and real-time loops. Use the LLM for writing, reasoning, and tasks needing text output.
Jev decisions average ~230–274ms (p50) in batched calls, 3–10× faster than calling the LLM for the same task, especially with enum constraints.
README (reference)
Source of truth, from the repository.
jev-use
English | 简体中文
The best way for Claude Code, Codex, and pi to work with Jev: hand the tasks that need no text output to Jev — faster steps, fewer tokens, tasks done sooner and better.
It makes the LLM and Jev true collaborators: when content needs to be written, the LLM takes over; when a step just needs a fast decision, Jev executes it.
Demos — real runs, 1× speed
<table> <tr> <td width="50%" valign="top"><b>Directions task: Jev clicks, the LLM types</b> — 10 decisions (p50 274 ms) · 4 writes; Jev rejects a wrong route, the LLM rewrites<br><img src="assets/collab.gif" alt="OpenStreetMap directions: Jev picks controls in green, the LLM types the locations in blue; a wrong 1809km geocode is rejected by Jev and repaired by the LLM, ending on the real 3.7km walking route" width="100%"></td> <td width="50%" valign="top"><b>Context compaction</b> — 200 messages judged in 7 calls, one LLM paragraph replaces the dropped pile; recall 3/3<br><img src="assets/compact.gif" alt="A real transcript fills the context window to 94%; Jev tints each message keep or drop, the LLM's summary paragraph replaces the dropped block, the window falls to 44% and three recall checks pass" width="100%"></td> </tr> <tr> <td width="50%" valign="top"><b>Pong: ball speed = decision latency</b> — 86 Jev decisions in 20 s vs 6 (haiku) and 3 (gemini) called the usual way; enum-constrain both and the gap is 3×<br><img src="assets/pong.gif" alt="Three Pong lanes replaying a live run at 1x: the Jev ball sweeps the field at ~224ms per decision while the LLM balls crawl" width="100%"></td> <td width="50%" valign="top"><b>Gate every shell command</b> — dangerous ones denied in ~230 ms with a reason, zero LLM tokens<br><img src="assets/gate.gif" alt="A 24-command dev session gated at 1x: dangerous commands denied at confidence 1.00, benign ones allowed" width="100%"></td> </tr> </table>Every demo is a rerunnable script in bench/examples/; all numbers, methodology, variance and caveats: bench/RESULTS.md · third-party measurements: docs/evidence.md.
Install
npx -y jev-use install # wires Claude Code, Codex, and pi — whichever it finds
Set one key in the environment your agent runs in (JEV_BACKEND=mock for
a keyless dry run):
| Provider | Env var |
|---|---|
| TypeSafe direct | TYPESAFE_API_KEY |
| OpenRouter | OPENROUTER_API_KEY |
| Vercel AI Gateway | AI_GATEWAY_API_KEY |
npx -y jev-use doctor checks the wiring. Judged state goes to the
provider you configure; JEV_BACKEND=mock stays local. Plugin form with
the routing skill and the PreToolUse gate:
harness/claude-code ·
harness/codex.
Use as a library
npm i jev-use — zero runtime dependencies on the judgment path:
import { Jev, check, pick, rate } from "jev-use";
const jev = new Jev();
const { answers } = await jev.judge(state, {
next: pick("Next action?", { merge: "all green", rerun: "looks flaky", hold: "needs attention" }),
risk: rate("How risky?", ["routine", "worth a look", "incident"]),
passed: check("Did the run fully succeed?"),
});
// answers.next → { answer: "merge", confidence: 0.93, confidenceFrom: "reported", escalate: false }
Anything Jev can't or shouldn't decide comes back with escalate: true
and a typed reason. Tools, verdict shape, escalation contract, CLI:
docs/reference.md.
Small enough to read
| File | Job |
|---|---|
| src/protocol.ts | Questions (check/pick/rate), verdicts, escalation reasons |
| src/dispatch.ts | Pre-call routing: what never reaches Jev |
| src/judge.ts | screen → backend → hand back what is unsure; gate |
| src/jev.ts | The Jev client over that engine |
| src/redact.ts | Credentials stripped from a gated action before it is sent |
| src/backends/ | TypeSafe, OpenRouter, Vercel, mock adapters |
| src/server.ts | The two MCP tools |
| src/cli.ts | install, serve, hook gate, doctor |
| skills/jev-use/SKILL.md | The routing rules the agent follows |
Development
$ npm run typecheck && npm test # unit tests incl. per-provider wire fixtures
$ npm run smoke # real MCP client ↔ built CLI over stdio
$ node bench/run.mjs # micro-benchmarks, your key and region
Substantially written with Claude Code (AI-assisted).
MIT © shitianfang
Related MCP servers
Manage AWS, Azure, GCP, Kubernetes, CI/CD, Docker, Terraform and more. 727 tools.
View repository →MCP server that cuts AI coding agent token usage via framework-aware context optimization

Seminara
Digital Teammate (Aura) hosting live interactive presentations, demos, and audience Q&A 24/7.

SMRITI Memory
Neuro-inspired long-term memory for AI agents with semantic graph and consolidation.
Privacy scores, merchant insights analytics.


