PluginBench
MCP Server
Active
MIT

io.github.ataberk-xyz/gossipcat MCP Server

io.github.ataberk-xyz/gossipcat

Multi-agent consensus code review that catches hallucinations by having AI reviewers verify each other's findings against your actual code.

What is the io.github.ataberk-xyz/gossipcat MCP server?

Gossipcat is an MCP server that runs multiple AI agents in parallel to review code, with each agent verifying its peers' findings against real file:line citations in your codebase. It filters out hallucinations through mechanical cross-review, assigns accuracy scores to agents, and auto-generates skill files from failure patterns to improve over time. It integrates with Claude Code and Cursor as a live orchestrator with a dashboard and two-way browser chat bridge.

Gossipcat solves the problem of confident AI hallucinations in code review by running several agents simultaneously and requiring each finding to survive peer verification against your actual source code. Findings are tagged as CONFIRMED (multiple agents verified), UNIQUE (one agent verified), DISPUTED (agents disagreed, re-checked), or UNVERIFIED. The system learns over time by routing work to agents with proven accuracy and auto-generating skill files to fix repeat failures.

How to install io.github.ataberk-xyz/gossipcat

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • GOOGLE_API_KEY
    secret

    Google Gemini API key for relay agents (optional — native agents need no API key)

  • GOSSIPCAT_PORT

    Fixed port for the relay/dashboard server (optional — defaults to OS-assigned with sticky file)

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "gossipcat": {
      "command": "npx",
      "args": [
        "-y",
        "gossipcat"
      ],
      "env": {
        "GOOGLE_API_KEY": "<YOUR_GOOGLE_API_KEY>",
        "GOSSIPCAT_PORT": "<YOUR_GOSSIPCAT_PORT>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • consensus_review — Run multiple agents in parallel to review code changes, with each agent cross-verifying peers' findings against actual file:line citations
  • agent_dispatch — Adaptive routing of review tasks to agents based on their measured competency scores in different categories
  • skill_generation — Auto-generate skill files from agent failure patterns and inject them into future prompts to improve accuracy
  • cross_verification — Mechanical verification of code review findings against real source code locations to confirm or reject claims
  • accuracy_scoring — Track per-agent competency scores updated by reward signals (confirmed findings) and penalty signals (hallucinations caught by peers)

Use cases

  • Run a consensus code review of recent changes with multiple agents verifying each other's findings
  • Set up a gossipcat team for a project and receive only high-confidence findings that survived peer cross-review
  • Identify which agents are most reliable for specific types of issues and let the system route work accordingly
  • Auto-generate skill files from repeated failure patterns to teach agents how to avoid common mistakes
  • Catch hallucinated bugs before they ship by requiring mechanical verification against actual code citations

io.github.ataberk-xyz/gossipcat MCP server FAQ

What is Gossipcat and how does it differ from single-agent code review?

Gossipcat runs multiple AI agents in parallel and has each one verify its peers' findings against your actual code. Solo reviewers hallucinate confidently with no second opinion; Gossipcat's consensus model filters those out mechanically by requiring findings to cite real file:line locations that peers verify. Only findings that survive cross-review are surfaced.

Is Gossipcat free to use?

The smallest working team (Sonnet reviewer + Haiku researcher) runs fully native on your existing Claude Code or Cursor subscription with zero API keys. Relay agents (Gemini, OpenAI, Grok, DeepSeek, Ollama) are optional and mix freely, but the core system is free.

How do I install Gossipcat in Claude Code or Cursor?

In Claude Code, run `npx skills add gossipcat-ai/gossipcat-ai` for the fastest setup, or manually install with `npm install -g gossipcat` then `claude mcp add gossipcat -s user -- gossipcat`. In Cursor, add `{ "gossipcat": { "command": "gossipcat" } }` to `.cursor/mcp.json`.

Does Gossipcat require API keys or authentication?

No API keys are required for the native team (Sonnet + Haiku agents running on your subscription). Relay agents are optional and only needed if you want to add external providers like OpenAI or Gemini.

How does Gossipcat improve over time?

The system tracks per-agent accuracy scores based on findings that survive peer verification (rewards) and hallucinations caught by peers (penalties). Agents with proven accuracy get routed more work. When an agent fails repeatedly in a category, Gossipcat auto-generates a skill file from that failure history and injects it into future prompts.

What do the finding tags (CONFIRMED, UNIQUE, DISPUTED, UNVERIFIED) mean?

CONFIRMED: multiple agents found and verified it — fix it. UNIQUE: one agent found and verified it — high signal. DISPUTED: agents disagreed, gossipcat re-checked the code — trust the verdict. UNVERIFIED: looks real but wasn't cross-checked yet — verify manually.

README (reference)

Source of truth, from the repository.

<p align="center"> <img src="https://raw.githubusercontent.com/gossipcat-ai/gossipcat-ai/master/packages/dashboard-v2/public/assets/banner.png" alt="Gossipcat" width="520" /> </p> <p align="center"> <strong>Multi-agent consensus code review.</strong><br/> AI reviewers lie confidently. Gossipcat makes them check each other — against your actual code. </p> <p align="center"> <code>TypeScript</code> · <code>MCP</code> · <code>Claude Code</code> · <code>Cursor</code> · <code>multi-agent</code> </p> <p align="center"> <a href="https://www.npmjs.com/package/gossipcat"><img src="https://img.shields.io/npm/v/gossipcat?color=0ea5e9" alt="npm" /></a> <a href="https://github.com/gossipcat-ai/gossipcat-ai/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/gossipcat-ai/gossipcat-ai/ci.yml?branch=master&label=tests" alt="tests" /></a> <a href="https://github.com/gossipcat-ai/gossipcat-ai/blob/master/LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue" alt="MIT" /></a> <a href="https://github.com/gossipcat-ai/gossipcat-ai/stargazers"><img src="https://img.shields.io/github/stars/gossipcat-ai/gossipcat-ai?style=social" alt="stars" /></a> </p> <p align="center"> <a href="#quick-start">Quick start</a> · <a href="#how-it-works">How it works</a> · <a href="docs/GUIDE.md">Guide</a> · <a href="docs/HANDBOOK.md">Handbook</a> · <a href="CHANGELOG.md">Changelog</a> </p>

A single AI reviewer will, with total confidence, report bugs that aren't there. You read the finding, you go look, you waste twenty minutes — the code was fine. No second opinion, no track record, no way to tell a real catch from a hallucination until you've paid for it.

Gossipcat runs several agents in parallel, has each one verify its peers' findings against your real file:line, and only surfaces what survives. When an agent invents a finding, a peer catches it and the agent's accuracy score drops — over time the system routes each kind of work to whoever is measurably reliable at it. The verdict comes from citation checks against your source, never from one model grading another.

It runs as an MCP server inside Claude Code and Cursor, with a live operator dashboard and a two-way browser chat bridge into the running orchestrator.

<p align="center"> <img src="https://raw.githubusercontent.com/gossipcat-ai/gossipcat-ai/master/packages/dashboard-v2/public/assets/dashboard-overview.png" alt="Gossipcat dashboard — live fleet view with per-agent accuracy rings, signal volume, and recent hallucination catches" width="880" /> </p>

Reading a report

Your whole job is four tags:

TagMeansWhat you do
CONFIRMEDMultiple agents found it and verified it against the codeFix it
UNIQUEOne agent found it, cross-checked and held upFix it — high signal
DISPUTEDAgents disagreed; gossipcat re-checked the codeTrust the verdict
UNVERIFIEDLooks real but wasn't cross-checked yetGlance, then verify

The DISPUTED false alarm that cross-review kills is the bug a solo reviewer would have shipped to you. That delta is the whole point.


How it works

flowchart LR
    A([agent review]) -->|cites file:line| B([peer cross-review])
    B -->|verifies against code| C{verdict}
    C -->|confirmed| D[reward signal]
    C -->|hallucination| E[penalty signal]
    D --> F[competency score]
    E --> F
    F -->|steer dispatch| G([next agent pick])
    E -->|≥3 in category| H[auto-generate skill]
    H -->|inject into prompt| A
    G --> A
    style A fill:#0ea5e9,stroke:#0369a1,color:#fff
    style H fill:#f59e0b,stroke:#b45309,color:#fff
    style D fill:#10b981,stroke:#047857,color:#fff
    style E fill:#ef4444,stroke:#b91c1c,color:#fff

Every finding must cite a real file:line. Peers verify the citation mechanically — agree, disagree, or new — and the verified outcomes become reward signals that update per-agent competency scores. An agent that keeps failing in one category gets a skill file auto-generated from its own failure history and injected into future prompts; skills that don't measurably help are statistically demoted. It's in-context reinforcement learning at the prompt layer: the reward is grounded in your source code, the "policy update" is a markdown file, and no weights are ever touched.

Since v0.8, skills also activate by task relevance instead of shipping wholesale, and agents can pull skills on demand mid-task — including your own Claude Code project skills from .claude/skills/, no duplication needed.


Quick start

Node 22+, and either Claude Code or Cursor.

npx skills add gossipcat-ai/gossipcat-ai   # fastest — installer skill walks you through it

or manually:

npm install -g gossipcat
claude mcp add gossipcat -s user -- gossipcat     # Claude Code
# Cursor: add { "gossipcat": { "command": "gossipcat" } } to .cursor/mcp.json

Then, in any project:

"Set up a gossipcat team for this project." "Do a consensus review of my recent changes."

The smallest working team — sonnet-reviewer + haiku-researcher — is fully native and needs zero API keys: it runs on your existing Claude Code / Cursor subscription. Relay agents (Gemini, OpenAI, Grok, DeepSeek, Ollama, any OpenAI-compatible endpoint) are optional and mix freely.

First run, daily recipes, dashboard, configuration, and troubleshooting: docs/GUIDE.md.


Compared to the alternatives

Filters hallucinationsImproves over time
Gossipcat — 3+ agents cross-review; confirmed bugs onlyYes — peers catch and penalize hallucinations mechanicallyYes — accuracy steers dispatch; skill files fix repeat failures
Single-agent review (IDE built-in)No — hallucinations ship as findingsNo feedback loop
Model-grades-model reviewPartial — the judge hallucinates tooScores aren't wired to dispatch
Lint-style PR botsNoNo

The difference is ground truth: findings are verified against actual file:line citations in your codebase, which is what makes the reward signal trustworthy enough to automate.


Architecture

gossipcat/
  apps/cli/               MCP server, host-aware native agent bridge, boot sequence
  packages/
    orchestrator/         Dispatch pipeline, consensus engine, memory, skills, scoring
    relay/                WebSocket relay server, dashboard REST/WS API
    dashboard-v2/         React + Vite + shadcn/ui frontend (see DESIGN.md)
    client/               WebSocket client for relay connections
    tools/                File / shell / git tools for worker agents
    types/                Shared types and message protocol

Native agents run as host subagents (Claude Code Agent() / Cursor Task()) on your subscription — no API key. Relay agents run as WebSocket workers against any provider. Both participate equally in consensus, memory, and skill development.

Reading this as a Claude Code or Cursor instance? Call gossip_status() — it boots your full operating rules. The internals and design invariants live in docs/HANDBOOK.md.


Docs

docs/GUIDE.mdOperator guide — first run, daily recipes, dashboard, config, tools, troubleshooting
docs/HANDBOOK.mdInternals — architectural invariants, the signal pipeline, why the design is shaped this way
CHANGELOG.mdReleases, with per-version upgrade steps
CLAUDE.mdThe operating rules gossipcat's own agents follow while developing gossipcat

Roadmap

Dashboard enrichment (graphs, trends, session history) · local Postgres migration · Windsurf / VS Code native parity · standalone CLI. Shipped work: releases.

Contributing

Bug reports, ideas, and PRs welcome — open an issue or ask in-session "file a gossipcat bug report about …". Fork, branch, npm test, conventional commits; details in CONTRIBUTING.md.

License

MIT

Related MCP servers

BIBilinc logo

Bilinc

Active

Hosted agent memory with provenance, contradiction surfacing, and snapshot rollback.

1
Python
View repository →
CLClawifi logo

Clawifi

Active

Agent-native internet gateway: typed search, fetch, scrape, crawl, and extract via the Clawifi API.

0
Python
MIT
View repository →

8-tool AI web intelligence suite: search, scrape, screenshot, SEO, docs, crypto, code.

View repository →

An MCP server that provides [describe what your server does]

1
Python
AGPL-3.0
View repository →

A MCP server provides SQL database access with support for PostgreSQL, MySQL, and SQLite.

1
Python
AGPL-3.0
View repository →

Deterministic dependency graph + 24 MCP tools for AI-safe code changes, impact analysis, and security scanning.

52
TypeScript
View repository →