PluginBench
MCP Server
Active
MIT-0

io.github.nickjlamb/redacta-mcp MCP Server

io.github.nickjlamb/redacta-mcp

Pseudonymise patient identifiers and PII in clinical text locally with HIPAA Safe Harbor support.

What is the io.github.nickjlamb/redacta-mcp MCP server?

Redacta is an MCP server that pseudonymises medical and clinical documents by replacing patient identifiers with labelled tokens like [PATIENT_NAME_1] and [NHS_NUMBER_1], while preserving clinical meaning. It uses deterministic pattern matching for structured identifiers (NHS numbers, dates, postcodes, emails) plus optional LLM reasoning for free-text names and addresses, and can restore original identifiers locally via a token map. The server supports HIPAA Safe Harbor mode for stricter US de-identification.

Redacta helps you safely process clinical documents with AI by redacting sensitive patient information before they leave your environment. It detects and replaces identifiers with tokens, returns a redaction report, and can reverse the process to restore original values—keeping identifiable text on your machine throughout. Use it to prepare patient data for AI analysis while maintaining privacy and regulatory compliance.

How to install io.github.nickjlamb/redacta-mcp

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "redacta-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "redacta-mcp"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • redact — Pseudonymise patient identifiers and PII in clinical text, replacing them with labelled tokens and returning a redaction report.
  • reinstate — Restore original patient identifiers from a token map generated by a prior redaction, reversing the pseudonymisation locally.
  • HIPAA Safe Harbor mode — Apply stricter de-identification rules covering all dates, specific ages, and additional HIPAA identifiers beyond standard redaction.

Use cases

  • Redact patient documents before sending them to Claude or other AI tools for analysis while keeping original identifiers secure locally.
  • Prepare clinical notes for AI-assisted diagnosis or treatment planning without exposing sensitive patient information.
  • Reverse redaction after AI processing to restore original patient details using the token map, completing a secure round-trip workflow.
  • Comply with HIPAA de-identification requirements by applying Safe Harbor rules to remove all regulated identifiers.
  • Integrate privacy-preserving redaction into clinical AI workflows running on Kubernetes or self-hosted infrastructure.

io.github.nickjlamb/redacta-mcp MCP server FAQ

What does Redacta do?

Redacta replaces patient identifiers (names, NHS numbers, dates of birth, postcodes, emails, etc.) with labelled tokens like [PATIENT_NAME_1], preserving clinical meaning while removing PII. It also reverses the process to restore original identifiers locally using a token map.

Is Redacta free?

Yes, Redacta is open-source under the MIT-0 license. The MCP server is available on npm as redacta-mcp. PharmaTools.AI also offers fixed-price design-partner integrations for teams deploying it in production on clinical data.

How do I install it in Claude or Cursor?

Install via npm: `npx -y redacta-mcp`. For Claude Desktop or Cursor, configure it in your MCP settings to connect to the server. For Claude Code, clone the repository as a skill.

Does it require authentication or network access?

No. The deterministic pattern layer runs entirely locally using Python standard library only, with no network calls. Identifiers never leave your machine unless you explicitly send redacted text to an AI tool.

What identifiers does it detect?

Redacta detects NHS numbers (Modulus-11 validated), UK National Insurance numbers, dates of birth, UK postcodes, phone numbers, emails, hospital/MRN numbers, US SSN, and US ZIP codes. It also uses reasoning to identify patient names, relatives, postal addresses, and identifying ages.

Can I run it on my own infrastructure?

Yes. Redacta includes a self-hosted HTTP service deployable to Kubernetes with plain YAML, keeping text pseudonymised inside your environment before reaching external AI services.

README (reference)

Source of truth, from the repository.

Redacta

CI DOI engine npm downloads redacta-mcp PyPI App Store Anthropic MCP Directory Glama MCP server Listed on mcpservers.org self-hosted

Pseudonymise medical and clinical documents before they're processed by AI or shared. Redacta replaces patient identifiers with labelled tokens — [PATIENT_NAME_1], [NHS_NUMBER_1], [DATE_OF_BIRTH_1], … — while leaving the clinical meaning intact, and returns a redaction report alongside the cleaned text.

It started as an Agent Skill and is now one engine shipped across eight surfaces — an iOS app, agent skill, MCP server, a self-hosted HTTP service with a Kubernetes deployment, two libraries, a CLI, and a FigJam whiteboard plugin.

Running this in production? Redacta offers a small number of fixed-price design-partner integrations for teams shipping AI agents on clinical or patient data — deployment in your environment, one real workflow integrated, and a data-flow document written for your DPO. Details →

One engine, many surfaces

<p align="center"> <img src="ios-app/docs/architecture.svg" width="100%" alt="One detection engine feeds eight surfaces: the iOS app, Share Extension and widget run it on-device via JavaScriptCore; the MCP server, CLI, TypeScript library and FigJam plugin consume it directly; a Python package mirrors it; and the agent skill adds LLM reasoning." /> </p>
SurfaceFolderGet it
iOS app — iPhone (app, Share Extension, widget)ios-app/Build with Xcode — see ios-app/README.md
Agent skill (Claude Code / apps / API)SKILL.md, scripts/openclaw skills install redacta (ClawHub)
MCP server (Claude Desktop, Cursor, …)mcp-server/npx -y redacta-mcp (npm · MCP Registry · Anthropic MCP Directory)
TypeScript librarynpm-package/npm i @pharmatools/redacta (npm)
Python librarypython-package/pip install redacta (PyPI)
Command-line toolcli-package/npx redacta-cli (npm)
Self-hosted HTTP service + Kubernetesgateway-service/docker build — see gateway-service/README.md
FigJam pluginfigjam-plugin/Figma Community

The detection logic lives in one place — the TypeScript engine (@pharmatools/redacta, in npm-package/), which the MCP server and the FigJam plugin consume, and which the iOS app runs on-device via JavaScriptCore. The Python package mirrors it for pip users; the agent skill adds LLM reasoning for free-text names on top of the deterministic patterns.

How it works

<picture> <source media="(prefers-color-scheme: dark)" srcset="docs/boundary-dark.svg"> <img src="docs/boundary-light.svg" alt="The Redacta privacy boundary: a clinical document is redacted inside your boundary — deterministic patterns plus reasoning plus a self-check — producing tokenised text and a token map. Only the tokenised text crosses to the AI tool; the token map never leaves. The processed output comes back and reinstate restores the original identifiers locally. Raw identifiers never cross the boundary." width="100%"> </picture>

Two layers:

  • Patterns (deterministic). A bundled script (scripts/redact_structured.py, Python standard library only, no network) matches fixed-format identifiers: NHS numbers (Modulus-11 validated), UK National Insurance numbers, dates of birth, UK postcodes, phone numbers, emails, and hospital/MRN numbers. US SSN and ZIP codes are also handled.
  • Reasoning (judgement). The skill then has the agent handle what patterns can't: patient names (told apart from the clinicians treating them), relatives and carers, postal addresses, and identifying ages.
  • Self-check. A final pass re-reads the output for any identifier that slipped through before the report is written.

It also works in reverse. Re-identification (scripts/reinstate.py) takes the token map from an earlier redaction and restores the original values — so you can redact a document, run it through another AI tool, and put the real details back locally. Redact → process → re-identify is a complete round trip, and identifiers only ever exist on your machine.

Safe Harbor mode. Ask for HIPAA Safe Harbor (or "US de-identification") and Redacta applies a stricter pass: all dates (not just the date of birth), all specific ages, and the remaining HIPAA identifiers — fax, certificate/licence, device serial, VIN, and health-plan beneficiary numbers.

Self-hosting on Kubernetes

Organisations that can't let identifiable text leave their environment can run Redacta inside their own infrastructure: a small HTTP service (gateway-service/) deployable into an existing Kubernetes cluster with plain YAML — two stateless replicas behind a Service for redact/reinstate, an optional single-replica session boundary for the protect → release loop, health probes, resource limits, restrictive security defaults, and no-PHI logging. Text is pseudonymised before it reaches any external AI service, and the processing boundary stays under your control. Walkthrough (local kind cluster included): gateway-service/k8s/README.md · concepts: docs/KUBERNETES.md. Deploying somewhere a DPO will ask questions? There's a one-page security & data-protection summary at pharmatools.ai/redacta-security.

Install

Claude Code

git clone https://github.com/nickjlamb/redacta ~/.claude/skills/redacta

Then invoke it with /redacta, or let it trigger automatically when you ask to redact or de-identify clinical text.

Claude apps / API

Zip the repository folder and upload it as a skill.

Contents

PathWhat it is
SKILL.mdThe skill — instructions plus metadata
reference.mdPattern specs, the Modulus-11 algorithm, NI prefix rules, the date-of-birth vs clinical-date rule, token vocabulary, limitations
scripts/redact_structured.pyThe deterministic pattern layer
scripts/reinstate.pyThe re-identification layer (restore originals from a token map)
scripts/test_redact_structured.pyTests for the pattern layer
scripts/test_reinstate.pyTests for the re-identification layer
evaluations.jsonExample evaluation scenarios

Run the tests:

python3 scripts/test_redact_structured.py
python3 scripts/test_reinstate.py

A note on limits

Redacta is a strong first line of defence, not a guarantee. It won't catch every possible identifier and isn't a substitute for formal data-protection processes. Always review the redaction report before sharing text.

License

MIT-0 (MIT No Attribution). Built by PharmaTools.AI.

Related MCP servers

Explain why two scientific papers disagree, every claim grounded in a verbatim source quote.

1
JavaScript
MIT
View repository →

Deployment-readiness checks for document-QA and extraction AI: grounding, extraction, CI gate.

1
TypeScript
View repository →

Check whether an AI answer is grounded in its context — deterministic, no LLM judge.

1
JavaScript
MIT
View repository →

PubMed, Europe PMC, FDA/UK drug labels, and clinical trials access for AI assistants.

13
TypeScript
MIT
View repository →

MCP server for GreyNoise API - Check if IPs are internet background noise or targeted attacks

0
JavaScript
MIT
View repository →
QUQueryPilot logo

QueryPilot

Maintained

Safe SQL gateway for AI agents: SELECT-only validation, access policies, masking, audit trail, evals

1
Python
MIT
View repository →