PluginBench
MCP Server
Maintained
MIT

Decompose MCP Server

io.github.echology-io/decompose

Deterministic text classification for AI agents—extract authority, risk, and attention scores without an LLM.

What is the Decompose MCP server?

The Decompose MCP server is a deterministic text classifier that breaks documents into structured semantic units, each labeled with authority level, risk category, attention score, and extracted entities. It requires no LLM, no setup, and no API keys—running pure regex and heuristics for fast, offline, reproducible classification.

Decompose preprocesses documents before they reach your LLM, classifying each sentence or clause by authority (mandatory, prohibitive, directive, etc.), risk (safety-critical, compliance, financial, etc.), and attention score (0–10). This lets your agent filter boilerplate, route high-risk content to specialized chains, and reduce token consumption by 60–80% while preserving critical information.

How to install Decompose

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "decompose": {
      "command": "uvx",
      "args": [
        "decompose-mcp"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • decompose_text — Decompose any text into classified semantic units with authority, risk, attention, type, irreducibility, and extracted entities.
  • decompose_url — Fetch a URL and decompose its content into classified semantic units.

Use cases

  • Filter contract or specification documents to extract only mandatory requirements and safety-critical constraints before passing to an LLM
  • Route financial or compliance-flagged content to specialized processing chains while skipping boilerplate
  • Reduce LLM token consumption by 60–80% by pre-filtering documents to keep only high-attention units
  • Extract and structure formal references (standards, codes, regulations) from technical documents for downstream processing
  • Measure document complexity and token cost reduction by counting high-attention units versus total units

Decompose MCP server FAQ

What does Decompose do?

Decompose classifies text into semantic units, labeling each with authority (mandatory, prohibitive, etc.), risk level (safety-critical, compliance, financial, etc.), attention score (0–10), and extracted entities. It works offline, deterministically, with no LLM.

Is Decompose free?

Yes. Decompose is MIT-licensed and free to use.

How do I install it in Cursor or Claude?

Install via pip (pip install decompose-mcp), then add to your MCP config: {"mcpServers": {"decompose": {"command": "uvx", "args": ["decompose-mcp", "--serve"]}}}

Does it require authentication or API keys?

No. Decompose runs entirely offline using regex and heuristics—no LLM, no API key, no GPU required.

What types of documents does it work on?

Decompose was battle-tested on engineering specs, contracts, RFIs, and inspection reports, but the methodology is domain-agnostic and works on any structured or semi-structured text.

How much faster is it than chunking + embedding + LLM retrieval?

Decompose typically reduces token consumption by 60–80% and processes 50-page specs in under 500ms, letting your agent focus compute on high-value units only.

README (reference)

Source of truth, from the repository.

Decompose

CI PyPI Python

<!-- mcp-name: io.github.echology-io/decompose -->

Stop prompting. Start decomposing.

Deterministic text classification for AI agents. Decompose turns any text into classified, structured semantic units — instantly. No LLM. No setup. One function call.


Before: your agent reads this

The contractor shall provide all materials per ASTM C150-20. Maximum load
shall not exceed 500 psf per ASCE 7-22. Notice to proceed within 14 calendar
days of contract execution. Retainage of 10% applies to all payments.
For general background, the project is located in Denver, CO...

After: your agent reads this

[
  {
    "text": "The contractor shall provide all materials per ASTM C150-20.",
    "authority": "mandatory",
    "risk": "compliance",
    "type": "requirement",
    "irreducible": true,
    "attention": 8.0,
    "entities": ["ASTM C150-20"]
  },
  {
    "text": "Maximum load shall not exceed 500 psf per ASCE 7-22.",
    "authority": "prohibitive",
    "risk": "safety_critical",
    "type": "constraint",
    "irreducible": true,
    "attention": 10.0,
    "entities": ["ASCE 7-22"]
  }
]

Every unit classified. Every standard extracted. Every risk scored. Your agent knows what matters.


Install

pip install decompose-mcp

Use as MCP Server

Add to your agent's MCP config (Claude Code, Cursor, Windsurf, etc.):

{
  "mcpServers": {
    "decompose": {
      "command": "uvx",
      "args": ["decompose-mcp", "--serve"]
    }
  }
}

Your agent gets two tools:

  • decompose_text — decompose any text
  • decompose_url — fetch a URL and decompose its content

OpenClaw

Install the skill from ClawHub or configure directly:

{
  "mcpServers": {
    "decompose": {
      "command": "python3",
      "args": ["-m", "decompose", "--serve"]
    }
  }
}

Or install the skill: clawdhub install decompose-mcp

Use as CLI

# Pipe text
cat spec.txt | decompose --pretty

# Inline
decompose --text "The contractor shall provide all materials per ASTM C150-20."

# Compact output (smaller JSON)
cat document.md | decompose --compact

Use as Library

from decompose import decompose_text, filter_for_llm

result = decompose_text("The contractor shall provide all materials per ASTM C150-20.")

for unit in result["units"]:
    print(f"[{unit['authority']}] [{unit['risk']}] {unit['text'][:60]}...")

# Pre-filter for LLM context — keep only high-value units
filtered = filter_for_llm(result, max_tokens=4000)
print(f"{filtered['meta']['reduction_pct']}% token reduction")
llm_input = filtered["text"]  # Ready for your LLM

What Each Field Means

FieldValuesWhat It Tells Your Agent
authoritymandatory, prohibitive, directive, permissive, conditional, informationalIs this a hard requirement or background?
risksafety_critical, security, compliance, financial, contractual, advisory, informationalHow much does this matter?
typerequirement, definition, reference, constraint, narrative, dataWhat kind of content is this?
irreducibletrue/falseMust this be preserved verbatim?
attention0.0 - 10.0How much compute should the agent spend here?
entitiesstandards, codes, regulationsWhat formal references are cited?
actionabletrue/falseDoes someone need to do something?

What to Build With This

Decompose is not the destination. It's the step before the LLM that most developers skip — not because it's hard, but because nobody showed them it exists. Documents have structure. That structure is classifiable. And classification should happen before reasoning.

Without:  document → chunk → embed → retrieve → LLM → answer  (100% of tokens)
With:     document → decompose → filter/route → LLM → answer  (20-40% of tokens)

Filter: built-in LLM pre-filter

filter_for_llm() keeps mandatory, safety-critical, financial, and compliance units — drops boilerplate before it reaches your LLM or vector store.

from decompose import decompose_text, filter_for_llm

result = decompose_text(open("contract.md").read())
filtered = filter_for_llm(result, max_tokens=4000)

# filtered["text"] = high-value units only, ready for LLM
# filtered["meta"]["reduction_pct"] = how much was dropped (typically 60-80%)

# Or use the units directly for embedding
for unit in filtered["units"]:
    embed_and_store(unit["text"], metadata={
        "authority": unit["authority"],
        "risk": unit["risk"],
        "attention": unit["attention"],
    })

Route: risk-based processing

Safety-critical content goes to one chain. Financial content goes to another. Boilerplate gets skipped.

from decompose import decompose_text

result = decompose_text(spec_text)

for unit in result["units"]:
    if unit["risk"] == "safety_critical":
        safety_chain.process(unit)       # Full analysis + human review
    elif unit["risk"] == "financial":
        audit_chain.process(unit)         # Flag for finance team
    elif unit["attention"] < 0.5:
        pass                              # Skip boilerplate
    else:
        general_chain.process(unit)       # Standard LLM analysis

Measure: token cost reduction

from decompose import decompose_text

result = decompose_text(spec_text)
total = len(result["units"])
high = [u for u in result["units"] if u["attention"] >= 1.0]

print(f"{len(high)}/{total} units need LLM analysis")
print(f"{100 - len(high) * 100 // total}% token reduction")

See examples/ for runnable scripts.


Why No LLM?

Decompose runs on pure regex and heuristics. No Ollama, no API key, no GPU, no inference cost.

This is intentional:

  • Fast: <500ms for a 50-page spec
  • Deterministic: Same input always produces same output
  • Offline: Works air-gapped, on a plane, on CI
  • Composable: Your agent's LLM reasons over the structured output — decompose handles the preprocessing

The LLM is what your agent uses. Decompose makes whatever model you're running work better.


Built by Echology

Decompose is built by Echology and extracted from AECai, a document intelligence platform for Architecture, Engineering, and Construction firms. The classification patterns, entity extraction, and irreducibility detection are battle-tested against thousands of real AEC documents — specs, contracts, RFIs, inspection reports, pay applications.

Decompose earned its independence — it started as AECai's text classification module, proved general enough to work across domains (insurance, trading, regulatory), and was released standalone. Free, MIT-licensed.

Case Study: Open Scripture Intelligence

The same chunking and entity extraction patterns that classify engineering specs also structure the Bible. Open Scripture Intelligence uses Decompose's Markdown-aware chunker and regex entity extraction to transform 31,100 verses into a knowledge graph with 344,799 cross-reference edges and semantic embeddings — proving the methodology is domain-agnostic.

Blog

License: MIT — Copyright (c) 2025-2026 Echology, Inc.

Related MCP servers

Real founder decisions, lessons and signals from 100+ podcasts — searchable by AI agents.

1
Python
MIT
View repository →

MCP server for the Horoshop e-commerce API — orders, catalog, customers, webhooks

1
TypeScript
MIT
View repository →

Hyperliquid perp funding data: top-20 rates, premia, OI, and noise-gated anomalies (15-min refresh).

View repository →

MCP server for Reolink cameras: snapshots, device state, AI detection, PTZ, and deterrence controls

2
Python
MIT
View repository →

Verifica RFCs contra las listas del SAT (69/69-B CFF, EFOS): veredicto de riesgo fiscal.

0
Python
View repository →
POPostgres MCP logo

PostgreSQL MCP wrapper with .env credential mapping, tool selection, and safe read-only defaults.

5
JavaScript
View repository →