PluginBench
MCP Server
Active
MIT

anymd MCP Server

io.github.SylphxAI/anymd

Convert any file to clean Markdown for AI agents: PDF, Office, EPUB, HTML, images, audio/video—fast, local, no API key.

What is the anymd MCP server?

The anymd MCP server converts PDFs, Word documents, PowerPoint, Excel, EPUB, HTML, images, and audio/video files into clean Markdown optimized for AI agents. It runs locally in Rust with no API keys required, handles OCR and transcription when tools are installed, and processes documents in parallel for speed.

anymd is a fast, accurate document-to-Markdown converter built for AI agents. It reads 12+ file formats—PDFs, Office documents, web pages, images, audio, and video—and outputs clean Markdown with page citations, compact tables, and token budgets to fit large documents in context. Everything runs locally on your machine with no uploads or API keys.

How to install anymd

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "anymd": {
      "command": "npx",
      "args": [
        "-y",
        "@sylphx/anymd"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • read — Convert a file, URL, or folder into Markdown. Supports pages selection, token limits, OCR, image extraction, and transcript generation.
  • search — Find text across files, folders, and URLs with literal or ranked (BM25) search modes and glob filtering.
  • inspect — Inspect PDF details: render pages, extract regions, OCR specific pages, extract structure as JSON with geometry, or compare documents.

Use cases

  • Extract text and tables from PDFs with correct reading order and formatting for analysis or summarization
  • Search across a folder of documents to find specific information or answer questions from multiple sources
  • Convert Word, PowerPoint, and Excel files to Markdown for use in AI workflows without losing formatting
  • Transcribe audio or video files and extract metadata like duration, chapters, and subtitles
  • OCR image-only PDFs or scanned documents to make them searchable and readable by AI agents

anymd MCP server FAQ

What file formats does anymd support?

PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls, .ods), CSV/TSV, EPUB, HTML, URLs, Markdown, text, JSON, images (with OCR), and audio/video (with transcripts and metadata).

Is anymd free?

Yes, anymd is open-source under the MIT license. It runs locally on your machine with no API keys or cloud services required.

How do I install anymd in Cursor or Claude?

Run `npx -y @sylphx/anymd setup` to add it to all MCP clients at once, or install manually by adding the command `npx -y @sylphx/anymd` to your client's MCP configuration.

Do I need to install OCR or transcription tools?

No, they are optional. anymd uses tesseract for OCR, ffmpeg/ffprobe for audio/video metadata, and whisper.cpp for transcripts—only if installed. It works without them.

Does anymd upload my documents?

No, everything runs locally on your machine. Documents never leave your device unless you pass a URL, and even then only that URL is fetched.

How fast is anymd compared to other tools?

On the AgentDocBench benchmark, anymd converts 38 documents in 22.9 seconds—143× faster than docling, 4× faster than markitdown, and 435× faster than marker.

README (reference)

Source of truth, from the repository.

<div align="center"> <img src="docs/public/og-image.png" alt="anymd — any file → clean Markdown for AI agents" width="820" /> <h1 hidden>anymd</h1> <!-- generated:lead -->

PDF, Word, PowerPoint, Excel, EPUB, HTML and web pages, images (OCR), audio and video (metadata, subtitles, transcripts). A fast Rust MCP server and CLI that runs on your machine. No API key.

<!-- /generated:lead -->

npm downloads stars MCP registry license OpenSSF Scorecard

<!-- repomap:agent-ready -->[![agent-ready 93/100](https://mark.sylphx.com/badge/agent--ready-93%2F100-brightgreen?style=flat-square&labelColor=0a0d07)](https://github.com/SylphxAI/repomap#agent-readiness-score)<!-- /repomap:agent-ready -->

Install · Benchmarks · Tools · CLI · Formats · Docs

<!-- generated:formerly -->

<sub>Formerly pdf-reader-mcp. Migrating from pdf-reader-mcp</sub>

<!-- /generated:formerly --> <img src="docs/public/demo.gif" alt="Real terminal session: anymd converts a PDF page with its table, searches a folder, reads a spreadsheet, then Claude Code answers from the PDF through the anymd MCP server" width="820" />

<sub>A real, unedited terminal recording (asciinema + agg, <a href="bench/demo">script</a>). The last command is Claude Code answering from the PDF through the anymd MCP server.</sub>

</div>

Why anymd

<!-- fast:start -->
  • Fast. Native Rust converts in parallel, page by page. On the 19 benchmark documents every tool converted, anymd takes 12.1 s in total; docling 1,723.3 s (143×), markitdown 45.4 s (4×), marker 5,256.3 s (435×).
<!-- fast:end -->
  • Accurate. A layout engine rebuilds words from glyph gaps, puts two-column papers in reading order, and recovers tables, including borderless ones. The text stays exactly as printed, with no glued words and no scrambled columns.
  • Lean on tokens. Pages come back as Markdown with <!-- page 3 --> citation anchors, a small front-matter header, and compact tables. A token budget and a cursor keep large documents within your agent's context.
  • Every format, one call. One tool reads every format listed below. It also accepts web URLs and whole directories, and search looks across all of them.
  • Local and private. Nothing is uploaded. OCR and transcripts use local tools you already have (tesseract, ffmpeg, whisper.cpp), and only when they are installed.

Install

Add anymd to every MCP client on your machine (Claude Code, Codex, Cursor, VS Code, Claude Desktop, Windsurf, Gemini CLI) with one command:

npx -y @sylphx/anymd setup     # --dry-run to preview, --remove to undo

Or add it by hand: every MCP client runs the same command, npx -y @sylphx/anymd. Node 18+ is the only requirement; npm installs the native binary for your platform.

<details open> <summary><b>Claude Code</b></summary>
claude mcp add anymd -- npx -y @sylphx/anymd

Or as a plugin, with the anymd skill: /plugin marketplace add SylphxAI/anymd, then /plugin install anymd@anymd.

</details> <details> <summary><b>Codex</b></summary>
codex mcp add anymd -- npx -y @sylphx/anymd

or in ~/.codex/config.toml:

[mcp_servers.anymd]
command = "npx"
args = ["-y", "@sylphx/anymd"]
</details> <details> <summary><b>Cursor</b></summary>

Add to Cursor

or in .cursor/mcp.json:

{ "mcpServers": { "anymd": { "command": "npx", "args": ["-y", "@sylphx/anymd"] } } }
</details> <details> <summary><b>VS Code</b></summary>

Install in VS Code with one click, or from a terminal:

code --add-mcp '{"name":"anymd","command":"npx","args":["-y","@sylphx/anymd"]}'

or in .vscode/mcp.json:

{ "servers": { "anymd": { "type": "stdio", "command": "npx", "args": ["-y", "@sylphx/anymd"] } } }
</details> <details> <summary><b>Claude Desktop</b></summary>

One click: download anymd-<version>.mcpb from the latest release and open it. Or, by hand:

Add to claude_desktop_config.json (Settings → Developer → Edit Config):

{ "mcpServers": { "anymd": { "command": "npx", "args": ["-y", "@sylphx/anymd"] } } }
</details> <details> <summary><b>Windsurf, Zed, Cline, and other clients</b></summary>

Any client that speaks MCP over stdio: command npx, args ["-y", "@sylphx/anymd"]. To keep the server inside one folder, add --allow-dir=/path/to/docs.

</details> <details> <summary><b>CLI only</b></summary>
npm install -g @sylphx/anymd     # or run it once with: npx -y @sylphx/anymd <file>

Python: uvx anymd report.pdf > report.md runs it once, pip install anymd installs it, and uvx anymd mcp starts the MCP server. The wheels carry the same prebuilt binary.

Docker (amd64 and arm64):

docker run --rm -v "$PWD:/data" ghcr.io/sylphxai/anymd report.pdf > report.md
docker run -i --rm ghcr.io/sylphxai/anymd        # MCP server on stdio

Or build it from crates.io (needs a Rust 1.92+ toolchain; OCR and transcripts still use tesseract/ffmpeg when installed):

cargo install anymd

npm, pip and Docker ship a prebuilt binary, while cargo install compiles one on your machine.

</details>

Benchmarks

AgentDocBench is an open benchmark for document → Markdown conversion for agents: license-clean documents in 12 categories (math papers, two-column papers, financial tables, forms, scans, CJK, slides, spreadsheets, Word, EPUB, HTML), scored on verbatim sentences, text F1, reading order, and table cells, with time and output tokens. Every tool runs on the same kind of GitHub-hosted runner (4 CPUs):

<!-- headline:start -->
anymddoclingkreuzbergunstructuredmarkitdownmarkerpdftotext
Overall score96.393.081.781.276.871.042.2
Table cells F192.289.938.438.457.260.90.0
Reading order98.894.496.893.985.576.852.0
Docs converted38/3838/3838/3838/3838/3830/3823/38
Time, all docs22.9 s2,432.4 s16.0 s346.5 s75.6 s7,104.5 s0.90 s
<!-- headline:end -->

The generated leaderboard, per-category scores (including where anymd loses), and method are in the benchmark guide. The corpus, ground truth, adapters, and raw results are in bench/, and the Benchmark workflow reruns everything; new tools can join with a single adapter file.

MCP tools

anymd exposes three tools.

ToolUse it toKey arguments
readTurn a file, URL, or folder into Markdownsource, pages ("1-5,8"), max_tokens (default 20000), cursor, ocr, images (refs · none), transcript, download_whisper_model
searchFind text across files, folders, and URLsquery, sources, mode (auto · literal · ranked), glob, max_results
inspectGo deeper on a PDFoperation: render_page, extract_regions, ocr_pages, structure (JSON with geometry), compare, inspect

A read answer looks like this:

---
source: papers/attention.pdf
title: Attention Is All You Need
pages: 15
showing: pages 1-9
---

<!-- page 1 -->

# Attention Is All You Need
…

<!-- page 8 -->

|Model|BLEU EN-DE|BLEU EN-FR|
|-|-|-|
|Transformer (big)|28.4|41.8|
…

<!-- Stopped at the 20000-token budget. Continue with cursor: "10", or pick pages, or raise max_tokens. -->

search answers with one line per hit:

5 matches for "masked language model" (2 files, 31 sections searched)

### papers/bert.pdf (5)
- p.1: …by using a “**masked language model**” (MLM) pre-training objective, inspired by the Cloze task…
- p.2: …In addition to the **masked language model**, we also use a “next sentence prediction” task…

If nothing matches exactly, search falls back to BM25-ranked passages, so a question like "how does bidirectional pretraining work" still finds the right page.

CLI

The same binary is a command-line converter, like MarkItDown but much faster:

anymd report.pdf > report.md                 # a file
anymd deck.pptx notes.docx budget.xlsx        # several files, each with a header
anymd https://example.com/article            # a web page (main content only)
cat scan.png | anymd - --ocr                 # stdin, with OCR
anymd paper.pdf --pages 1-3 --max-tokens 4000
anymd search "indemnification" contracts/ --glob '*.pdf'
anymd doctor                                 # lists the optional tools anymd found

Run with no arguments from an MCP client (piped stdin), or as anymd mcp, and it serves MCP over stdio.

Formats

InputWhat you get
PDFReading-order Markdown: headings, paragraphs, lists, tables, sub/superscripts, <!-- page N --> markers, bookmarks as an outline. Running headers and page numbers are removed. Image-only pages are OCR'd when tesseract is installed. Embedded figures are saved to the anymd cache and marked in place with their caption (images: "refs", the default).
Word .docxHeadings, bold/italic, links, nested lists, tables with merged cells, footnotes, equations as LaTeX, embedded pictures as image files
PowerPoint .pptxOne section per slide in deck order, titles, bullets, tables, chart data, speaker notes, pictures as image files
Excel .xlsx .xls .ods · CSV/TSVOne table per sheet, dates as ISO strings, capped at 2,000 rows per sheet
EPUBOne section per chapter in spine order, plus title and author; pictures as image files
HTML and URLsThe main article only: navigation, cookie banners, and sidebars are dropped. Relative links are resolved, and code keeps its language.
Markdown, text, JSONReturned unchanged, with pagination
ImagesDimensions and EXIF (camera, date, GPS), plus OCR text when tesseract is installed
Audio / videoDuration, streams, chapters, embedded and sidecar subtitles (via ffprobe/ffmpeg). Local whisper.cpp transcript with transcript: true; download_whisper_model: true fetches a verified model on first use.

How it works

For PDFs, anymd reads glyph positions rather than text runs. Glyphs are grouped into lines by baseline, which tolerates super- and subscripts. Word spaces come from the gaps between glyphs, measured against the font size and adjusted for letter tracking. A column-aware XY cut finds gutters between running text. Tables come from drawn lines where a table has them (a missing line between two cells makes a merged cell) and from aligned columns of whitespace where it does not. Wrapped cell text stays in its cell, stacked header lines become one header, and a header over several columns is kept with each of them. Text a reader cannot see (invisible text, or text in the colour of the box behind it) is left out. Pages are processed in parallel and isolated from each other, so one malformed page never fails the whole document. The other formats are parsed natively in Rust (zip/XML, calamine, html5ever); no Python, LibreOffice, or cloud service is involved.

Security

  • Local-first: documents never leave your machine unless you pass a URL, and even then only that URL is fetched.
  • URL fetches block private and loopback addresses, and every redirect hop is checked again, pinned to its resolved address.
  • --allow-dir=<path> (repeatable) or MCP_PDF_ALLOWED_DIRS confines the server to the directories you list.
  • Embedded images are written only to anymd's own cache directory (ANYMD_CACHE_DIR, else the platform cache), never next to the source document, and refused over 50 megapixels.
  • External tools (tesseract, ffprobe, whisper.cpp) are optional. anymd runs them without a shell, with a timeout and an output cap.

See SECURITY.md to report a vulnerability.

Also from Sylphx

<!-- generated:also-from -->
  • repomap: A map of your codebase for AI agents: code graph, search, call paths and change impact.
  • lockdocs: Exact-version library docs from your lockfile. Local, offline, no rate limits.
  • skills: Battle-tested agent skills for Claude Code and Codex, installed in one command.
  • readme-mark: Beautiful README images from one URL: banners, badges, icons and stats cards.
<!-- /generated:also-from -->

More from Sylphx: https://sylphx.com/open-source

Star history

Star History Chart

License

MIT © Sylphx

Related MCP servers

CUCue logo

Cue

Active

Cue — video answers with timestamp-level proof. Scenes and frames stay off until asked.

3
TypeScript
MIT
View repository →
IRIris logo

Iris

Active

Iris — image facts with pixel-level proof. OCR runs only when requested.

2
TypeScript
MIT
View repository →
LOlockdocs logo

lockdocs

Active

Exact-version library docs from your lockfile — local, offline, no rate limits.

1
Rust
MIT
View repository →
LOLookout logo

Lookout

Active

Lookout — web answers with source-level proof. Search and fetch citeable excerpts, no API key.

1
TypeScript
MIT
View repository →
RErepomap logo

repomap

Active

A map of your codebase for AI agents: code graph, search, call paths and change impact. No API key.

2
Rust
MIT
View repository →

Symbiotic CLI MCP Server for security scanning and analysis

0
TypeScript
MIT
View repository →