PluginBench
MCP Server
Maintained
MIT

MCP Document Reader MCP Server

io.github.xt765/mcp_documents_reader

Read DOCX, PDF, Excel, and TXT files directly in Claude and Cursor with a unified MCP tool.

What is the MCP Document Reader MCP server?

The MCP Document Reader is an MCP server that enables AI assistants to read and extract text from multiple document formats including DOCX, PDF, Excel (XLSX/XLS), and TXT files. It provides a single unified interface for document access, allowing Claude and other AI agents to process documents directly from the file system.

This server exposes a read_document tool that lets AI assistants extract and analyze content from common business and text documents. Use it when you need Claude to understand the contents of Word documents, PDFs, spreadsheets, or plain text files during a conversation.

How to install MCP Document Reader

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "mcp_documents_reader": {
      "command": "uvx",
      "args": [
        "mcp-documents-reader"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • read_document — Read any supported document type (DOCX, PDF, Excel, TXT) with a unified interface. Takes a filename parameter (absolute or relative path) and returns the extracted text content.

Use cases

  • Extract and summarize text from PDF reports or Word documents
  • Analyze data from Excel spreadsheets by reading cell contents
  • Process plain text files for content analysis or transformation
  • Enable AI assistants to understand document contents during conversations
  • Batch read multiple documents for comparison or synthesis

MCP Document Reader MCP server FAQ

What document formats does this server support?

It supports DOCX (Word), PDF, Excel (XLSX/XLS), and TXT (plain text) files. Each format is automatically detected and processed with the appropriate reader.

Is this server free to use?

Yes, the MCP Document Reader is open-source under the MIT License and available for free via PyPI.

How do I install it in Claude Desktop or Cursor?

Add the server to your MCP configuration file using `uvx mcp-documents-reader` (PyPI method) or point to the GitHub repository. See the README for three configuration options including a faster Gitee mirror for China.

Does it require authentication or API keys?

No, it reads documents directly from your local file system without requiring any authentication or external API keys.

Can I use it as a Python library outside of MCP?

Yes, you can import DocumentReaderFactory from mcp_documents_reader and use it programmatically in Python code to read documents.

What Python version is required?

Python 3.10 or higher is required to run this server.

README (reference)

Source of truth, from the repository.

<h1 align="center">MCP Document Reader</h1> <!-- mcp-name: io.github.xt765/mcp_documents_reader --> <p align="center"><strong>MCP (Model Context Protocol) Document Reader - A powerful MCP tool for reading documents in multiple formats, enabling AI agents to truly "read" your documents.</strong></p> <p align="center">🌐 <strong>Language</strong>: <a href="README.md">English</a> | <a href="README.zh-CN.md">中文</a></p> <p align="center"> <a href="https://blog.csdn.net/Yunyi_Chi"><img src="https://img.shields.io/badge/CSDN-玄同765-orange.svg?style=flat&logo=csdn" alt="CSDN"></a> <a href="https://github.com/xt765/mcp_documents_reader"><img src="https://img.shields.io/badge/GitHub-mcp_documents_reader-black.svg?style=flat&logo=github" alt="GitHub"></a> <a href="https://gitee.com/xt765/mcp_documents_reader"><img src="https://img.shields.io/badge/Gitee-mcp_documents_reader-red.svg?style=flat&logo=gitee" alt="Gitee"></a> </p> <p align="center"> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue.svg?style=flat&logo=opensourceinitiative" alt="License"></a> <a href="https://www.python.org/downloads/"><img src="https://img.shields.io/badge/python-3.10+-blue.svg?style=flat&logo=python" alt="Python"></a> <a href="https://pypi.org/project/mcp-documents-reader/"><img src="https://img.shields.io/pypi/v/mcp-documents-reader.svg?logo=pypi" alt="PyPI Version"></a> <a href="https://pepy.tech/project/mcp-documents-reader"><img src="https://img.shields.io/pepy/dt/mcp-documents-reader.svg?logo=pypi&label=PyPI%20Downloads" alt="PyPI Downloads"></a> <a href="https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.xt765/mcp_documents_reader"><img src="https://img.shields.io/badge/MCP-Registry-blue?logo=modelcontextprotocol" alt="MCP Registry"></a> <a href="https://mcp-marketplace.io/server/io-github-xt765-mcp-documents-reader"><img src="https://img.shields.io/badge/MCP-Marketplace-22c55e.svg?style=flat&logo=shopify&logoColor=white" alt="MCP Marketplace"></a> </p>

Features

  • Multi-format Support: Supports 4 mainstream document formats: Excel (XLSX/XLS), DOCX, PDF, and TXT
  • MCP Protocol: Compliant with MCP standards, can be used as a tool for AI assistants like Trae IDE
  • Easy Integration: Simple configuration for immediate use
  • Reliable Performance: Successfully tested and running in Trae IDE
  • File System Support: Reads documents directly from the file system

📚 Documentation

User Guide · API Reference · Contributing · Changelog · License


Architecture

graph TB
    A[AI Assistant / User] -->|Call read_document| B[MCP Document Reader]
    B -->|Detect file type| C{File Type?}
    C -->|.docx| D[DOCX Reader]
    C -->|.pdf| E[PDF Reader]
    C -->|.xlsx/.xls| F[Excel Reader]
    C -->|.txt| G[Text Reader]
    D -->|Extract text| H[Return Content]
    E -->|Extract text| H
    F -->|Extract text| H
    G -->|Extract text| H
    H -->|Text content| A
    
    style A fill:#e1f5ff
    style B fill:#fff4e1
    style C fill:#f0f0f0
    style D fill:#e8f5e9
    style E fill:#e8f5e9
    style F fill:#e8f5e9
    style G fill:#e8f5e9
    style H fill:#fff9c4

Supported Formats

FormatExtensionsMIME TypeFeatures
Excel.xlsx, .xlsapplication/vnd.openxmlformats-officedocument.spreadsheetml.sheetSheet and cell data extraction
DOCX.docxapplication/vnd.openxmlformats-officedocument.wordprocessingml.documentText and structure extraction
PDF.pdfapplication/pdfText extraction
Text.txttext/plainPlain text reading

Installation

Using pip (Recommended)

pip install mcp-documents-reader

From Source

git clone https://github.com/xt765/mcp_documents_reader.git
cd mcp_documents_reader
pip install -e .

MCP Tools

This server provides the following tool:

read_document

Read any supported document type with a unified interface.

Arguments:

  • filename (string, required): Document file path, supports absolute or relative paths.

Configuration

Using in Trae IDE / Claude Desktop

Add the following to your MCP configuration file:

Option 1: Using PyPI (Recommended)

{
  "mcpServers": {
    "mcp-document-reader": {
      "command": "uvx",
      "args": [
        "mcp-documents-reader"
      ]
    }
  }
}

Option 2: Using GitHub repository

{
  "mcpServers": {
    "mcp-document-reader": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/xt765/mcp_documents_reader",
        "mcp_documents_reader"
      ]
    }
  }
}

Option 3: Using Gitee repository (Faster access in China)

{
  "mcpServers": {
    "mcp-document-reader": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://gitee.com/xt765/mcp_documents_reader",
        "mcp_documents_reader"
      ]
    }
  }
}

Usage

As an MCP Tool

After configuration, AI assistants can directly call the following tool:

# Read a DOCX file
read_document(filename="example.docx")

# Read a PDF file
read_document(filename="example.pdf")

# Read an Excel file
read_document(filename="example.xlsx")

# Read a text file
read_document(filename="example.txt")

As a Python Library

from mcp_documents_reader import DocumentReaderFactory

# Using factory (recommended)
reader = DocumentReaderFactory.get_reader("document.pdf")
content = reader.read("/path/to/document.pdf")

# Check if format is supported
if DocumentReaderFactory.is_supported("file.xlsx"):
    reader = DocumentReaderFactory.get_reader("file.xlsx")
    content = reader.read("/path/to/file.xlsx")

Tool Interface Details

read_document

Read any supported document type.

Parameters:

ParameterTypeRequiredDescription
filenamestring✅Document file path, supports absolute or relative paths

Dependencies

Core Dependencies

  • mcp >= 1.26.0 - MCP protocol implementation
  • python-docx >= 1.2.0 - DOCX file reading
  • pypdf >= 6.8.0 - PDF file reading (replaces PyPDF2)
  • openpyxl >= 3.1.5 - Excel file reading

Development Dependencies

  • pytest >= 8.0.0 - Testing framework
  • pytest-asyncio >= 0.24.0 - Async testing support
  • pytest-cov >= 6.0.0 - Coverage reporting
  • basedpyright >= 0.28.0 - Type checking
  • ruff >= 0.8.0 - Linting and formatting

License

MIT License

Contributing

Issues and Pull Requests are welcome!

Related Projects

Related MCP servers

Convert between PDF, DOCX, HTML, Markdown, and Text documents for AI assistant context injection.

12
Python
MIT
View repository →

Ukrainian business directory: companies, contacts and registry details (xtrust.info).

0
JavaScript
MIT
View repository →

Cybersecurity MCP server: 323 prompts + 7 workflows for red team, blue team, SOC, cloud, OSINT.

0
JavaScript
View repository →

Markdown memory for AI agents. Files you can read, edit, grep, and commit. Not a database.

1
JavaScript
MIT
View repository →

Deterministic, resumable multi-agent workflows across Codex, Claude, Cursor, and Kimi.

5
JavaScript
MIT
View repository →
OCocular logo

ocular

Active

Vision tools for coding agents: screenshots, OCR, UI diffs, errors, tables, and charts.

0
TypeScript
MIT
View repository →