MCP Document Reader MCP Server
io.github.xt765/mcp_documents_reader
Read DOCX, PDF, Excel, and TXT files directly in Claude and Cursor with a unified MCP tool.
What is the MCP Document Reader MCP server?
The MCP Document Reader is an MCP server that enables AI assistants to read and extract text from multiple document formats including DOCX, PDF, Excel (XLSX/XLS), and TXT files. It provides a single unified interface for document access, allowing Claude and other AI agents to process documents directly from the file system.
This server exposes a read_document tool that lets AI assistants extract and analyze content from common business and text documents. Use it when you need Claude to understand the contents of Word documents, PDFs, spreadsheets, or plain text files during a conversation.
How to install MCP Document Reader
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
read_document— Read any supported document type (DOCX, PDF, Excel, TXT) with a unified interface. Takes a filename parameter (absolute or relative path) and returns the extracted text content.
Use cases
- Extract and summarize text from PDF reports or Word documents
- Analyze data from Excel spreadsheets by reading cell contents
- Process plain text files for content analysis or transformation
- Enable AI assistants to understand document contents during conversations
- Batch read multiple documents for comparison or synthesis
MCP Document Reader MCP server FAQ
It supports DOCX (Word), PDF, Excel (XLSX/XLS), and TXT (plain text) files. Each format is automatically detected and processed with the appropriate reader.
Yes, the MCP Document Reader is open-source under the MIT License and available for free via PyPI.
Add the server to your MCP configuration file using `uvx mcp-documents-reader` (PyPI method) or point to the GitHub repository. See the README for three configuration options including a faster Gitee mirror for China.
No, it reads documents directly from your local file system without requiring any authentication or external API keys.
Yes, you can import DocumentReaderFactory from mcp_documents_reader and use it programmatically in Python code to read documents.
Python 3.10 or higher is required to run this server.
README (reference)
Source of truth, from the repository.
Features
- Multi-format Support: Supports 4 mainstream document formats: Excel (XLSX/XLS), DOCX, PDF, and TXT
- MCP Protocol: Compliant with MCP standards, can be used as a tool for AI assistants like Trae IDE
- Easy Integration: Simple configuration for immediate use
- Reliable Performance: Successfully tested and running in Trae IDE
- File System Support: Reads documents directly from the file system
📚 Documentation
User Guide · API Reference · Contributing · Changelog · License
Architecture
graph TB
A[AI Assistant / User] -->|Call read_document| B[MCP Document Reader]
B -->|Detect file type| C{File Type?}
C -->|.docx| D[DOCX Reader]
C -->|.pdf| E[PDF Reader]
C -->|.xlsx/.xls| F[Excel Reader]
C -->|.txt| G[Text Reader]
D -->|Extract text| H[Return Content]
E -->|Extract text| H
F -->|Extract text| H
G -->|Extract text| H
H -->|Text content| A
style A fill:#e1f5ff
style B fill:#fff4e1
style C fill:#f0f0f0
style D fill:#e8f5e9
style E fill:#e8f5e9
style F fill:#e8f5e9
style G fill:#e8f5e9
style H fill:#fff9c4
Supported Formats
| Format | Extensions | MIME Type | Features |
|---|---|---|---|
| Excel | .xlsx, .xls | application/vnd.openxmlformats-officedocument.spreadsheetml.sheet | Sheet and cell data extraction |
| DOCX | .docx | application/vnd.openxmlformats-officedocument.wordprocessingml.document | Text and structure extraction |
| application/pdf | Text extraction | ||
| Text | .txt | text/plain | Plain text reading |
Installation
Using pip (Recommended)
pip install mcp-documents-reader
From Source
git clone https://github.com/xt765/mcp_documents_reader.git
cd mcp_documents_reader
pip install -e .
MCP Tools
This server provides the following tool:
read_document
Read any supported document type with a unified interface.
Arguments:
filename(string, required): Document file path, supports absolute or relative paths.
Configuration
Using in Trae IDE / Claude Desktop
Add the following to your MCP configuration file:
Option 1: Using PyPI (Recommended)
{
"mcpServers": {
"mcp-document-reader": {
"command": "uvx",
"args": [
"mcp-documents-reader"
]
}
}
}
Option 2: Using GitHub repository
{
"mcpServers": {
"mcp-document-reader": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/xt765/mcp_documents_reader",
"mcp_documents_reader"
]
}
}
}
Option 3: Using Gitee repository (Faster access in China)
{
"mcpServers": {
"mcp-document-reader": {
"command": "uvx",
"args": [
"--from",
"git+https://gitee.com/xt765/mcp_documents_reader",
"mcp_documents_reader"
]
}
}
}
Usage
As an MCP Tool
After configuration, AI assistants can directly call the following tool:
# Read a DOCX file
read_document(filename="example.docx")
# Read a PDF file
read_document(filename="example.pdf")
# Read an Excel file
read_document(filename="example.xlsx")
# Read a text file
read_document(filename="example.txt")
As a Python Library
from mcp_documents_reader import DocumentReaderFactory
# Using factory (recommended)
reader = DocumentReaderFactory.get_reader("document.pdf")
content = reader.read("/path/to/document.pdf")
# Check if format is supported
if DocumentReaderFactory.is_supported("file.xlsx"):
reader = DocumentReaderFactory.get_reader("file.xlsx")
content = reader.read("/path/to/file.xlsx")
Tool Interface Details
read_document
Read any supported document type.
Parameters:
| Parameter | Type | Required | Description |
|---|---|---|---|
| filename | string | ✅ | Document file path, supports absolute or relative paths |
Dependencies
Core Dependencies
mcp>= 1.26.0 - MCP protocol implementationpython-docx>= 1.2.0 - DOCX file readingpypdf>= 6.8.0 - PDF file reading (replaces PyPDF2)openpyxl>= 3.1.5 - Excel file reading
Development Dependencies
pytest>= 8.0.0 - Testing frameworkpytest-asyncio>= 0.24.0 - Async testing supportpytest-cov>= 6.0.0 - Coverage reportingbasedpyright>= 0.28.0 - Type checkingruff>= 0.8.0 - Linting and formatting
License
MIT License
Contributing
Issues and Pull Requests are welcome!
Related Projects
- MCP Document Converter - MCP document converter supporting multiple format conversions
- Model Context Protocol - Official Model Context Protocol documentation
Related MCP servers

io.github.xt765/mcp-document-converter
Convert between PDF, DOCX, HTML, Markdown, and Text documents for AI assistant context injection.
Ukrainian business directory: companies, contacts and registry details (xtrust.info).

io.github.xu-c0/cybersec-mcp
Cybersecurity MCP server: 323 prompts + 7 workflows for red team, blue team, SOC, cloud, OSINT.

io.github.xultrax-web/agent-memory-mcp
Markdown memory for AI agents. Files you can read, edit, grep, and commit. Not a database.

Deterministic, resumable multi-agent workflows across Codex, Claude, Cursor, and Kimi.

ocular
Vision tools for coding agents: screenshots, OCR, UI diffs, errors, tables, and charts.
