io.github.blazickjp/arxiv-mcp-server MCP Server
io.github.blazickjp/arxiv-mcp-server
Search, download, and analyze arXiv papers with semantic search, citation graphs, and research alerts via MCP.
What is the io.github.blazickjp/arxiv-mcp-server MCP server?
The ArXiv MCP Server is a bridge between AI assistants and arXiv's research repository through the Model Context Protocol. It enables AI models to search for papers, download full text, perform semantic search across local collections, explore citation graphs, and set up research alerts—all programmatically.
This server connects Claude and other AI assistants to arXiv's research database. You can search papers by keyword with filters for date ranges and categories, download full text (HTML or PDF), read papers locally, perform semantic similarity searches across your downloaded collection, explore citation relationships via Semantic Scholar, and set up alerts for new papers matching your research interests. It's designed for researchers, students, and AI agents that need to access and analyze academic papers at scale.
How to install io.github.blazickjp/arxiv-mcp-server
Copy-paste configuration for popular MCP clients.
ARXIV_STORAGE_PATHOptional path for storing downloaded papers locally.
Tools & capabilities
Tools this server exposes to the agent.
search_papers— Query arXiv papers with optional filters for date ranges, categories, and boolean search operators. Enforces arXiv's 3-second rate limit automatically.download_paper— Download a paper by arXiv ID in HTML or PDF format. Stores locally for later access. Supports pagination for very large papers.list_papers— List all papers downloaded locally. Returns arXiv IDs only.read_paper— Read the full text of a locally downloaded paper in markdown format. Supports pagination for large papers.semantic_search— Semantic similarity search across locally downloaded papers only. Requires [pro] dependencies.citation_graph— Fetch references and citing papers via Semantic Scholar for any arXiv ID. Requires [pro] dependencies.watch_topic— Register a topic watch to monitor for newly published papers matching a query. Requires [pro] dependencies.check_alerts— Poll all registered topic watches for newly published papers since last check. Requires [pro] dependencies.
Use cases
- Search for papers on specific topics (e.g., Kolmogorov-Arnold Networks) and download full text for analysis
- Perform semantic similarity searches across your local paper collection to find related work
- Explore citation networks to understand how papers reference each other and identify influential work
- Set up research alerts to automatically track new papers published in your areas of interest
- Analyze and summarize academic papers with AI assistance by feeding full text into Claude's context
io.github.blazickjp/arxiv-mcp-server MCP server FAQ
It's an MCP server that gives AI assistants like Claude programmatic access to arXiv's research repository. You can search papers, download full text, perform semantic searches, explore citations, and set up alerts for new papers.
Yes. The server itself is open-source (Apache 2.0 license) and free. arXiv access is also free, though the server enforces arXiv's 3-second rate limit between requests.
The easiest way is via Smithery: run `npx -y @smithery/cli install arxiv-mcp-server --client claude`. Alternatively, download the `.mcpb` bundle from the latest release and drag it into Claude Desktop's Extensions panel. For manual setup, use `uv tool install arxiv-mcp-server` and add the MCP config to your Claude Desktop settings.
No. The server accesses arXiv's public API directly, which requires no authentication. For citation graph features (Semantic Scholar), no API key is required either.
arXiv papers are user-generated, untrusted content. A maliciously crafted paper could contain prompt injection attacks designed to manipulate the AI's behavior. The README recommends using read-only MCP configurations, reviewing AI summaries before acting on them, and being cautious in multi-tool setups.
Base install handles HTML papers (most arXiv papers). The `[pdf]` extra adds support for older PDF-only papers. The `[pro]` extra enables semantic search, citation graphs, and research alerts.
README (reference)
Source of truth, from the repository.
ArXiv MCP Server
<!-- mcp-name: io.github.blazickjp/arxiv-mcp-server -->🔍 Enable AI assistants to search and access arXiv papers through a simple MCP interface.
The ArXiv MCP Server provides a bridge between AI assistants and arXiv's research repository through the Model Context Protocol (MCP). It allows AI models to search for papers and access their content in a programmatic way.
<div align="center">🤝 Contribute • 📝 Report Bug
<a href="https://www.pulsemcp.com/servers/blazickjp-arxiv-mcp-server"><img src="https://www.pulsemcp.com/badge/top-pick/blazickjp-arxiv-mcp-server" width="400" alt="Pulse MCP Badge"></a>
</div>✨ Core Features
- 🔎 Paper Search: Query arXiv papers with filters for date ranges and categories
- 📄 Paper Access: Download and read paper content
- 📋 Paper Listing: View all downloaded papers
- 🗃️ Local Storage: Papers are saved locally for faster access
- 📝 Prompts: A set of research prompts for paper analysis
🔒 Security
Prompt Injection Risk
Paper content retrieved from arXiv is untrusted external input.
When an AI assistant downloads or reads a paper through this server, the paper's text is passed directly into the model's context. A maliciously crafted paper could embed adversarial instructions designed to hijack the AI's behavior — for example, instructing it to exfiltrate data, invoke other tools with unintended arguments, or override system-level instructions. This is a known class of attack described by OWASP as LLM01: Prompt Injection and by the OWASP Agentic AI framework as AG01: Prompt Injection in LLM-Integrated Systems.
Recommended Mitigations
- Use read-only MCP configurations — where possible, configure the MCP client so that the arxiv-mcp-server cannot trigger write operations or invoke other tools on your behalf.
- Review paper content before acting on AI summaries — if an AI summary asks you to run commands or visit external URLs that were not part of your original request, treat that as a red flag.
- Be cautious in multi-tool setups — agentic pipelines that combine this server with filesystem, shell, or browser tools are higher risk; a prompt injection in a paper could chain tool calls unexpectedly.
- Treat AI-generated summaries as data, not instructions — always apply human judgment before executing any action the AI recommends after reading a paper.
References
🚀 Quick Start
Installing via Smithery
To install ArXiv Server for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install arxiv-mcp-server --client claude
Installing via Claude Desktop (.mcpb)
The .mcpb bundle is the one-click install path for Claude Desktop on macOS. It bundles the server code and Python package dependencies, so users do not need uv, pip, or manual MCP JSON configuration. Python 3.11+ must still be available on the user's machine.
- Download the artifact matching your Mac from the latest release:
- Apple Silicon:
arxiv-mcp-server-darwin-arm64-<version>.mcpb - Intel:
arxiv-mcp-server-darwin-x86_64-<version>.mcpb
- Apple Silicon:
- In Claude Desktop open Settings → Extensions (or drag-and-drop the file onto the Claude Desktop window).
- Click Install and, when prompted, set your preferred paper storage directory (defaults to
~/.arxiv-mcp-server/papers).
Claude Desktop launches the bundled server over stdio — no configuration file edits needed.
Installing Manually
Important — use
uv tool install, not npm/pnpm oruv pip installThis project publishes the supported server as a Python package on PyPI. Do not install
arxiv-mcp-serverwithnpm install,pnpm add, ornpx arxiv-mcp-server: the npm package with this name is an unrelated third-party package and has its own Python-detection wrapper.Running
uv pip install arxiv-mcp-serverinstalls the package into the current virtual environment but does not place thearxiv-mcp-serverexecutable on yourPATH. You must useuv tool installso that uv creates an isolated environment and exposes the executable globally:
uv tool install arxiv-mcp-server
After this, the arxiv-mcp-server command will be available on your PATH.
PDF fallback (older papers): Most arXiv papers have an HTML version which the base install handles automatically. For older papers that only have a PDF, the server needs the
[pdf]extra (pymupdf4llm). Install it with:uv tool install 'arxiv-mcp-server[pdf]'
You can verify it with:
arxiv-mcp-server --help
If you previously ran uv pip install arxiv-mcp-server and the command is
missing, uninstall it and re-install with uv tool install as shown above.
For development:
# Clone and set up development environment
git clone https://github.com/blazickjp/arxiv-mcp-server.git
cd arxiv-mcp-server
# Create and activate virtual environment
uv venv
source .venv/bin/activate
# Install with test dependencies (development only — no global executable)
uv pip install -e ".[test]"
🤖 Codex Plugin Integration
This repository now includes a Codex plugin manifest at .codex-plugin/plugin.json
and a portable MCP config at .mcp.json so Codex-oriented tooling can discover
the server without inventing its own install recipe.
The Codex integration uses the same stdio launch path documented elsewhere in this README:
{
"mcpServers": {
"arxiv": {
"command": "uvx",
"args": ["arxiv-mcp-server"]
}
}
}
If your Codex client supports plugin manifests, point it at
./.codex-plugin/plugin.json. If it only supports raw MCP configuration, use
./.mcp.json directly.
🔌 MCP Integration
Add this configuration to your MCP client config file:
{
"mcpServers": {
"arxiv-mcp-server": {
"command": "uv",
"args": [
"tool",
"run",
"arxiv-mcp-server",
"--storage-path", "/path/to/paper/storage"
]
}
}
}
For Development:
{
"mcpServers": {
"arxiv-mcp-server": {
"command": "uv",
"args": [
"--directory",
"path/to/cloned/arxiv-mcp-server",
"run",
"arxiv-mcp-server",
"--storage-path", "/path/to/paper/storage"
]
}
}
}
HTTP Transport
For server deployments where stdio is not practical, run the server with Streamable HTTP:
TRANSPORT=http HOST=127.0.0.1 PORT=8080 arxiv-mcp-server --storage-path /path/to/papers
Then configure an MCP client that supports Streamable HTTP:
{
"mcpServers": {
"arxiv-mcp-server": {
"type": "http",
"url": "http://127.0.0.1:8080/mcp"
}
}
}
The default HTTP bind host is 127.0.0.1. Streamable HTTP enables MCP DNS rebinding protection by default and allows loopback hosts for the configured port. If exposing the server through a reverse proxy, keep it bound to localhost unless you have added authentication and network controls upstream; set ALLOWED_HOSTS and ALLOWED_ORIGINS to the external host/origin values your proxy forwards.
🔒 Security Note
arXiv papers are user-generated, untrusted content. Paper text returned by this server may contain prompt injection attempts — crafted text designed to manipulate an AI assistant's behavior. Treat all paper content as untrusted input.
In production environments, apply appropriate sandboxing and avoid feeding raw paper content into agentic pipelines that have access to sensitive tools or data without review. See SECURITY.md for the full security policy.
💡 Available Tools
Core Workflow
The typical workflow for deep paper research is:
search_papers → download_paper → read_paper
list_papers shows what you have locally. semantic_search searches across your local collection.
1. Paper Search
Search arXiv with optional category, date, and boolean filters. Enforces arXiv's 3-second rate limit automatically. If rate limited, wait 60 seconds before retrying.
result = await call_tool("search_papers", {
"query": "\"KAN\" OR \"Kolmogorov-Arnold Networks\"",
"max_results": 10,
"date_from": "2024-01-01",
"categories": ["cs.LG", "cs.AI"],
"sort_by": "date" # or "relevance" (default)
})
Supported categories include cs.AI, cs.LG, cs.CL, cs.CV, cs.NE, stat.ML, math.OC, quant-ph, eess.SP, and more. See tool description for the full list.
2. Paper Download
Download a paper by its arXiv ID. Tries HTML first, falls back to PDF. Stores the paper locally for read_paper and semantic_search. The response includes content_length, returned_chars, next_start, and is_truncated so clients can safely page through very large papers without mistaking client-side output caps for failed downloads.
result = await call_tool("download_paper", {
"paper_id": "2401.12345"
})
# For very large papers, request bounded chunks:
result = await call_tool("download_paper", {
"paper_id": "2401.12345",
"start": 0,
"max_chars": 50000
})
For older papers that only have a PDF, install the
[pdf]extra:uv tool install 'arxiv-mcp-server[pdf]'
3. List Papers
List all papers downloaded locally. Returns arXiv IDs only — use read_paper to access content.
result = await call_tool("list_papers", {})
4. Read Paper
Read the full text of a locally downloaded paper in markdown. Requires download_paper to be called first. Use start and max_chars with the returned next_start value to page through large papers.
result = await call_tool("read_paper", {
"paper_id": "2401.12345"
})
result = await call_tool("read_paper", {
"paper_id": "2401.12345",
"start": 50000,
"max_chars": 50000
})
📝 Research Prompts
The server offers specialized prompts to help analyze academic papers:
Paper Analysis Prompt
A comprehensive workflow for analyzing academic papers that only requires a paper ID:
result = await call_prompt("deep-paper-analysis", {
"paper_id": "2401.12345"
})
This prompt includes:
- Detailed instructions for using available tools (list_papers, download_paper, read_paper, search_papers)
- A systematic workflow for paper analysis
- Comprehensive analysis structure covering:
- Executive summary
- Research context
- Methodology analysis
- Results evaluation
- Practical and theoretical implications
- Future research directions
- Broader impacts
Pro Prompt Pack
summarize_paper: concise structured summary for one paper.compare_papers: side-by-side technical comparison across paper IDs.literature_review: thematic synthesis across a topic and optional paper set.
⚙️ Configuration
Configure through command-line options and environment variables:
| Setting | Purpose | Default |
|---|---|---|
--storage-path | Paper storage location | ~/.arxiv-mcp-server/papers |
MAX_RESULTS | Maximum search results | 50 |
REQUEST_TIMEOUT | API timeout in seconds | 60 |
TRANSPORT | Transport type: stdio, http, or streamable-http | stdio |
HOST | Host to bind to in HTTP mode | 127.0.0.1 |
PORT | Port to listen on in HTTP mode | 8000 |
ALLOWED_HOSTS | Comma-separated extra allowed Host header values for Streamable HTTP DNS rebinding protection | empty |
ALLOWED_ORIGINS | Comma-separated extra allowed Origin header values for Streamable HTTP DNS rebinding protection | empty |
🧪 Testing
Run the test suite:
python -m pytest
🧪 Experimental Features
These features are not yet fully tested and may behave unexpectedly. Use with caution.
The following tools require additional dependencies and are under active development:
uv pip install -e ".[pro]"
Semantic Search
Semantic similarity search over your locally downloaded papers only. Returns empty results if no papers have been downloaded yet. Requires [pro] dependencies.
result = await call_tool("semantic_search", {
"query": "test-time adaptation in multimodal transformers",
"max_results": 5
})
# or find papers similar to a known paper:
result = await call_tool("semantic_search", {
"paper_id": "2404.19756",
"max_results": 5
})
Citation Graph
Fetch references and citing papers via Semantic Scholar. Works on any arXiv ID — no local download required.
result = await call_tool("citation_graph", {
"paper_id": "2401.12345"
})
Research Alerts
Save topic watches and poll for newly published papers since the last check. Uses the same query syntax as search_papers.
# Register a watch (idempotent — calling again updates the existing watch)
await call_tool("watch_topic", {
"topic": "\"multi-agent reinforcement learning\"",
"categories": ["cs.AI", "cs.LG"],
"max_results": 10
})
# Check all watches — returns only papers published since last check
result = await call_tool("check_alerts", {})
# Check a single watch
result = await call_tool("check_alerts", {"topic": "\"multi-agent reinforcement learning\""})
Advanced Prompts
summarize_paper, compare_papers, and literature_review for deeper research workflows. Requires [pro] dependencies.
📄 License
Released under the Apache License 2.0. See the LICENSE file for details.
<div align="center">
Made with ❤️ by the Pearl Labs Team
<a href="https://glama.ai/mcp/servers/04dtxi5i5n"><img width="380" height="200" src="https://glama.ai/mcp/servers/04dtxi5i5n/badge" alt="ArXiv Server MCP server" /></a>
</div>Related MCP servers

Dear User
Tells you how you and your Claude agent actually work together. Local-only, no API keys.

A content exploration toolkit that helps LLMs surface high signal, unsummarized web content.

io.github.blipemail/email
MCP server for Blip disposable email — create inboxes, receive emails, extract OTP codes

io.github.blitzdotdev/blitz
Automate iOS App Store Connect submissions—simulators, IAPs, screenshots, and review—via AI agents.

Model Ledger
Model inventory and governance: trace deployed models and their dependencies as an append-only log
Crypto MCP — 99 tools across news, market, on-chain, sentiment, indicators, agent generation.
