PyreCrawl MCP Server
io.github.SanggonBoy/PyreCrawl
Web scraping + crawling toolkit for AI agents with anti-bot bypass, no API keys or rate limits.
What is the PyreCrawl MCP server?
PyreCrawl is an MCP server that gives AI agents like Claude and Cursor the ability to scrape, crawl, search, and monitor web pages. It uses a smart auto-fallback ladder (HTTP → stealth browser → deep processing) to bypass Cloudflare and other protections, and returns LLM-ready markdown without requiring API keys or subscriptions.
PyreCrawl enables AI agents to browse and extract data from the web autonomously. It handles static pages, JavaScript-heavy sites, Cloudflare protection, and structured data extraction. The server is self-hosted, free, and includes tools for single-page scraping, multi-page crawling, web search, academic paper search, content monitoring, and persistent browser sessions for login flows.
How to install PyreCrawl
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
scrape— Fetch a single URL and return LLM-ready markdown with auto-escalation past Cloudflare and JS renderingextract— Scrape a URL and perform structured extraction using a JsonCss schema to return JSONmap_site— Enumerate all internal URLs from a root domain with optional pattern filteringcrawl— Multi-page crawl with BFS depth, path filters, and bulk scraping of discovered pagesbatch_scrape— Fetch many URLs in parallel in a single call with deduplication and cache awarenesssearch— Web search via DuckDuckGo HTML with anti-bot bypass, no API key requiredsearch_papers— Academic paper search via arXiv and Crossref, no API key requireddeep_research— Combined search, scrape, and citation tool that returns evidence packs with source citationsdocument— Extract text from PDF, DOCX, and PPTX URLs and convert to markdownmonitor— Track a URL for content changes over time with persisted snapshots and unified diffsession— Persistent browser session with cookies for login walls and multi-step flowscache— Inspect, clear, enable, or disable the HTTP response cachehealth— Verify engine availability and check version information
Use cases
- Scrape and summarize web pages for research or content aggregation
- Extract structured data (products, prices, listings) from websites using CSS schemas
- Monitor competitor websites or news sites for price or content changes
- Crawl documentation sites or knowledge bases to build RAG datasets
- Search academic papers and web content with citations for primary research
PyreCrawl MCP server FAQ
PyreCrawl is a self-hosted MCP server that gives AI agents web scraping, crawling, and search capabilities with built-in Cloudflare bypass and anti-bot protection. It requires no API keys or subscriptions.
Yes, PyreCrawl is completely free and open-source (MIT license). It is self-hosted, so you run it locally without paying for API calls or subscriptions.
Run `pyrecrawl install` after installing the package via `uv tool install pyrecrawl` or `pipx install pyrecrawl`. This auto-detects your agent and writes the MCP config. Then restart your agent.
No. PyreCrawl uses DuckDuckGo for web search and arXiv/Crossref for academic papers—all without API keys. It is fully self-contained.
PyreCrawl automatically escalates through a three-tier ladder: fast HTTP → stealth browser with Cloudflare solver → deep processing. You don't have to choose; `prefer='auto'` handles it.
Yes. PyreCrawl returns markdown and structured data that any LLM can process. It works with Ollama, local Claude instances, or any MCP-compatible agent.
README (reference)
Source of truth, from the repository.
🔥 PyreCrawl — Web Browsing Superpowers for Your AI Agent
One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search — self-hosted, no API keys, no rate limits, no subscription.
PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.
A smart auto-fallback ladder always picks the cheapest method that succeeds:
fast HTTP
│ (403/503/Cloudflare challenge or empty body)
▼
stealth browser (real Chromium + Cloudflare solver)
│ (still blocked, or the page needs full JS rendering)
▼
deep processing (LLM-ready markdown, citations, structured extraction)
⚡ Tools exposed
| Tool | What it does |
|---|---|
scrape(url, prefer="auto") | Single URL → LLM-ready markdown |
extract(url, schema) | Scrape + structured extraction (JsonCss schema) |
map_site(root, include_pattern=None, limit=200) | Enumerate all internal URLs |
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0) | Multi-page crawl with path filters + true BFS depth |
document(url) | PDF/DOCX/PPTX → markdown (no browser, optional [docs] extras) |
search(query, limit=10) | Web search via DuckDuckGo HTML (no API key) |
search_papers(query, limit=8, source="arxiv", category=None) | Academic search via arXiv + Crossref (no API key) — feed pdf_url into document |
batch_scrape(urls[], ...) | Many URLs in ONE call — parallel, deduped, cache-aware |
deep_research(query, limit=5, scrape_top=3) | Search → evidence pack with [n] citations (no LLM synthesis — your agent does that) |
monitor(url, action, css_selector=None) | Change detection with persisted snapshots + unified diff |
session(session, action, ...) | Persistent browser session (cookies kept) — login walls, multi-step flows, screenshots |
cache(action) | Inspect/clear/enable/disable the HTTP response cache |
health() | Versions + import sanity check |
MCP Resources (read-only state without a tool call):
pyrecrawl://cache/stats · pyrecrawl://sessions · pyrecrawl://monitors
MCP Prompts (ready-made playbooks): research(topic) · rag_ingest(site) · watch_page(url)
Env flags
| Variable | Default | Effect |
|---|---|---|
PYRECRAWL_CACHE | off | 1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode) |
PYRECRAWL_CACHE_TTL | 900 | Cache entry lifetime in seconds |
PYRECRAWL_MONITOR_DIR | ~/.pyrecrawl/monitors | Where monitor snapshots persist |
PYRECRAWL_NO_TELEMETRY | off | 1 = disable the anonymous startup ping (also honors DO_NOT_TRACK=1) |
prefer options: "auto" (default ladder) · "fast" (HTTP only) · "stealth" (CF bypass) · "llm" (deep processing).
🚀 Install & Use (one-liner)
1. Install
UV (recommended — one command, zero Python setup)
UV is a fast Python package manager that handles Python itself — no need to install Python separately. Get it once:
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
Then run PyreCrawl directly — no venv, no pip install, no Python download:
uvx pyrecrawl@latest
Or via uv tool install (persistent, recommended for regular use)
uv tool install pyrecrawl
Or via pipx (alternative)
pipx install pyrecrawl
Or via pip into a venv
pip install pyrecrawl
2. One-time browser engines
pyrecrawl setup
This installs Chromium + stealth browser engines (~2 min, one-time).
3. Register with your AI agent
# Auto-detect installed agents and write their MCP configs
pyrecrawl install
# Or target specific agents
pyrecrawl install claude-desktop cursor
# Dry-run to preview what would change
pyrecrawl install --dry-run
Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.
4. Start chatting
After installing + registering, restart your agent (or start a new session). Then ask:
"Scrape https://example.com and summarize it."
Available tools:
| Tool | What it does |
|---|---|
scrape | Fetch a single URL → markdown (auto-escalates past Cloudflare) |
extract | Scrape + structured extraction via CSS schema → JSON |
map_site | Enumerate all internal URLs from a root |
crawl | Multi-page crawl: discover + scrape in bulk |
batch_scrape | Fetch many URLs in one parallel call |
search | Web search via DuckDuckGo with anti-bot bypass |
search_papers | Academic paper search (arXiv / Crossref) |
deep_research | Search + scrape + citations in one call — primary research tool |
document | Extract text from PDF/DOCX/PPTX URLs |
monitor | Track a URL for content changes over time |
session | Persistent browser session for login walls |
cache | Inspect or clear the response cache |
health | Verify engine availability + version |
Plus 3 guided prompts: research, rag_ingest, watch_page.
Quick examples
Ask your agent naturally — no special syntax needed:
| You say | Agent uses |
|---|---|
| "Scrape https://example.com and summarize it" | scrape → returns markdown → agent summarizes |
| "Research Rust memory safety vulnerabilities" | deep_research → search + scrape + citations |
| "Deep research on AI regulation worldwide" | deep_research(iterations=3) → multi-pass with refined queries |
| "Extract all product names and prices from this page" | extract → CSS schema → structured JSON |
| "Crawl https://docs.example.com and give me an overview" | crawl → multi-page → summary |
| "Monitor this page for price changes" | monitor → baseline snapshot → periodic diff |
| "Find papers about transformer attention" | search_papers → arXiv results |
| "What's the current cache hit rate?" | cache → stats |
🧠 Skills — Maximize Your Agent's Research Quality
PyreCrawl tools give your agent hands (scrape, crawl, search). But the agent still needs a brain — instructions on when to use which tool, how to chain research passes, and what anti-hallucination rules to follow.
That's what PyreCrawl Skills provides.
| MCP Tools (this repo) | Skills (pyrecrawl-skills) | |
|---|---|---|
| Role | Execute web operations | Tell the agent how to use them |
| Analogy | Hands | Brain |
| Example | deep_research(query, iterations=3) | "Run 3 passes, check gaps after each, cite everything" |
| Required? | Yes (the engine) | Optional (but recommended for research quality) |
Quick setup:
# 1. Install the tools (you already have this)
uvx pyrecrawl@latest
# 2. Add the research skill to your project
git clone https://github.com/SanggonBoy/pyrecrawl-skills.git /tmp/pyrecrawl-skills
cp /tmp/pyrecrawl-skills/pyrecrawl-research/SKILL.md ./CLAUDE.md # or .cursorrules / AGENTS.md
Without skills: Your agent has powerful tools but improvises usage. With skills: Your agent follows a proven research protocol with anti-hallucination guardrails.
📚 Manual config (if pyrecrawl install doesn't match your setup)
Claude Desktop
Config file
- Linux:
~/.config/Claude/claude_desktop_config.json - macOS:
~/Library/Application Support/Claude/claude_desktop_config.json - Windows:
%AppData%\Claude\claude_desktop_config.json
{
"mcpServers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
}
}
}
Claude Code
Config file: project-scoped .mcp.json
{
"mcpServers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
}
}
}
Cursor
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
}
}
}
VS Code / Copilot
Config file: .vscode/mcp.json (project-scoped)
{
"servers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
"type": "stdio"
}
}
}
Codex CLI
Config file: ~/.codex/config.toml
[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]
OpenCode
Config file: ~/.config/opencode/opencode.json
{
"mcp": {
"pyrecrawl": {
"type": "local",
"command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
"enabled": true
}
}
}
Hermes
Config file
- Linux/macOS:
~/.hermes/config.yaml - Windows:
%LocalAppData%\hermes\config.yaml
mcp_servers:
pyrecrawl:
command: uvx
args:
- --from
- pyrecrawl
- pyrecrawl
- serve
enabled: true
Windows note:
uvxmust be on PATH. If not, use the full path touvx.exe(e.g.C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).
🧠 How the ladder chooses
PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:
| Concern | Fast tier | Stealth tier | Deep tier |
|---|---|---|---|
| Static HTML page | ✅ ~200ms | — | — |
| Cloudflare-protected | ❌ | ✅ Turnstile solver | — |
| JS-heavy SPA | ❌ | ✅ real Chromium | — |
Live DOM data (input .value, JS state) | ❌ | ✅ js param | — |
| LLM-ready markdown + citations | — | — | ✅ BM25, fit-markdown |
| Structured extraction (CSS schema) | — | — | ✅ |
| Deep crawl (BFS/DFS/BestFirst) | — | — | ✅ adaptive |
The agent never has to pick. prefer="auto" does it every call.
Live DOM data with js and wait_for
Some sites keep the data you want in a DOM property (e.g. an <input>'s .value)
that JS writes after an XHR — it never appears in the serialized HTML. The
scrape tool accepts two stealth-tier params for exactly this:
{
"url": "https://temp-mail.org/id",
"prefer": "stealth",
"wait_for": "document.getElementById('mail').value.includes('@')",
"js": "document.getElementById('mail').value"
}
wait_for— a JS predicate expression polled until truthy (bounded bytimeout). Use it instead of guessing a sleep for anything that arrives asynchronously.js— a JS expression evaluated once the page settles; the value comes back inmeta.js_result. Errors are captured inmeta.js_error(the page result is still returned, never a crash).
📊 Compared to Firecrawl (hosted)
| Firecrawl | PyreCrawl | |
|---|---|---|
| Cost | Free 1k/mo, then $16–333/mo | Free, self-hosted |
| Local LLM support | ❌ | ✅ Ollama / any LLM |
| Cloudflare bypass | ✅ (Fire-Engine, paid) | ✅ (free, built-in) |
| Markdown + BM25 | ✅ | ✅ |
| Self-host | ❌ | ✅ |
| Academic paper search | ❌ | ✅ arXiv + Crossref (search_papers) |
| Hosted search API | ✅ /search | ⚠️ DuckDuckGo HTML + arXiv/Crossref (no key) |
🔧 Development
git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install
Run tests
python scripts/selfcheck.py # real-network smoke test (13 tools + engines)
python scripts/probe_stdio.py # stdio JSON-RPC probe
python scripts/test_ladder_bug.py # SPA-shell ladder escalation regression
python scripts/test_js_eval.py # stealth js/wait_for params regression
python scripts/test_scope_selector.py # crawl css_selector/max_depth wiring
python scripts/test_link_harvest.py # map/BFS link purity regression
📦 Publish
Maintainers only:
git tag vX.Y.Z
git push origin vX.Y.Z
GitHub Actions builds + uploads to PyPI via trusted publishing.
🔔 Stay up to date
PyreCrawl checks PyPI on every startup and reports the latest version — your
MCP agent sees this automatically via the health() tool response and can
notify you inline.
To check manually:
pyrecrawl version
To upgrade:
pyrecrawl update # runs: uv tool upgrade pyrecrawl
Get notified of new releases: click Watch → Releases only at the GitHub repo to receive email notifications when a new version is published.
[!NOTE] PyreCrawl sends one anonymous usage ping per 24 h at server startup — see Privacy for exactly what's sent and how to opt out.
🔒 Privacy — anonymous usage ping
PyreCrawl phones home once per 24 h with a tiny anonymous ping when the MCP server starts, so we can count real users (DAU/MAU) instead of raw downloads.
| Sent (4 fields, ~100 bytes) | Never sent |
|---|---|
| Hashed machine id (SHA-256 of hostname+MAC — not reversible) | Your IP (not stored) |
| PyreCrawl version | Any URL you scrape |
| Python version | Any page content or search queries |
OS family (windows / linux / darwin) | Anything else |
Client code: src/pyrecrawl/telemetry.py (~90 lines, stdlib only) ·
Collector: workers/telemetry/ — a self-hostable Cloudflare Worker + D1, no third-party analytics service.
Opt out any time:
export PYRECRAWL_NO_TELEMETRY=1 # or the industry-standard DO_NOT_TRACK=1
📜 Uninstall
# Remove from all agent configs
pyrecrawl uninstall
# Remove the package
uv tool uninstall pyrecrawl
🛡️ License
MIT — see LICENSE.
<!-- mcp-name: io.github.SanggonBoy/PyreCrawl -->Related MCP servers
Search, organize, and chat with your saved Reddit posts from Claude, Cursor, and any MCP client.

MCP server for fetching URLs, safe against SSRF, DNS rebinding, and redirect-to-internal attacks.

AIppocampus
Local-first, source-backed continuity for AI agents via stdio MCP.
Etch is a signed audit chain for AI agent decisions, offline-verifiable against pinned public keys.
Persistent memory + post-quantum-signed audit trail for AI coding agents, fully local and offline-verifiable.
Manage iOS, macOS, tvOS, and visionOS apps on App Store Connect directly from Claude or Cursor.


