PluginBench
MCP Server
Active
MIT

PyreCrawl MCP Server

io.github.SanggonBoy/PyreCrawl

Web scraping + crawling toolkit for AI agents with anti-bot bypass, no API keys or rate limits.

What is the PyreCrawl MCP server?

PyreCrawl is an MCP server that gives AI agents like Claude and Cursor the ability to scrape, crawl, search, and monitor web pages. It uses a smart auto-fallback ladder (HTTP → stealth browser → deep processing) to bypass Cloudflare and other protections, and returns LLM-ready markdown without requiring API keys or subscriptions.

PyreCrawl enables AI agents to browse and extract data from the web autonomously. It handles static pages, JavaScript-heavy sites, Cloudflare protection, and structured data extraction. The server is self-hosted, free, and includes tools for single-page scraping, multi-page crawling, web search, academic paper search, content monitoring, and persistent browser sessions for login flows.

How to install PyreCrawl

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "PyreCrawl": {
      "command": "uvx",
      "args": [
        "pyrecrawl",
        "serve"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • scrape — Fetch a single URL and return LLM-ready markdown with auto-escalation past Cloudflare and JS rendering
  • extract — Scrape a URL and perform structured extraction using a JsonCss schema to return JSON
  • map_site — Enumerate all internal URLs from a root domain with optional pattern filtering
  • crawl — Multi-page crawl with BFS depth, path filters, and bulk scraping of discovered pages
  • batch_scrape — Fetch many URLs in parallel in a single call with deduplication and cache awareness
  • search — Web search via DuckDuckGo HTML with anti-bot bypass, no API key required
  • search_papers — Academic paper search via arXiv and Crossref, no API key required
  • deep_research — Combined search, scrape, and citation tool that returns evidence packs with source citations
  • document — Extract text from PDF, DOCX, and PPTX URLs and convert to markdown
  • monitor — Track a URL for content changes over time with persisted snapshots and unified diff
  • session — Persistent browser session with cookies for login walls and multi-step flows
  • cache — Inspect, clear, enable, or disable the HTTP response cache
  • health — Verify engine availability and check version information

Use cases

  • Scrape and summarize web pages for research or content aggregation
  • Extract structured data (products, prices, listings) from websites using CSS schemas
  • Monitor competitor websites or news sites for price or content changes
  • Crawl documentation sites or knowledge bases to build RAG datasets
  • Search academic papers and web content with citations for primary research

PyreCrawl MCP server FAQ

What is PyreCrawl?

PyreCrawl is a self-hosted MCP server that gives AI agents web scraping, crawling, and search capabilities with built-in Cloudflare bypass and anti-bot protection. It requires no API keys or subscriptions.

Is PyreCrawl free?

Yes, PyreCrawl is completely free and open-source (MIT license). It is self-hosted, so you run it locally without paying for API calls or subscriptions.

How do I install PyreCrawl in Claude or Cursor?

Run `pyrecrawl install` after installing the package via `uv tool install pyrecrawl` or `pipx install pyrecrawl`. This auto-detects your agent and writes the MCP config. Then restart your agent.

Do I need API keys or authentication?

No. PyreCrawl uses DuckDuckGo for web search and arXiv/Crossref for academic papers—all without API keys. It is fully self-contained.

What happens if a page is protected by Cloudflare?

PyreCrawl automatically escalates through a three-tier ladder: fast HTTP → stealth browser with Cloudflare solver → deep processing. You don't have to choose; `prefer='auto'` handles it.

Can I use PyreCrawl with local LLMs?

Yes. PyreCrawl returns markdown and structured data that any LLM can process. It works with Ollama, local Claude instances, or any MCP-compatible agent.

README (reference)

Source of truth, from the repository.

🔥 PyreCrawl — Web Browsing Superpowers for Your AI Agent

License: MIT MCP Python 3.10+ PyPI GitHub stars Downloads / 30d Active users / 30d

One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search — self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A smart auto-fallback ladder always picks the cheapest method that succeeds:

fast HTTP
    │  (403/503/Cloudflare challenge or empty body)
    ▼
stealth browser (real Chromium + Cloudflare solver)
    │  (still blocked, or the page needs full JS rendering)
    ▼
deep processing (LLM-ready markdown, citations, structured extraction)

⚡ Tools exposed

ToolWhat it does
scrape(url, prefer="auto")Single URL → LLM-ready markdown
extract(url, schema)Scrape + structured extraction (JsonCss schema)
map_site(root, include_pattern=None, limit=200)Enumerate all internal URLs
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0)Multi-page crawl with path filters + true BFS depth
document(url)PDF/DOCX/PPTX → markdown (no browser, optional [docs] extras)
search(query, limit=10)Web search via DuckDuckGo HTML (no API key)
search_papers(query, limit=8, source="arxiv", category=None)Academic search via arXiv + Crossref (no API key) — feed pdf_url into document
batch_scrape(urls[], ...)Many URLs in ONE call — parallel, deduped, cache-aware
deep_research(query, limit=5, scrape_top=3)Search → evidence pack with [n] citations (no LLM synthesis — your agent does that)
monitor(url, action, css_selector=None)Change detection with persisted snapshots + unified diff
session(session, action, ...)Persistent browser session (cookies kept) — login walls, multi-step flows, screenshots
cache(action)Inspect/clear/enable/disable the HTTP response cache
health()Versions + import sanity check

MCP Resources (read-only state without a tool call): pyrecrawl://cache/stats · pyrecrawl://sessions · pyrecrawl://monitors

MCP Prompts (ready-made playbooks): research(topic) · rag_ingest(site) · watch_page(url)

Env flags

VariableDefaultEffect
PYRECRAWL_CACHEoff1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)
PYRECRAWL_CACHE_TTL900Cache entry lifetime in seconds
PYRECRAWL_MONITOR_DIR~/.pyrecrawl/monitorsWhere monitor snapshots persist
PYRECRAWL_NO_TELEMETRYoff1 = disable the anonymous startup ping (also honors DO_NOT_TRACK=1)

prefer options: "auto" (default ladder) · "fast" (HTTP only) · "stealth" (CF bypass) · "llm" (deep processing).


🚀 Install & Use (one-liner)

1. Install

UV (recommended — one command, zero Python setup)

UV is a fast Python package manager that handles Python itself — no need to install Python separately. Get it once:

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Learn more about UV →

Then run PyreCrawl directly — no venv, no pip install, no Python download:

uvx pyrecrawl@latest

Or via uv tool install (persistent, recommended for regular use)

uv tool install pyrecrawl

Or via pipx (alternative)

pipx install pyrecrawl

Or via pip into a venv

pip install pyrecrawl

2. One-time browser engines

pyrecrawl setup

This installs Chromium + stealth browser engines (~2 min, one-time).

3. Register with your AI agent

# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run

Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.

4. Start chatting

After installing + registering, restart your agent (or start a new session). Then ask:

"Scrape https://example.com and summarize it."

Available tools:

ToolWhat it does
scrapeFetch a single URL → markdown (auto-escalates past Cloudflare)
extractScrape + structured extraction via CSS schema → JSON
map_siteEnumerate all internal URLs from a root
crawlMulti-page crawl: discover + scrape in bulk
batch_scrapeFetch many URLs in one parallel call
searchWeb search via DuckDuckGo with anti-bot bypass
search_papersAcademic paper search (arXiv / Crossref)
deep_researchSearch + scrape + citations in one call — primary research tool
documentExtract text from PDF/DOCX/PPTX URLs
monitorTrack a URL for content changes over time
sessionPersistent browser session for login walls
cacheInspect or clear the response cache
healthVerify engine availability + version

Plus 3 guided prompts: research, rag_ingest, watch_page.

Quick examples

Ask your agent naturally — no special syntax needed:

You sayAgent uses
"Scrape https://example.com and summarize it"scrape → returns markdown → agent summarizes
"Research Rust memory safety vulnerabilities"deep_research → search + scrape + citations
"Deep research on AI regulation worldwide"deep_research(iterations=3) → multi-pass with refined queries
"Extract all product names and prices from this page"extract → CSS schema → structured JSON
"Crawl https://docs.example.com and give me an overview"crawl → multi-page → summary
"Monitor this page for price changes"monitor → baseline snapshot → periodic diff
"Find papers about transformer attention"search_papers → arXiv results
"What's the current cache hit rate?"cache → stats

🧠 Skills — Maximize Your Agent's Research Quality

PyreCrawl tools give your agent hands (scrape, crawl, search). But the agent still needs a brain — instructions on when to use which tool, how to chain research passes, and what anti-hallucination rules to follow.

That's what PyreCrawl Skills provides.

MCP Tools (this repo)Skills (pyrecrawl-skills)
RoleExecute web operationsTell the agent how to use them
AnalogyHandsBrain
Exampledeep_research(query, iterations=3)"Run 3 passes, check gaps after each, cite everything"
Required?Yes (the engine)Optional (but recommended for research quality)

Quick setup:

# 1. Install the tools (you already have this)
uvx pyrecrawl@latest

# 2. Add the research skill to your project
git clone https://github.com/SanggonBoy/pyrecrawl-skills.git /tmp/pyrecrawl-skills
cp /tmp/pyrecrawl-skills/pyrecrawl-research/SKILL.md ./CLAUDE.md  # or .cursorrules / AGENTS.md

Without skills: Your agent has powerful tools but improvises usage. With skills: Your agent follows a proven research protocol with anti-hallucination guardrails.


📚 Manual config (if pyrecrawl install doesn't match your setup)

Claude Desktop

Config file

  • Linux: ~/.config/Claude/claude_desktop_config.json
  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %AppData%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Claude Code

Config file: project-scoped .mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Cursor

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

VS Code / Copilot

Config file: .vscode/mcp.json (project-scoped)

{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}

Codex CLI

Config file: ~/.codex/config.toml

[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]

OpenCode

Config file: ~/.config/opencode/opencode.json

{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}

Hermes

Config file

  • Linux/macOS: ~/.hermes/config.yaml
  • Windows: %LocalAppData%\hermes\config.yaml
mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true

Windows note: uvx must be on PATH. If not, use the full path to uvx.exe (e.g. C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).


🧠 How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:

ConcernFast tierStealth tierDeep tier
Static HTML page✅ ~200ms——
Cloudflare-protected❌✅ Turnstile solver—
JS-heavy SPA❌✅ real Chromium—
Live DOM data (input .value, JS state)❌✅ js param—
LLM-ready markdown + citations——✅ BM25, fit-markdown
Structured extraction (CSS schema)——✅
Deep crawl (BFS/DFS/BestFirst)——✅ adaptive

The agent never has to pick. prefer="auto" does it every call.

Live DOM data with js and wait_for

Some sites keep the data you want in a DOM property (e.g. an <input>'s .value) that JS writes after an XHR — it never appears in the serialized HTML. The scrape tool accepts two stealth-tier params for exactly this:

{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
  • wait_for — a JS predicate expression polled until truthy (bounded by timeout). Use it instead of guessing a sleep for anything that arrives asynchronously.
  • js — a JS expression evaluated once the page settles; the value comes back in meta.js_result. Errors are captured in meta.js_error (the page result is still returned, never a crash).

📊 Compared to Firecrawl (hosted)

FirecrawlPyreCrawl
CostFree 1k/mo, then $16–333/moFree, self-hosted
Local LLM support❌✅ Ollama / any LLM
Cloudflare bypass✅ (Fire-Engine, paid)✅ (free, built-in)
Markdown + BM25✅✅
Self-host❌✅
Academic paper search❌✅ arXiv + Crossref (search_papers)
Hosted search API✅ /search⚠️ DuckDuckGo HTML + arXiv/Crossref (no key)

🔧 Development

git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install

Run tests

python scripts/selfcheck.py        # real-network smoke test (13 tools + engines)
python scripts/probe_stdio.py      # stdio JSON-RPC probe
python scripts/test_ladder_bug.py  # SPA-shell ladder escalation regression
python scripts/test_js_eval.py     # stealth js/wait_for params regression
python scripts/test_scope_selector.py  # crawl css_selector/max_depth wiring
python scripts/test_link_harvest.py    # map/BFS link purity regression

📦 Publish

Maintainers only:

git tag vX.Y.Z
git push origin vX.Y.Z

GitHub Actions builds + uploads to PyPI via trusted publishing.


🔔 Stay up to date

PyreCrawl checks PyPI on every startup and reports the latest version — your MCP agent sees this automatically via the health() tool response and can notify you inline.

To check manually:

pyrecrawl version

To upgrade:

pyrecrawl update   # runs: uv tool upgrade pyrecrawl

Get notified of new releases: click Watch → Releases only at the GitHub repo to receive email notifications when a new version is published.


[!NOTE] PyreCrawl sends one anonymous usage ping per 24 h at server startup — see Privacy for exactly what's sent and how to opt out.

🔒 Privacy — anonymous usage ping

PyreCrawl phones home once per 24 h with a tiny anonymous ping when the MCP server starts, so we can count real users (DAU/MAU) instead of raw downloads.

Sent (4 fields, ~100 bytes)Never sent
Hashed machine id (SHA-256 of hostname+MAC — not reversible)Your IP (not stored)
PyreCrawl versionAny URL you scrape
Python versionAny page content or search queries
OS family (windows / linux / darwin)Anything else

Client code: src/pyrecrawl/telemetry.py (~90 lines, stdlib only) · Collector: workers/telemetry/ — a self-hostable Cloudflare Worker + D1, no third-party analytics service.

Opt out any time:

export PYRECRAWL_NO_TELEMETRY=1   # or the industry-standard DO_NOT_TRACK=1

📜 Uninstall

# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl

🛡️ License

MIT — see LICENSE.

<!-- mcp-name: io.github.SanggonBoy/PyreCrawl -->

Related MCP servers

Search, organize, and chat with your saved Reddit posts from Claude, Cursor, and any MCP client.

0
JavaScript
View repository →

MCP server for fetching URLs, safe against SSRF, DNS rebinding, and redirect-to-internal attacks.

1
TypeScript
MIT
View repository →
AIAIppocampus logo

Local-first, source-backed continuity for AI agents via stdio MCP.

5
Python
Apache-2.0
View repository →

Etch is a signed audit chain for AI agent decisions, offline-verifiable against pinned public keys.

0
Python
MIT
View repository →

Persistent memory + post-quantum-signed audit trail for AI coding agents, fully local and offline-verifiable.

9
Python
MIT
View repository →

Manage iOS, macOS, tvOS, and visionOS apps on App Store Connect directly from Claude or Cursor.

24
TypeScript
MIT
View repository →