PluginBench
MCP Server
Active
MIT

Crawlberg MCP Server

io.github.xberg-io/crawlberg

Scrape and crawl websites into clean Markdown or JSON with a single command or API call.

What is the Crawlberg MCP server?

Crawlberg is a web scraping and crawling engine that converts websites into structured Markdown, metadata, and links. It handles JavaScript-heavy pages, respects robots.txt, detects bot filters, and provides an MCP server for AI agents plus language bindings for Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, Kotlin, and WebAssembly.

Crawlberg turns raw HTML into clean, structured data. Point it at a URL and get back Markdown with metadata (titles, links, images, JSON-LD, Open Graph), either from a single page or an entire crawled site. It handles the hard parts: JavaScript rendering via headless browser, bot-filter detection, rate limiting, SSRF protection, and robots.txt compliance. Use it from code, an AI agent, REST API, or CLI.

How to install Crawlberg

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "crawlberg": {
      "command": "docker",
      "args": [
        "run",
        "-i",
        "--rm",
        "ghcr.io/xberg-io/crawlberg:1.8.0",
        "mcp"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • crawl — Crawl a website with configurable depth, page limits, and concurrency; supports depth-first, breadth-first, or best-first traversal.
  • scrape — Scrape a single URL or batch of URLs and extract structured data: text, metadata, links, images, assets, JSON-LD, Open Graph, hreflang, favicons, headings, and response headers.
  • markdown_conversion — Convert HTML to clean Markdown with citations, document structure, and fit-content mode.
  • url_filtering — Filter URLs by include/exclude patterns, BM25 relevance scoring, and sitemap discovery.
  • browser_rendering — Optional headless browser rendering for JavaScript-heavy single-page applications with WAF detection and bypass.
  • authentication — Support for HTTP Basic, Bearer token, and custom-header authentication with cookie jar management.
  • rate_limiting — Per-domain request throttling and concurrent crawl control.

Use cases

  • Extract structured content from a website and convert it to Markdown for documentation or knowledge bases.
  • Crawl an entire site to build a searchable index or feed data into a vector database for RAG.
  • Scrape product listings, pricing, or metadata from e-commerce sites for competitive analysis.
  • Fetch and parse JSON-LD, Open Graph, and microdata from multiple pages for SEO or metadata extraction.
  • Monitor website changes by crawling and comparing page content over time.

Crawlberg MCP server FAQ

What is Crawlberg?

Crawlberg is a web scraping and crawling library that converts websites into clean Markdown and structured metadata. It handles JavaScript rendering, bot detection, rate limiting, and SSRF protection out of the box.

Is Crawlberg free?

Yes, Crawlberg is open-source under the MIT License. There are also commercial offerings (Xberg Pro and Xberg Enterprise) for managed extras like proxy pools, authenticated sessions, and scheduling.

How do I use Crawlberg with Claude or Cursor?

Install the Crawlberg plugin from the marketplace in Claude Code, Cursor, or other AI coding assistants. It ships the MCP server and agent skills for site crawling and HTML-to-Markdown conversion.

What languages does Crawlberg support?

Crawlberg has bindings for Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, Kotlin (Android), and WebAssembly, plus a CLI and REST API.

Does Crawlberg require authentication?

No authentication is required to use Crawlberg. It supports optional HTTP Basic, Bearer token, and custom-header auth for crawling protected resources.

How does Crawlberg handle JavaScript-heavy sites?

Crawlberg falls back to a real headless browser for JavaScript-heavy single-page applications and includes WAF detection and bot-evasion strategies.

README (reference)

Source of truth, from the repository.

<p align="center"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://cdn.jsdelivr.net/gh/xberg-io/assets@v1/banner/readme-banner-dark.svg"> <img alt="Xberg" width="420" src="https://cdn.jsdelivr.net/gh/xberg-io/assets@v1/banner/readme-banner-light.svg"> </picture> </p>

Crawlberg

<div align="center" style="display: flex; flex-wrap: wrap; gap: 8px; justify-content: center; margin: 20px 0;"> <a href="https://github.com/xberg-io/alef"> <img src="https://img.shields.io/badge/built%20with-alef%20%D7%90-007ec6" alt="Built with alef"> </a> <!-- Language Bindings --> <a href="https://crates.io/crates/crawlberg"> <img src="https://img.shields.io/crates/v/crawlberg?label=Rust&color=007ec6" alt="Rust"> </a> <a href="https://pypi.org/project/crawlberg/"> <img src="https://img.shields.io/pypi/v/crawlberg?label=Python&color=007ec6" alt="Python"> </a> <a href="https://www.npmjs.com/package/@xberg-io/crawlberg"> <img src="https://img.shields.io/npm/v/@xberg-io/crawlberg?label=Node.js&color=007ec6" alt="Node.js"> </a> <a href="https://www.npmjs.com/package/@xberg-io/crawlberg-wasm"> <img src="https://img.shields.io/npm/v/@xberg-io/crawlberg-wasm?label=WASM&color=007ec6" alt="WASM"> </a> <a href="https://central.sonatype.com/artifact/io.xberg.crawlberg/crawlberg"> <img src="https://img.shields.io/maven-central/v/io.xberg.crawlberg/crawlberg?label=Java&color=007ec6" alt="Java"> </a> <a href="https://pkg.go.dev/github.com/xberg-io/crawlberg/packages/go"> <img src="https://img.shields.io/github/v/tag/xberg-io/crawlberg?label=Go&color=007ec6" alt="Go"> </a> <a href="https://www.nuget.org/packages/XbergIo.Crawlberg/"> <img src="https://img.shields.io/nuget/v/XbergIo.Crawlberg?label=C%23&color=007ec6" alt="C#"> </a> <a href="https://packagist.org/packages/xberg-io/crawlberg"> <img src="https://img.shields.io/packagist/v/xberg-io/crawlberg?label=PHP&color=007ec6" alt="PHP"> </a> <a href="https://rubygems.org/gems/crawlberg"> <img src="https://img.shields.io/gem/v/crawlberg?label=Ruby&color=007ec6" alt="Ruby"> </a> <a href="https://hex.pm/packages/crawlberg"> <img src="https://img.shields.io/hexpm/v/crawlberg?label=Elixir&color=007ec6" alt="Elixir"> </a> <a href="https://pub.dev/packages/crawlberg"> <img src="https://img.shields.io/pub/v/crawlberg?label=Dart&color=007ec6" alt="Dart"> </a> <a href="https://central.sonatype.com/artifact/io.xberg.crawlberg.android/crawlberg-android"> <img src="https://img.shields.io/maven-central/v/io.xberg.crawlberg.android/crawlberg-android?label=Kotlin&color=007ec6" alt="Kotlin"> </a> <a href="https://github.com/xberg-io/crawlberg/tree/main/packages/swift"> <img src="https://img.shields.io/badge/Swift-SPM-007ec6" alt="Swift"> </a> <a href="https://github.com/xberg-io/crawlberg/tree/main/packages/zig"> <img src="https://img.shields.io/badge/Zig-package-007ec6" alt="Zig"> </a> <a href="https://github.com/xberg-io/crawlberg/releases"> <img src="https://img.shields.io/badge/C-FFI-007ec6" alt="C FFI"> </a> <a href="https://github.com/xberg-io/crawlberg/pkgs/container/crawlberg"> <img src="https://img.shields.io/badge/Docker-ghcr.io-007ec6?logo=docker&logoColor=white" alt="Docker"> </a> <!-- Project Info --> <a href="https://github.com/xberg-io/crawlberg/blob/main/LICENSE"> <img src="https://img.shields.io/badge/License-MIT-007ec6" alt="License"> </a> <a href="https://docs.crawlberg.xberg.io"> <img src="https://img.shields.io/badge/Docs-crawlberg-007ec6" alt="Documentation"> </a> </div> <div align="center" style="display: flex; flex-wrap: wrap; gap: 12px; justify-content: center; margin: 28px 0 24px;"> <a href="https://discord.gg/xt9WY3GnKR"> <img height="22" src="https://img.shields.io/badge/Discord-Chat-007ec6?logo=discord&logoColor=white" alt="Join Discord"> </a> </div>

Turn any website into clean, structured data. Point Crawlberg at a URL and get back Markdown, metadata, and links — from a single page or a whole site — in the language you already use.

What and Why?

You need data that lives on the web, and raw HTML is not it. Crawlberg does the crawling, scraping, and cleanup end-to-end: it fetches pages, follows links, converts each one to Markdown, and hands you structured metadata (titles, links, images, social-card and JSON-LD data) — so you skip the parsing and go straight to the content.

It runs from a single Rust core with identical results across 14 language bindings, and it handles the awkward parts for you: JavaScript-heavy pages fall back to a real headless browser, bot filters are detected and worked around, robots and sitemaps are respected, requests are throttled per domain, and requests to private or internal addresses are refused by default. Drive it from your code, an AI agent, a REST service, or the CLI.

Every part of the pipeline is a trait you can swap — the crawl frontier, rate limiter, storage, event stream, and content filters — so you can plug in your own behavior. Managed extras like proxy pools, tuned bot-evasion, authenticated sessions, scheduling, and billing live in xberg-enterprise.

Features

FeatureDescription
Structured extractionText, metadata, links, images, assets, JSON-LD, Open Graph, hreflang, favicons, headings, response headers
Markdown conversionClean Markdown with citations, document structure, and fit-content mode
Concurrent crawlingDepth-first, breadth-first, or best-first traversal with configurable depth, page limits, and concurrency
14 language bindingsRust, Python, Node.js, Ruby, Go, Java, Kotlin (Android), C#, PHP, Elixir, Dart, Swift, Zig, and WebAssembly
Smart filteringBM25 relevance scoring, URL include/exclude patterns, robots.txt compliance, sitemap discovery
Browser renderingOptional headless browser for JavaScript-heavy SPAs with WAF detection and bypass
Batch & streamingScrape or crawl hundreds of URLs concurrently; real-time crawl events via async streams
SSRF-safe by defaultRefuses loopback, private, link-local, and cloud-metadata addresses; opt out via env var or CrawlConfig
Auth & rate limitingHTTP Basic, Bearer, and custom-header auth with cookie jars; per-domain request throttling
MCP server & REST APIModel Context Protocol integration for AI agents plus an HTTP server with OpenAPI spec

Supported Platforms

Precompiled binaries for Linux (x86_64/aarch64), macOS (ARM64), and Windows (x64) across every binding. See the platform support reference for the full matrix.

<p align="center"><strong>⭐ Star this repo to show your support — it helps others discover Crawlberg.</strong></p>

Quick Start

Language Packages

<details open> <summary><strong>Python</strong></summary>
pip install crawlberg

See Python README for full documentation.

</details> <details> <summary><strong>Node.js</strong></summary>
npm install @xberg-io/crawlberg

See Node.js README for full documentation.

</details> <details> <summary><strong>Rust</strong></summary>
cargo add crawlberg

See Rust README for full documentation.

</details> <details> <summary><strong>Go</strong></summary>
go get github.com/xberg-io/crawlberg/packages/go

See Go README for full documentation.

</details> <details> <summary><strong>Java</strong></summary>

Available on Maven Central as io.xberg.crawlberg:crawlberg. See Java README for the dependency snippet and current version.

</details> <details> <summary><strong>C#</strong></summary>
dotnet add package XbergIo.Crawlberg

See C# README for full documentation.

</details> <details> <summary><strong>Ruby</strong></summary>
gem install crawlberg

See Ruby README for full documentation.

</details> <details> <summary><strong>PHP</strong></summary>
composer require xberg-io/crawlberg

See PHP README for full documentation.

</details> <details> <summary><strong>Elixir</strong></summary>

Add {:crawlberg, "~> 0.3"} to your mix.exs dependencies. See Elixir README for full documentation.

</details> <details> <summary><strong>Dart / Flutter</strong></summary>
dart pub add crawlberg

See Dart README for full documentation.

</details> <details> <summary><strong>Kotlin (Android)</strong></summary>

Available on Maven Central as io.xberg.crawlberg.android:crawlberg-android. See Kotlin README for the dependency snippet and current version.

</details> <details> <summary><strong>Swift</strong></summary>

Add via Swift Package Manager. See Swift README for full documentation.

</details> <details> <summary><strong>Zig</strong></summary>

See Zig README for installation and usage.

</details> <details> <summary><strong>WebAssembly</strong></summary>
npm install @xberg-io/crawlberg-wasm

See WebAssembly README for full documentation.

</details> <details> <summary><strong>C/C++ (FFI)</strong></summary>

C header + shared library from GitHub Releases. See FFI crate for full documentation.

</details> <details> <summary><strong>CLI</strong></summary>
cargo install crawlberg-cli
# Or install the prebuilt binary via cargo-binstall:
cargo binstall crawlberg-cli
brew install xberg-io/tap/crawlberg

See CLI README for full documentation.

</details>

AI Coding Assistants

Install the Crawlberg plugin from xberg-io/crawlberg. It ships the Crawlberg agent skills (site crawling, HTML→Markdown scraping, headless-Chrome fallback) plus the crawlberg MCP server, and works with every major coding agent — expand your harness below.

<details open> <summary><strong>Claude Code</strong></summary>
/plugin marketplace add xberg-io/crawlberg
/plugin install crawlberg@crawlberg
</details> <details> <summary><strong>Codex CLI</strong></summary>
/plugins add https://github.com/xberg-io/crawlberg

Then search for crawlberg and select Install Plugin.

</details> <details> <summary><strong>Cursor</strong></summary>

Settings → Plugins → Add from URL → https://github.com/xberg-io/crawlberg, then select crawlberg.

</details> <details> <summary><strong>Gemini CLI</strong></summary>
gemini extensions install https://github.com/xberg-io/crawlberg
</details> <details> <summary><strong>Factory Droid</strong></summary>
droid plugin marketplace add https://github.com/xberg-io/crawlberg
droid plugin install crawlberg@crawlberg
</details> <details> <summary><strong>GitHub Copilot CLI</strong></summary>
copilot plugin marketplace add https://github.com/xberg-io/crawlberg
copilot plugin install crawlberg@crawlberg
</details> <details> <summary><strong>opencode</strong></summary>

Add the package to opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@xberg-io/opencode-crawlberg"]
}
</details>

Documentation

Full guides, per-language API references, the substrate/operational model, antibot strategy, and observability live at docs.crawlberg.xberg.io.

Contributing

Contributions are welcome! See our Contributing Guide.

Part of Xberg.io

  • Xberg — the open-source content-intelligence engine: text, tables, and metadata from 101 formats (115 file extensions), with OCR, transcription, and code intelligence. MIT.
  • Xberg Pro — a complete self-hosted content-intelligence backend in a single container. Commercial.
  • Xberg Enterprise — the distributed, governed content-intelligence platform, scaled on Kubernetes with team governance and support. Commercial.
  • crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
  • html-to-markdown — fast, lossless HTML→Markdown engine.
  • liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
  • tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
  • alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

License

MIT License

Links

Related MCP servers

HIPAA compliance AI agent — scan, grade, SRA, and generate compliance docs.

View repository →

Real eyes and hands on the web for AI agents. MCP-first. Local-first. Vendor-agnostic.

6
TypeScript
MIT
View repository →

Feature-rich MCP for controlling Jira through LLM clients. Safe for corporate environments.

7
Python
MIT
View repository →
TOToken Enhancer logo

Token Enhancer

Maintained

Local proxy that strips web pages to clean text, reducing AI token costs by 86-99.6%.

68
Python
MIT
View repository →

Xenarch — x402 MCP server for AI agent payments. Non-custodial, USDC on Base L2.

1
TypeScript
MIT
View repository →

MCP server exposing live Helldivers 2 galactic war data.

2
TypeScript
MIT
View repository →