Kreuzcrawl MCP Server
io.github.kreuzberg-dev/kreuzcrawl
Turn any website into clean Markdown or JSON with intelligent crawling, scraping, and structured data extraction.
What is the Kreuzcrawl MCP server?
The Kreuzcrawl MCP server is a web crawling and scraping tool that converts websites into structured Markdown, JSON, and metadata. It fetches pages, follows links, extracts text and metadata (titles, links, images, JSON-LD, Open Graph), and handles JavaScript-heavy sites with optional headless browser rendering.
Kreuzcrawl lets you point an AI agent at any URL and get back clean, structured data instead of raw HTML. It handles the hard parts: JavaScript rendering, bot detection, rate limiting, robots.txt compliance, and SSRF safety. You can scrape single pages or crawl entire sites with configurable depth and concurrency, extracting text, metadata, links, images, and more.
How to install Kreuzcrawl
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
crawl— Crawl websites with configurable depth, breadth-first or depth-first traversal, and concurrent requestsscrape— Extract structured data from a single page: Markdown, metadata, links, images, JSON-LD, Open Graph tags, and headersextract_markdown— Convert HTML to clean Markdown with citations and document structureextract_metadata— Extract titles, links, images, social-card data, JSON-LD, hreflang, and faviconsbatch_crawl— Scrape or crawl hundreds of URLs concurrently with streaming results
Use cases
- Convert website content to Markdown for AI training or documentation
- Extract structured metadata and links from a site for indexing or analysis
- Crawl multi-page sites with depth limits and relevance filtering
- Scrape JavaScript-heavy single-page applications with headless browser fallback
- Build a web-to-data pipeline for content intelligence or research
Kreuzcrawl MCP server FAQ
Kreuzcrawl is an MCP server that crawls and scrapes websites, converting HTML to clean Markdown and extracting structured metadata. It handles JavaScript rendering, respects robots.txt, detects bot filters, and provides concurrent crawling with configurable depth and filtering.
Yes, Kreuzcrawl is open-source under the MIT License. The core library and MCP server are free to use. Managed extras like proxy pools, authenticated sessions, and scheduling are available in xberg-enterprise (commercial).
Install the Crawlberg plugin from the marketplace (xberg-io/crawlberg), which includes the Kreuzcrawl MCP server. In Cursor, go to Settings → Plugins → Add from URL and paste https://github.com/xberg-io/crawlberg. In Claude Code, use `/plugin marketplace add xberg-io/crawlberg`.
No authentication is required for basic crawling. The server supports optional HTTP Basic, Bearer, and custom-header authentication for protected sites, plus cookie jar management.
Kreuzcrawl outputs Markdown (with citations and structure), JSON (structured metadata), and raw extracted data including titles, links, images, JSON-LD, Open Graph tags, and response headers.
Yes, Kreuzcrawl includes optional headless browser rendering for JavaScript-heavy SPAs, with WAF detection and bypass strategies. It falls back to browser rendering when needed.
README (reference)
Source of truth, from the repository.
Crawlberg
<div align="center" style="display: flex; flex-wrap: wrap; gap: 8px; justify-content: center; margin: 20px 0;"> <a href="https://github.com/xberg-io/alef"> <img src="https://img.shields.io/badge/built%20with-alef%20%D7%90-007ec6" alt="Built with alef"> </a> <!-- Language Bindings --> <a href="https://crates.io/crates/crawlberg"> <img src="https://img.shields.io/crates/v/crawlberg?label=Rust&color=007ec6" alt="Rust"> </a> <a href="https://pypi.org/project/crawlberg/"> <img src="https://img.shields.io/pypi/v/crawlberg?label=Python&color=007ec6" alt="Python"> </a> <a href="https://www.npmjs.com/package/@xberg-io/crawlberg"> <img src="https://img.shields.io/npm/v/@xberg-io/crawlberg?label=Node.js&color=007ec6" alt="Node.js"> </a> <a href="https://www.npmjs.com/package/@xberg-io/crawlberg-wasm"> <img src="https://img.shields.io/npm/v/@xberg-io/crawlberg-wasm?label=WASM&color=007ec6" alt="WASM"> </a> <a href="https://central.sonatype.com/artifact/io.xberg.crawlberg/crawlberg"> <img src="https://img.shields.io/maven-central/v/io.xberg.crawlberg/crawlberg?label=Java&color=007ec6" alt="Java"> </a> <a href="https://pkg.go.dev/github.com/xberg-io/crawlberg/packages/go"> <img src="https://img.shields.io/github/v/tag/xberg-io/crawlberg?label=Go&color=007ec6" alt="Go"> </a> <a href="https://www.nuget.org/packages/XbergIo.Crawlberg/"> <img src="https://img.shields.io/nuget/v/XbergIo.Crawlberg?label=C%23&color=007ec6" alt="C#"> </a> <a href="https://packagist.org/packages/xberg-io/crawlberg"> <img src="https://img.shields.io/packagist/v/xberg-io/crawlberg?label=PHP&color=007ec6" alt="PHP"> </a> <a href="https://rubygems.org/gems/crawlberg"> <img src="https://img.shields.io/gem/v/crawlberg?label=Ruby&color=007ec6" alt="Ruby"> </a> <a href="https://hex.pm/packages/crawlberg"> <img src="https://img.shields.io/hexpm/v/crawlberg?label=Elixir&color=007ec6" alt="Elixir"> </a> <a href="https://pub.dev/packages/crawlberg"> <img src="https://img.shields.io/pub/v/crawlberg?label=Dart&color=007ec6" alt="Dart"> </a> <a href="https://central.sonatype.com/artifact/io.xberg.crawlberg.android/crawlberg-android"> <img src="https://img.shields.io/maven-central/v/io.xberg.crawlberg.android/crawlberg-android?label=Kotlin&color=007ec6" alt="Kotlin"> </a> <a href="https://github.com/xberg-io/crawlberg/tree/main/packages/swift"> <img src="https://img.shields.io/badge/Swift-SPM-007ec6" alt="Swift"> </a> <a href="https://github.com/xberg-io/crawlberg/tree/main/packages/zig"> <img src="https://img.shields.io/badge/Zig-package-007ec6" alt="Zig"> </a> <a href="https://github.com/xberg-io/crawlberg/releases"> <img src="https://img.shields.io/badge/C-FFI-007ec6" alt="C FFI"> </a> <a href="https://github.com/xberg-io/crawlberg/pkgs/container/crawlberg"> <img src="https://img.shields.io/badge/Docker-ghcr.io-007ec6?logo=docker&logoColor=white" alt="Docker"> </a> <!-- Project Info --> <a href="https://github.com/xberg-io/crawlberg/blob/main/LICENSE"> <img src="https://img.shields.io/badge/License-MIT-007ec6" alt="License"> </a> <a href="https://docs.crawlberg.xberg.io"> <img src="https://img.shields.io/badge/Docs-crawlberg-007ec6" alt="Documentation"> </a> </div> <div align="center" style="display: flex; flex-wrap: wrap; gap: 12px; justify-content: center; margin: 28px 0 24px;"> <a href="https://discord.gg/xt9WY3GnKR"> <img height="22" src="https://img.shields.io/badge/Discord-Chat-007ec6?logo=discord&logoColor=white" alt="Join Discord"> </a> </div>Turn any website into clean, structured data. Point Crawlberg at a URL and get back Markdown, metadata, and links — from a single page or a whole site — in the language you already use.
What and Why?
You need data that lives on the web, and raw HTML is not it. Crawlberg does the crawling, scraping, and cleanup end-to-end: it fetches pages, follows links, converts each one to Markdown, and hands you structured metadata (titles, links, images, social-card and JSON-LD data) — so you skip the parsing and go straight to the content.
It runs from a single Rust core with identical results across 14 language bindings, and it handles the awkward parts for you: JavaScript-heavy pages fall back to a real headless browser, bot filters are detected and worked around, robots and sitemaps are respected, requests are throttled per domain, and requests to private or internal addresses are refused by default. Drive it from your code, an AI agent, a REST service, or the CLI.
Every part of the pipeline is a trait you can swap — the crawl frontier, rate limiter, storage, event stream, and content filters — so you can plug in your own behavior. Managed extras like proxy pools, tuned bot-evasion, authenticated sessions, scheduling, and billing live in xberg-enterprise.
Features
| Feature | Description |
|---|---|
| Structured extraction | Text, metadata, links, images, assets, JSON-LD, Open Graph, hreflang, favicons, headings, response headers |
| Markdown conversion | Clean Markdown with citations, document structure, and fit-content mode |
| Concurrent crawling | Depth-first, breadth-first, or best-first traversal with configurable depth, page limits, and concurrency |
| 14 language bindings | Rust, Python, Node.js, Ruby, Go, Java, Kotlin (Android), C#, PHP, Elixir, Dart, Swift, Zig, and WebAssembly |
| Smart filtering | BM25 relevance scoring, URL include/exclude patterns, robots.txt compliance, sitemap discovery |
| Browser rendering | Optional headless browser for JavaScript-heavy SPAs with WAF detection and bypass |
| Batch & streaming | Scrape or crawl hundreds of URLs concurrently; real-time crawl events via async streams |
| SSRF-safe by default | Refuses loopback, private, link-local, and cloud-metadata addresses; opt out via env var or CrawlConfig |
| Auth & rate limiting | HTTP Basic, Bearer, and custom-header auth with cookie jars; per-domain request throttling |
| MCP server & REST API | Model Context Protocol integration for AI agents plus an HTTP server with OpenAPI spec |
Supported Platforms
Precompiled binaries for Linux (x86_64/aarch64), macOS (ARM64), and Windows (x64) across every binding. See the platform support reference for the full matrix.
<p align="center"><strong>⭐ Star this repo to show your support — it helps others discover Crawlberg.</strong></p>Quick Start
Language Packages
<details open> <summary><strong>Python</strong></summary>pip install crawlberg
See Python README for full documentation.
</details> <details> <summary><strong>Node.js</strong></summary>npm install @xberg-io/crawlberg
See Node.js README for full documentation.
</details> <details> <summary><strong>Rust</strong></summary>cargo add crawlberg
See Rust README for full documentation.
</details> <details> <summary><strong>Go</strong></summary>go get github.com/xberg-io/crawlberg/packages/go
See Go README for full documentation.
</details> <details> <summary><strong>Java</strong></summary>Available on Maven Central as io.xberg.crawlberg:crawlberg. See Java README for the dependency snippet and current version.
dotnet add package XbergIo.Crawlberg
See C# README for full documentation.
</details> <details> <summary><strong>Ruby</strong></summary>gem install crawlberg
See Ruby README for full documentation.
</details> <details> <summary><strong>PHP</strong></summary>composer require xberg-io/crawlberg
See PHP README for full documentation.
</details> <details> <summary><strong>Elixir</strong></summary>Add {:crawlberg, "~> 0.3"} to your mix.exs dependencies. See Elixir README for full documentation.
dart pub add crawlberg
See Dart README for full documentation.
</details> <details> <summary><strong>Kotlin (Android)</strong></summary>Available on Maven Central as io.xberg.crawlberg.android:crawlberg-android. See Kotlin README for the dependency snippet and current version.
Add via Swift Package Manager. See Swift README for full documentation.
</details> <details> <summary><strong>Zig</strong></summary>See Zig README for installation and usage.
</details> <details> <summary><strong>WebAssembly</strong></summary>npm install @xberg-io/crawlberg-wasm
See WebAssembly README for full documentation.
</details> <details> <summary><strong>C/C++ (FFI)</strong></summary>C header + shared library from GitHub Releases. See FFI crate for full documentation.
</details> <details> <summary><strong>CLI</strong></summary>cargo install crawlberg-cli
# Or install the prebuilt binary via cargo-binstall:
cargo binstall crawlberg-cli
brew install xberg-io/tap/crawlberg
See CLI README for full documentation.
</details>AI Coding Assistants
Install the Crawlberg plugin from xberg-io/crawlberg. It ships the Crawlberg agent skills (site crawling, HTML→Markdown scraping, headless-Chrome fallback) plus the crawlberg MCP server, and works with every major coding agent — expand your harness below.
/plugin marketplace add xberg-io/crawlberg
/plugin install crawlberg@crawlberg
</details>
<details>
<summary><strong>Codex CLI</strong></summary>
/plugins add https://github.com/xberg-io/crawlberg
Then search for crawlberg and select Install Plugin.
Settings → Plugins → Add from URL → https://github.com/xberg-io/crawlberg, then select crawlberg.
gemini extensions install https://github.com/xberg-io/crawlberg
</details>
<details>
<summary><strong>Factory Droid</strong></summary>
droid plugin marketplace add https://github.com/xberg-io/crawlberg
droid plugin install crawlberg@crawlberg
</details>
<details>
<summary><strong>GitHub Copilot CLI</strong></summary>
copilot plugin marketplace add https://github.com/xberg-io/crawlberg
copilot plugin install crawlberg@crawlberg
</details>
<details>
<summary><strong>opencode</strong></summary>
Add the package to opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["@xberg-io/opencode-crawlberg"]
}
</details>
Documentation
Full guides, per-language API references, the substrate/operational model, antibot strategy, and observability live at docs.crawlberg.xberg.io.
Contributing
Contributions are welcome! See our Contributing Guide.
Part of Xberg.io
- Xberg — the open-source content-intelligence engine: text, tables, and metadata from 101 formats (115 file extensions), with OCR, transcription, and code intelligence. MIT.
- Xberg Pro — a complete self-hosted content-intelligence backend in a single container. Commercial.
- Xberg Enterprise — the distributed, governed content-intelligence platform, scaled on Kubernetes with team governance and support. Commercial.
- crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- html-to-markdown — fast, lossless HTML→Markdown engine.
- liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
- tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
- alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.
License
Links
- Documentation
- GitHub Repository
- Issue Tracker
- Changelog
- Discord — community, roadmap, announcements.
Related MCP servers
Security-hardened file organizer with smart categorization and duplicate detection

io.github.krimto-labs/krimto
Memory for AI coding agents — markdown in your own git, not a vendor database. Apache-2.0.
AI visibility checker: GPTBot vs OAI-SearchBot, robots.txt audit, raw-HTML schema. 165 crawlers.
View repository →
Indian-accurate nutrition logging for your AI: IFCT 2017 + USDA, by text or photo.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA.

