PluginBench
MCP Server
Active
MIT

haiku.rag MCP Server

io.github.ggozad/haiku-rag

Local-first agentic RAG with hybrid search, reranking, and multimodal document retrieval with citations.

What is the haiku.rag MCP server?

The haiku.rag MCP server is an agentic retrieval-augmented generation (RAG) system that answers questions about your own documents with citations to page numbers and section headings. It runs locally on an embedded LanceDB database with no server required, and supports hybrid search, multimodal embeddings, vision QA, reranking, and code-based analysis.

haiku.rag lets you index and search your own documents locally, then ask questions and get answers with precise citations. It combines vector and full-text search via Reciprocal Rank Fusion, supports multimodal embeddings and vision-capable QA models, includes optional reranking and evidence compaction, and exposes all capabilities as MCP tools for use with Claude Desktop and other AI assistants.

How to install haiku.rag

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "haiku-rag": {
      "command": "uvx",
      "args": [
        "haiku-rag",
        "mcp",
        "--stdio"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • add-src — Index documents from local files, HTTP URLs, S3, or WebDAV sources
  • search — Hybrid vector + full-text search with Reciprocal Rank Fusion, returns chunks with page numbers and section headings
  • ask — Question answering with citations; supports attaching images for vision-capable models
  • analyze — Complex analytical tasks via sandboxed Python code execution for aggregation, computation, and multi-document analysis
  • chat — Multi-turn conversational RAG with session memory
  • tag — Name and manage database states for rollback and versioning

Use cases

  • Index PDFs and web documents, then ask questions with precise citations to page numbers and sections
  • Search documents using both semantic similarity and keyword matching simultaneously
  • Analyze multi-document datasets with Python code execution (e.g., count mentions, aggregate metrics)
  • Ask questions about embedded figures and images using vision-capable models
  • Set up continuous document ingestion from file systems, HTTP endpoints, S3, or WebDAV sources

haiku.rag MCP server FAQ

What is haiku.rag?

haiku.rag is a local-first RAG system that indexes your documents and answers questions with citations. It combines vector search, full-text search, and optional reranking, and runs entirely on your machine using an embedded LanceDB database.

Is haiku.rag free?

Yes, haiku.rag is open-source under the MIT License. You can use it for free, though you may need to provide your own embedding model (via Ollama, OpenAI, VoyageAI, Cohere, or vLLM) and LLM for QA.

How do I use haiku.rag as an MCP server in Claude Desktop?

Install haiku-rag via pip, then add it to your Claude Desktop configuration by setting the command to 'haiku-rag' with args ['mcp', '--stdio']. This exposes document management, search, QA, and analysis tools to Claude.

What embedding and LLM providers does haiku.rag support?

Embeddings: Ollama, OpenAI, VoyageAI, Cohere, LM Studio, and vLLM (with multimodal support on vLLM, VoyageAI, and Cohere). QA: any model supported by Pydantic AI, including local models via Ollama or vLLM.

Can haiku.rag process images and figures?

Yes. It captures embedded figures during document ingestion and supports vision-capable QA models that receive figure bytes alongside text. You can also attach images to questions directly via the ask command or MCP tools.

Does haiku.rag require a server?

No. It runs entirely locally with an embedded LanceDB database. It also optionally supports remote storage (S3, GCS, Azure, LanceDB Cloud) and includes a production ingester service for continuous indexing.

README (reference)

Source of truth, from the repository.

haiku.rag

PyPI Python Downloads Docs Tests codecov

Agentic RAG that answers questions about your own documents with citations to page numbers and section headings. Runs locally on an embedded database, no server required.

Built on LanceDB, Pydantic AI, and Docling. Full documentation at ggozad.github.io/haiku.rag.

New: vision and multimodal search. Picture-aware ingestion captures embedded figure bytes; vision-capable QA models receive them alongside text. Multimodal embedders put picture vectors in the same space as text, enabling text-as-query → figure hits and image-as-query retrieval.

Features

  • Hybrid search — Vector + full-text with Reciprocal Rank Fusion
  • Multimodal & cross-modal search — Multimodal embedders (vLLM, VoyageAI, Cohere) put picture vectors in the same space as text; supports text-as-query → figure hits and image-as-query
  • Question answering — RAG capability with citations (page numbers, section headings)
  • Vision QA — Vision-capable models receive figure bytes alongside chunk text; attach your own images to questions in ask, analyze, MCP, and the chat TUI
  • Reranking — local cross-encoders, Cohere, Zero Entropy, or vLLM
  • Analysis capability — Complex analytical tasks via sandboxed Python code execution (aggregation, computation, multi-document analysis)
  • Evidence compaction — Optional capability that replaces earlier questions' search results on the request with the evidence they cited, so long conversations stop resending everything they retrieved
  • Citation policy — Optional capability that requires every answer to declare what grounds it, including declaring that nothing does
  • Conversational RAG — Chat TUI and web application for multi-turn conversations with session memory
  • Document structure — Stores full DoclingDocument, enabling structure-aware context expansion
  • Multiple providers — Embeddings: Ollama, OpenAI, VoyageAI, Cohere, LM Studio, vLLM (multimodal via multimodal: true on vLLM/VoyageAI/Cohere). QA: any model supported by Pydantic AI
  • Local-first — Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud
  • CLI & Python API — Full functionality from command line or code
  • MCP server — Expose as tools for AI assistants (Claude Desktop, etc.)
  • Visual grounding — View chunks highlighted on original page images
  • Production ingester — Long-lived haiku-ingester service with persistent SQLite queue, async worker pool with retries and a dead-letter queue, FS / HTTP / S3 / WebDAV source adapters, FastAPI control plane, and a browser dashboard for operators. See docs/ingester.md.
  • Tags — Name database states with haiku-rag tag and roll back to them
  • Inspector — TUI for browsing documents, chunks, and search results

Installation

Python 3.12 or newer required

Full Package (Recommended)

pip install haiku.rag

Includes all features: document processing, all embedding providers, and rerankers.

Using uv? uv pip install haiku.rag

Slim Package (Minimal Dependencies)

pip install haiku.rag-slim

Install only the extras you need. See the Installation documentation for available options.

Quick Start

Note: Requires an embedding provider (Ollama, OpenAI, etc.). See the Tutorial for setup instructions.

# Index a PDF
haiku-rag add-src paper.pdf

# Search
haiku-rag search "attention mechanism"

# Ask questions with citations
haiku-rag ask "What datasets were used for evaluation?"

# Ask about an image (vision-capable model)
haiku-rag ask "Does this figure match the spec in the design doc?" --image figure.png

# Analyze — complex analytical tasks via code execution
haiku-rag analyze "How many documents mention transformers?"

# Interactive chat — multi-turn conversations with memory
haiku-rag chat

# Continuously ingest from configured sources (FS, HTTP, S3, WebDAV)
haiku-ingester serve

See Configuration for customization options.

Python API

from haiku.rag.client import HaikuRAG

async with HaikuRAG("knowledge.lancedb", create=True) as rag:
    # Index documents
    await rag.create_document_from_source("paper.pdf")
    await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")

    # Search — returns chunks with provenance
    results = await rag.search("self-attention")
    for result in results:
        print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")

    # QA with citations
    answer, citations = await rag.ask("What is the complexity of self-attention?")
    print(answer)
    for cite in citations:
        print(f"  [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")

For direct agent composition, see the capabilities documentation.

MCP Server

Use with AI assistants like Claude Desktop:

haiku-rag mcp --stdio

Add to your Claude Desktop configuration:

{
  "mcpServers": {
    "haiku-rag": {
      "command": "haiku-rag",
      "args": ["mcp", "--stdio"]
    }
  }
}

Provides tools for document management, search, QA, and analysis directly in your AI assistant.

Examples

See the examples directory for working examples:

  • Docker Setup - Complete Docker deployment with continuous ingestion (haiku-ingester) and MCP server
  • Web Application - Full-stack conversational RAG with CopilotKit frontend

Documentation

Full documentation at: https://ggozad.github.io/haiku.rag/

License

This project is licensed under the MIT License.

<!-- mcp-name is used by the MCP registry to identify this server -->

mcp-name: io.github.ggozad/haiku-rag

Related MCP servers

Deterministic Korean Saju / BaZi Four Pillars MCP. Day Master, five elements, compatibility.

Deterministic Pythagorean numerology MCP. Life Path, Destiny, Soul Urge, compatibility. No API key.

Korean Four Pillars of Destiny (Saju/Bazi): calculate, interpret, compatibility & daily fortune.

0
TypeScript
View repository →

Give any LLM agent a real Android or iPhone as its body with 62 MCP tools for mobile automation.

314
Python
MIT
View repository →
GHGhostchars logo

Find and remove invisible Unicode: zero-width, tag smuggling, bidi, homoglyphs. Offline.

0
TypeScript
View repository →

Minimal MCP server for Ghost Security API - compatible with all MCP clients