io.github.shinpr/mcp-local-rag MCP Server
io.github.shinpr/mcp-local-rag
Local RAG server for searching private documents without sending them to APIs
What is the io.github.shinpr/mcp-local-rag MCP server?
MCP Local RAG is a local retrieval-augmented generation server that indexes and searches PDF, DOCX, Markdown, and text files on your machine using semantic embeddings and keyword matching. It runs entirely offline after initial setup, with no API keys or external services required.
mcp-local-rag lets you search private documents using hybrid semantic and keyword search without uploading them to cloud services. It indexes files locally, combines semantic similarity with exact-term matching for technical documentation, and works both as an MCP server for AI coding tools and as a CLI utility. Perfect for keeping confidential documents searchable while maintaining complete privacy.
How to install io.github.shinpr/mcp-local-rag
Copy-paste configuration for popular MCP clients.
BASE_DIRBase directory for document storage (defaults to current working directory). Ignored when BASE_DIRS is set.
BASE_DIRSJSON array of base directories (e.g. '["/a","/b"]'). Takes precedence over BASE_DIR.
DB_PATHPath to LanceDB database directory (defaults to ./lancedb/)
CACHE_DIRDirectory where Transformers.js models are cached (defaults to ./models/)
HF_ENDPOINTHugging Face model download endpoint. Set this to a mirror URL when direct downloads are blocked (defaults to https://huggingface.co).
MODEL_NAMEEmbedding model name (defaults to Xenova/all-MiniLM-L6-v2)
MAX_FILE_SIZEMaximum file size in bytes (defaults to 104857600 / 100MB)
RAG_MAX_DISTANCEMaximum distance threshold for filtering search results. Results with distance greater than this value will be excluded. Lower values mean stricter filtering (e.g., 0.5 for high relevance only)
RAG_GROUPINGGrouping mode for quality filtering. 'similar' returns only the most similar group (stops at first distance jump). 'related' includes related groups (stops at second distance jump). Unset means no grouping filter
RAG_MAX_FILESMaximum number of files to keep in search results. Results are filtered to include only chunks from the top N best-scoring files. For example, 1 returns only the single best-matching file's chunks. Unset means no file filtering.
CHUNK_MIN_LENGTHMinimum chunk length in characters (1-10000, defaults to 50). Chunks shorter than this threshold are filtered out during ingestion.
STORE_IMAGESStore supported PDF and DOCX images during ingestion and return them with matched chunks (defaults to false).
RAG_DEVICEExecution device for the embedder (defaults to cpu). Passed straight to ONNX Runtime; see the Transformers.js device source for the supported backend names. If the requested device fails to initialize, the server throws an error.
RAG_DTYPEEmbedding quantization dtype for the embedder (defaults to fp32). Opt-in and pass-through; accepts any dtype the chosen model provides (fp32, fp16, q8, int8, ...). If the model has no variant for the requested dtype, the server throws an error. Changing this changes the embedding space — re-ingest existing data.
RAG_HYBRID_WEIGHTKeyword boost factor for hybrid search (0.0-1.0, defaults to 0.6). 0 means semantic similarity only; higher values increase the keyword-match contribution to the final score.
RAG_RERANK_CMDExternal reranker command template. Use {query} and {top} for query text and result count; unset disables reranking.
RAG_RERANK_TIMEOUT_MSTime budget per rerank call in milliseconds (100-600000, defaults to 10000). On timeout the spawned command is killed and the pre-rerank ordering is returned.
Tools & capabilities
Tools this server exposes to the agent.
sync_start— Reconcile the index with all configured roots or one pathsync_status— Poll a running sync jobingest_file— Ingest or replace one file (PDF, DOCX, TXT, Markdown)ingest_data— Ingest text, Markdown, or HTML already held by the clientquery_documents— Search with semantic matching and keyword boostread_chunk_neighbors— Read surrounding chunks from a search resultlist_files— Show supported files and their ingestion statedelete_file— Delete an indexed file or an ingest_data itemstatus— Show index and search status
Use cases
- Search API documentation for authentication details and error codes without uploading to external services
- Index confidential internal documentation and search it semantically while keeping it on your machine
- Find exact technical terms (API names, class names, error codes) in large document sets using hybrid keyword and semantic search
- Ingest HTML pages fetched by your AI assistant and search them alongside local documents
- Retrieve surrounding context from search results to get complete answers from technical documentation
io.github.shinpr/mcp-local-rag MCP server FAQ
MCP Local RAG is a local retrieval-augmented generation server that indexes and searches documents (PDF, DOCX, Markdown, text) using semantic embeddings and keyword matching, all running on your machine without sending data to external APIs.
Yes, MCP Local RAG is open source under the MIT license with no API costs. You only need Node.js 22+ and internet access on first use to download the embedding model.
Add to ~/.cursor/mcp.json: {"mcpServers": {"local-rag": {"command": "npx", "args": ["-y", "mcp-local-rag"], "env": {"BASE_DIR": "/absolute/path/to/your/documents"}}}}. Then restart Cursor and ask it to sync documents.
Run: claude mcp add local-rag --scope user --env BASE_DIR=/absolute/path/to/your/documents -- npx -y mcp-local-rag
No, MCP Local RAG requires no API keys, Docker, Python, or external database. It runs entirely locally after downloading the embedding model on first use.
Supported formats for file ingestion are PDF, DOCX, TXT, and Markdown. HTML can be ingested if fetched by the client. Excel, PowerPoint, and source-code files are not supported.
README (reference)
Source of truth, from the repository.
MCP Local RAG
<p align="center"> <strong>English</strong> | <a href="README.zh-CN.md">简体中文</a> | <a href="README.de.md">Deutsch</a> | <a href="README.es.md">Español</a> | <a href="README.pt-BR.md">Português (Brasil)</a> | <a href="README.fr.md">Français</a> </p>Search private documents from an MCP client or the terminal without sending them to an embedding API.
mcp-local-rag indexes PDF, DOCX, Markdown, and text files on your machine. Search combines semantic similarity with keyword matching, so queries can match both intent and exact technical terms such as API names, class names, and error codes.
Features
- Runs locally: Document parsing, embeddings, storage, and search run on your machine. After the initial model download, text ingestion and search work offline.
- Hybrid search: Semantic retrieval finds related concepts, while keyword matching boosts exact technical terms.
- Configurable embeddings: Choose a Hugging Face embedding model that fits the language and domain of your documents.
- Semantic chunking: Documents are split at topic boundaries instead of fixed character counts. Markdown code blocks stay intact.
- MCP and CLI: Use the same index from an AI coding tool or directly from the terminal.
No API key, Docker, Python, or external database is required.
Quick Start
Requirements
- Node.js 22 or later
- Internet access on first use to download the npm package and embedding model
- A directory containing the documents you want to search
Set BASE_DIR to that directory. It is also the security boundary for file operations. Replace
/absolute/path/to/your/documents below with the directory's absolute path.
mcp-local-rag uses the standard MCP protocol over a local stdio server, so it works with AI coding tools and other MCP hosts that support local MCP servers.
Use one of the examples below, or register npx -y mcp-local-rag and set BASE_DIR using your
client's MCP configuration format.
For Claude Code: Run this command:
claude mcp add local-rag --scope user --env BASE_DIR=/absolute/path/to/your/documents -- npx -y mcp-local-rag
For Codex: Add to ~/.codex/config.toml:
[mcp_servers.local-rag]
command = "npx"
args = ["-y", "mcp-local-rag"]
[mcp_servers.local-rag.env]
BASE_DIR = "/absolute/path/to/your/documents"
For OpenCode: Add to ~/.config/opencode/opencode.json (or opencode.jsonc):
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"local-rag": {
"type": "local",
"command": ["npx", "-y", "mcp-local-rag"],
"environment": {
"BASE_DIR": "/absolute/path/to/your/documents"
}
}
}
}
For Cursor: Add to ~/.cursor/mcp.json:
{
"mcpServers": {
"local-rag": {
"command": "npx",
"args": ["-y", "mcp-local-rag"],
"env": {
"BASE_DIR": "/absolute/path/to/your/documents"
}
}
}
}
Restart the client, then ask it to build the index:
Sync all documents in the configured root and wait until it finishes.
The first sync downloads the default embedding model (about 90 MB) and may take 1–2 minutes before ingestion starts. Later runs use the local cache.
Once the sync completes:
What does the API documentation say about authentication?
CLI Quick Start
To use the CLI without an MCP client:
npx mcp-local-rag ingest ./docs/
npx mcp-local-rag query "authentication API"
The CLI uses the current directory as its document root by default. Run both commands from the
same directory so they use the same default index, or set BASE_DIR and DB_PATH explicitly.
Why This Exists
Some document sets cannot be sent to a hosted embedding service because of confidentiality or organizational policy. Keeping the index local makes them searchable without adding a per-query API cost.
Semantic search alone can miss exact identifiers that matter in technical documentation. Keyword reranking keeps those terms visible without giving up natural-language retrieval.
Supported Content
| Input | How to ingest |
|---|---|
| PDF, DOCX, TXT, Markdown | File ingestion or directory sync |
| HTML already fetched by the client | ingest_data; cleaned with Readability and converted to Markdown |
| Plain text or Markdown held in memory | ingest_data with a stable source identifier |
HTML fetching is not built into the server. An MCP client can fetch a page and pass its HTML to
ingest_data.
Excel, PowerPoint, standalone images, and source-code file extensions are not supported by file ingestion. PDFs can optionally use a local vision model to describe figures, but this is not OCR or image search.
MCP Tools
| Tool | Purpose |
|---|---|
sync_start | Reconcile the index with all configured roots or one path |
sync_status | Poll a running sync job |
ingest_file | Ingest or replace one file |
ingest_data | Ingest text, Markdown, or HTML already held by the client |
query_documents | Search with semantic matching and keyword boost |
read_chunk_neighbors | Read surrounding chunks from a search result |
list_files | Show supported files and their ingestion state |
delete_file | Delete an indexed file or an ingest_data item |
status | Show index and search status |
Syncing a Document Root
sync_start ingests new and changed files, skips byte-identical files, and removes index entries
for files that no longer exist:
Sync everything under the configured document roots and wait for completion.
The tool returns a jobId immediately. Clients should poll sync_status until its state becomes
succeeded or failed. There is no visual mode during sync; changed PDFs are ingested as text.
Only one sync job is retained by the server process. A newer job replaces a finished record, and restarting the server discards it.
Ingesting One File
ingest_file accepts PDF, DOCX, TXT, and Markdown. MCP file paths must be absolute and must stay
inside a configured document root:
Ingest the document at /Users/me/docs/api-spec.pdf.
Re-ingesting the same path replaces its existing chunks.
Searching and Reading More Context
What does the API documentation say about authentication?
Find the documented behavior of ERR_CONNECTION_REFUSED.
Results contain the text, source path, title, chunk index, and relevance score. Pass the
chunkIndex and either filePath or source from a result to read_chunk_neighbors when the
answer needs more context:
Read the surrounding chunks for that authentication result.
Both query_documents and list_files accept an optional absolute scope path prefix, or a
list of prefixes. A prefix matches the exact path and its descendants.
Ingesting HTML
Use ingest_data after the MCP client fetches a page:
Fetch https://example.com/docs and ingest the HTML.
The server extracts the main article, converts it to Markdown, and stores it under the supplied source identifier. Reusing the same source updates the existing content.
Respect the source site's terms and copyright when indexing external content.
PDF Figures
Visual mode adds a generated caption for figure-heavy PDF pages. It is opt-in and does not load a vision model during normal ingestion.
Ingest /Users/me/docs/research-paper.pdf with visual: true.
npx mcp-local-rag ingest ./docs/research-paper.pdf --visual
| Profile | Model cache | Use case |
|---|---|---|
fast (default) | about 250 MB | Lightweight visual indexing |
quality | about 2.9 GB | Figures containing labels, annotations, or other in-image text |
Select the larger model with visualQuality: "quality" over MCP or
--visual-quality quality over CLI. Measured CPU inference was about twice as slow as fast,
though results depend on hardware and model updates.
Captions are auxiliary text, not faithful transcriptions. Treat retrieved captions and document text as untrusted input rather than instructions.
CLI
The CLI uses the same parser, embedder, and vector store without an MCP client:
npx mcp-local-rag ingest ./docs/
npx mcp-local-rag sync ./docs/
npx mcp-local-rag query "authentication API"
npx mcp-local-rag query "auth" --scope /docs/api --scope /docs/guide
npx mcp-local-rag read-neighbors --file-path /abs/path.md --chunk-index 5
npx mcp-local-rag list
npx mcp-local-rag status
npx mcp-local-rag delete ./docs/old.pdf
npx mcp-local-rag delete --source "https://example.com/docs"
Global options such as --db-path, --cache-dir, and --model-name go before the subcommand.
Subcommand options go after it:
npx mcp-local-rag --db-path ./my-db query "authentication"
Run npx mcp-local-rag --help for the complete command reference.
The CLI does not read MCP client configuration. Set the same environment variables or flags if
both interfaces should share an index. In particular, MODEL_NAME and the CLI --model-name
must match for a shared database.
Search Tuning
Keyword boost is enabled by default. Relevance-gap grouping and the distance and file filters are optional controls for corpora that need tighter result selection.
| Variable | Default | Description |
|---|---|---|
RAG_HYBRID_WEIGHT | 0.6 | Keyword boost factor (0.0–1.0). 0 disables keyword reranking; 1 applies the maximum boost. |
RAG_GROUPING | (not set) | similar keeps the first relevance group; related keeps up to two, using significant vector-distance gaps as boundaries. |
RAG_MAX_DISTANCE | (not set) | Filter out low-relevance results (e.g., 0.5). |
RAG_MAX_FILES | (not set) | Limit results to top N files (e.g., 1 for single best file). |
For API specifications and other documents containing many identifiers, a stronger keyword weight can improve exact-term ranking:
"env": {
"RAG_HYBRID_WEIGHT": "0.7"
}
0.7: slightly stronger exact-term reranking than the default1.0: maximum keyword boost
How It Works
During ingestion:
- The parser extracts text for the input format.
- The semantic chunker finds topic boundaries and preserves Markdown code blocks.
- Transformers.js creates embeddings locally.
- LanceDB stores the chunks, metadata, vectors, and full-text index.
During search:
- The query is embedded with the same model.
- Vector search retrieves semantically related chunks.
- Optional distance and relevance-group filters narrow the candidates when configured.
- Full-text matches boost exact query terms.
Agent Skills
Agent Skills provide query and ingestion guidance for AI assistants:
npx mcp-local-rag skills install --claude-code
npx mcp-local-rag skills install --claude-code --global
npx mcp-local-rag skills install --codex
Installed skills cover query formulation, result refinement, and HTML ingestion. Ask the assistant to use the mcp-local-rag skill explicitly if it does not activate automatically.
Configuration
The MCP server reads environment variables. The CLI accepts the same variables plus the listed flags, with CLI flags taking precedence.
| Environment Variable | CLI Flag | Default | Description |
|---|---|---|---|
BASE_DIR | --base-dir | Current directory | One document root; the CLI flag is repeatable on ingest, list, and sync |
BASE_DIRS | N/A | (unset) | JSON array of document roots; takes precedence over BASE_DIR |
DB_PATH | --db-path | ./lancedb/ | Vector database location |
CACHE_DIR | --cache-dir | ./models/ | Model cache directory |
MODEL_NAME | --model-name | Xenova/all-MiniLM-L6-v2 | Hugging Face embedding model |
MAX_FILE_SIZE | --max-file-size | 104857600 (100MB) | Maximum file size in bytes |
CHUNK_MIN_LENGTH | --chunk-min-length | 50 | Minimum chunk length in characters (1–10000) |
RAG_DEVICE | N/A | cpu | ONNX Runtime execution device |
RAG_DTYPE | N/A | fp32 | Embedding dtype passed to the selected model |
Document Roots (BASE_DIR and BASE_DIRS)
mcp-local-rag only allows file operations inside configured roots. For multiple roots,
BASE_DIRS must be a JSON array of non-empty paths:
export BASE_DIRS='["/Users/me/Documents/work","/Users/me/Projects/specs"]'
Root configuration is resolved in this order:
- CLI
--base-dir <path>flags (repeatable oningest,list, andsync) BASE_DIRSBASE_DIR- Current directory
Each source replaces the lower-priority source rather than merging with it. Invalid BASE_DIRS
configuration fails instead of falling back to BASE_DIR or the current directory. status
remains available in MCP so the client can report the configuration error.
npx mcp-local-rag ingest --base-dir /Users/me/work --base-dir /Users/me/specs /Users/me/work/readme.md
npx mcp-local-rag list --base-dir /Users/me/work --base-dir /Users/me/specs
npx mcp-local-rag sync --base-dir /Users/me/work --base-dir /Users/me/specs
BASE_DIRS='["/Users/me/work","/Users/me/specs"]' npx mcp-local-rag list
Storage and Models
DB_PATH and CACHE_DIR are relative to the process working directory by default. Set absolute
paths when the MCP client may start the server from different project directories.
Set MODEL_NAME or pass --model-name to choose a Hugging Face embedding model that fits the
language and domain of your documents.
mcp-local-rag generates embeddings with mean pooling and L2 normalization. When choosing a model, check whether these settings match its recommended inference setup, since the pooling method can affect retrieval quality.
Changing MODEL_NAME, RAG_DEVICE, or RAG_DTYPE can make existing vectors incompatible.
Use a new DB_PATH or delete the existing index and re-ingest after changing the embedding
configuration.
An example model for English documents is Xenova/bge-small-en-v1.5.
Security and Operation
- File access is restricted to
BASE_DIR,BASE_DIRS, or CLI--base-dirroots. - Symlinks that resolve outside every configured root are rejected.
- Document processing and search make no network requests after the required models are cached.
- The server is designed for one local user and does not provide authentication or access control.
- Do not run multiple CLI or MCP writers against the same
DB_PATH. Read-only queries can run while a sync is active. - Back up an index by copying its
DB_PATHdirectory while no writer is active.
"No results found"
Documents must be ingested first. Run "List all ingested files" to verify.
Model download failed
Check internet connection. If behind a proxy, configure network settings. The model can also be downloaded manually.
"File too large"
Default limit is 100MB. Split large files or increase MAX_FILE_SIZE.
Slow queries
Check chunk count with status. Large documents with many chunks may slow queries. Consider splitting very large files.
"Path outside BASE_DIR"
Ensure file paths are within one of the configured roots (BASE_DIR, any BASE_DIRS entry, or any CLI --base-dir). Use absolute paths.
"BASE_DIRS must be a JSON array..."
BASE_DIRS accepts a JSON array of one or more non-empty path strings:
- Valid:
BASE_DIRS='["/Users/me/work","/Users/me/specs"]' - Invalid:
BASE_DIRS=/a:/b(delimiter syntax not supported) - Invalid:
BASE_DIRS='[]'(empty array)
MCP client doesn't see tools
- Verify config file syntax
- Restart client completely (Cmd+Q on Mac for Cursor)
- Test directly:
npx mcp-local-ragshould run without errors
Contributing
Contributions welcome! See CONTRIBUTING.md for setup and guidelines.
License
MIT License. Free for personal and commercial use.
Blog Posts
- Building a Local RAG for Agentic Coding: Technical deep-dive into the semantic chunking and hybrid search design.
Acknowledgments
Built with Model Context Protocol by Anthropic, LanceDB, and Transformers.js.
Related MCP servers

Run task-specific AI sub-agents across Cursor, Claude, Codex, Gemini, and other tools via MCP.

Generate and edit images with AI-powered prompt optimization across Gemini, OpenAI, and BytePlus providers.
Local-first codebase intelligence: cited answers, audits, reports. EN/FR.
Track any parcel by tracking number — real-time status, event history, and carrier detection.
View repository →Shipmail MCP server for AI agent custom-domain email inboxes with REST API and webhooks.
AI-powered laytime, demurrage and despatch calculator for chartering professionals.


