io.github.kapillamba4/code-memory MCP Server
io.github.kapillamba4/code-memory
Local semantic code search with Git history—no API key, runs offline, saves 50% tokens.
What is the io.github.kapillamba4/code-memory MCP server?
The code-memory MCP server is a deterministic code intelligence layer that performs semantic search across your codebase using local embeddings and Git history, with zero telemetry and no API key required. It indexes code with tree-sitter AST parsing and sentence-transformers embeddings, then routes queries through purpose-built tools for definitions, architecture, and debugging.
code-memory helps you find the right code context from large codebases without dumping entire files into prompts. It uses hybrid retrieval (BM25 + dense vectors) to search code semantically, extract documentation, and trace Git history—all running locally on your machine. Ideal for proprietary codebases, air-gapped environments, and reducing token waste.
How to install io.github.kapillamba4/code-memory
Copy-paste configuration for popular MCP clients.
CODE_MEMORY_LOG_LEVELLogging verbosity (DEBUG, INFO, WARNING, ERROR)
EMBEDDING_MODELHuggingFace model ID for embeddings
Tools & capabilities
Tools this server exposes to the agent.
index_codebase— Indexes source files and documentation in a directory using tree-sitter AST parsing and sentence-transformers embeddings. Stores results in local SQLite database.search_code— Performs semantic and structural code search with hybrid retrieval (BM25 + vector embeddings). Finds definitions, references, and file structure across indexed code.search_docs— Searches markdown documentation, READMEs, and extracted docstrings to understand architectural patterns and conceptual workflows.search_history— Searches Git history to debug regressions and understand developer intent. Supports commit search, file history, and blame queries.
Use cases
- Find where a function is defined or called across a large codebase without grepping manually
- Understand architectural patterns and workflows by searching documentation and docstrings semantically
- Debug regressions by searching Git commit history and blame information for specific files
- Reduce token usage by retrieving only relevant code snippets instead of dumping entire files into prompts
- Index and search proprietary codebases offline without sending code to external APIs
io.github.kapillamba4/code-memory MCP server FAQ
code-memory is an MCP server that indexes your codebase locally and provides semantic search across code, documentation, and Git history. It uses tree-sitter for AST parsing and sentence-transformers for embeddings—all running offline on your machine.
Yes, code-memory is open-source (MIT license) and free to use. It requires no API keys or cloud services.
Install via `pip install code-memory` or `uvx code-memory`, then add to your Claude Desktop config at `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or `%APPDATA%\Claude\claude_desktop_config.json` (Windows) with the command `uvx code-memory`.
No. code-memory runs entirely locally with no external API calls, no telemetry, and no authentication required.
Full AST support for Python, JavaScript/TypeScript, Java, Go, Rust, C/C++, Ruby, and Kotlin. Fallback whole-file indexing for C#, Swift, Scala, Lua, Shell, YAML/TOML/JSON, HTML/CSS, SQL, and Markdown.
Yes. Download the embedding model once on a connected machine (cached to ~/.cache/huggingface/), then transfer the binary and cache to your air-gapped machine. No internet required after setup.
README (reference)
Source of truth, from the repository.
code-memory
<!-- mcp-name: io.github.kapillamba4/code-memory --> <img src="assets/logo.png" alt="code-memory logo" width="100%">A deterministic, high-precision code intelligence layer exposed as a Model Context Protocol (MCP) server.
- Zero telemetry — your code never leaves your machine
- No API key required — runs entirely locally with sentence-transformers
- 1 min setup — just
uvx code-memoryand you're ready - Token saving by 50% — precise code retrieval instead of dumping entire files
Please help star code-memory if you like this project!
Why code-memory?
Finding the right context from a large codebase is expensive, inaccurate, and limited by context windows. Dumping files into prompts wastes tokens, and LLMs lose track of the actual task as context fills up.
Instead of manually hunting with grep/find or dumping raw file text, code-memory runs semantic searches against a locally indexed codebase. Inspired by claude-context, but designed from the ground up for large-scale local search.
Supported Languages
Full AST Support (structural parsing with symbol extraction): Python, JavaScript/TypeScript, Java, Go, Rust, C/C++, Ruby, Kotlin
Fallback Support (whole-file indexing): C#, Swift, Scala, Lua, Shell, Config (yaml/toml/json), Web (html/css), SQL, Markdown
Files matching
.gitignorepatterns are automatically skipped.
Architecture: Progressive Disclosure
Instead of a single monolithic search, code-memory routes queries through three purpose-built tools:
| Question Type | Tool | Data Source |
|---|---|---|
| "Where / What / How?" — find definitions, references, structure, semantic search | search_code | BM25 + Dense Vector (SQLite vec) |
| "Architecture / Patterns" — understand architecture, explain workflows | search_docs | Semantic / Fuzzy |
| "Who / Why?" — debug regressions, understand intent | search_history | Git + BM25 + Dense Vector (SQLite vec) |
| "Setup / Prepare" — index parsing & embedding generation | index_codebase | AST Parser + sentence-transformers |
This forces the LLM to pick the right retrieval strategy before any data is fetched.
Installation
From PyPI (Recommended)
# Install with pip
pip install code-memory
# Or with uvx (for MCP hosts)
uvx code-memory
From Source
# Clone the repo
git clone https://github.com/kapillamba4/code-memory.git
cd code-memory
# Install dependencies
uv sync
# Run the MCP server (stdio transport)
uv run mcp run code_memory/server.py
Pre-built Binaries (Standalone)
Download standalone executables from GitHub Releases — no Python installation required.
| Platform | Architecture | File |
|---|---|---|
| Linux | x86_64 | code-memory-linux-x86_64 |
| macOS | x86_64 (Intel) | code-memory-macos-x86_64 |
| macOS | ARM64 (Apple Silicon) | code-memory-macos-arm64 |
| Windows | x86_64 | code-memory-windows-x86_64.exe |
# Linux/macOS: Download and make executable
chmod +x code-memory-*
./code-memory-*
# Windows: Run directly
code-memory-windows-x86_64.exe
Note: The first run will download the embedding model (~600MB) to ~/.cache/huggingface/. Subsequent runs use the cached model.
Quickstart
Prerequisites
- Python ≥ 3.13
uvpackage manager (recommended) or pip
Install uv if you don't have it:
curl -LsSf https://astral.sh/uv/install.sh | sh
Install & Run
# Install from PyPI
pip install code-memory
# Or run directly with uvx
uvx code-memory
Development
# Run with the MCP Inspector for interactive debugging
uv run mcp dev code_memory/server.py
# Run tests
uv run pytest tests/ -v
# Lint and format
uv run ruff check .
uv run ruff format .
# Build package
uv build
# Build standalone binary (requires pyinstaller)
pip install pyinstaller
pyinstaller --clean code-memory.spec
# Binary output: dist/code-memory
Configure Your MCP Host
You can use either uvx (requires Python) or the standalone binary (no dependencies).
Using uvx (Python required)
Gemini CLI / Gemini Code Assist
Add to your MCP settings (e.g. ~/.gemini/settings.json):
{
"mcpServers": {
"code-memory": {
"command": "uvx",
"args": ["code-memory"]
}
}
}
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"code-memory": {
"command": "uvx",
"args": ["code-memory"]
}
}
}
Claude Code (CLI)
Add to .mcp.json in your project root or ~/.mcp.json for global access:
{
"mcpServers": {
"code-memory": {
"command": "uvx",
"args": ["code-memory"]
}
}
}
VS Code (Copilot / Continue)
Add to .vscode/mcp.json in your workspace:
{
"servers": {
"code-memory": {
"command": "uvx",
"args": ["code-memory"]
}
}
}
Using Standalone Binary (No Python required)
Replace the path with the location of your downloaded binary:
{
"mcpServers": {
"code-memory": {
"command": "/path/to/code-memory-linux-x86_64"
}
}
}
For Windows:
{
"mcpServers": {
"code-memory": {
"command": "C:\\path\\to\\code-memory-windows-x86_64.exe"
}
}
}
Shared SSE Server (Reduce Memory Usage)
By default, each MCP host project launches its own code-memory process, which loads the embedding model (~1–2 GB) once per project. To avoid this, you can run a single shared instance over SSE (Server-Sent Events) and point all your MCP hosts at it.
Start the shared server
# Using uvx (recommended)
uvx code-memory --transport sse
# Custom port and host
uvx code-memory --transport sse --port 8765 --host 127.0.0.1
# Using standalone binary
./code-memory-linux-x86_64 --transport sse
The server listens on http://127.0.0.1:8765/sse by default.
Configure MCP hosts to use the shared server
Instead of launching a new process, point your MCP host at the running SSE endpoint.
Claude Desktop
{
"mcpServers": {
"code-memory": {
"url": "http://127.0.0.1:8765/sse"
}
}
}
VS Code (Copilot / Continue)
{
"servers": {
"code-memory": {
"url": "http://127.0.0.1:8765/sse"
}
}
}
Claude Code (CLI) — .mcp.json
{
"mcpServers": {
"code-memory": {
"url": "http://127.0.0.1:8765/sse"
}
}
}
Tip: Configure
uvx code-memory --transport sseto start via a single-instance service manager (e.g. systemd user service, launchd agent, or another one-time login/startup mechanism) so the shared server starts automatically.
Security: The SSE endpoint is unauthenticated. Keep the default
--host 127.0.0.1so only local processes can connect; do not bind to0.0.0.0or a public interface unless you've put authentication in front of it.
Configuration
CLI Options
| Option | Description | Default |
|---|---|---|
--transport | Transport protocol: stdio or sse | stdio |
--port | Port for SSE transport (only when --transport sse is used) | 8765 |
--host | Host/bind address for SSE transport (only when --transport sse is used) | 127.0.0.1 |
Environment Variables
| Variable | Description | Default |
|---|---|---|
CODE_MEMORY_LOG_LEVEL | Logging verbosity (DEBUG, INFO, WARNING, ERROR) | INFO |
EMBEDDING_MODEL | HuggingFace model ID for embeddings | jinaai/jina-code-embeddings-0.5b |
Example:
CODE_MEMORY_LOG_LEVEL=DEBUG uvx code-memory
Custom Embedding Model
You can use a different embedding model by setting the EMBEDDING_MODEL environment variable:
EMBEDDING_MODEL="BAAI/bge-small-en-v1.5" uvx code-memory
For MCP hosts, add the environment variable to your configuration:
{
"mcpServers": {
"code-memory": {
"command": "uvx",
"args": ["code-memory"],
"env": {
"EMBEDDING_MODEL": "BAAI/bge-small-en-v1.5"
}
}
}
}
Note: Changing the embedding model will invalidate existing indexes. You'll need to re-run
index_codebaseafter switching models.
Tools
index_codebase
Indexes or re-indexes source files and documentation in the given directory. Run this before using search_code or search_docs to ensure the database is up to date. Uses tree-sitter for language-agnostic structural extraction and generates dense vector embeddings using sentence-transformers (runs locally, in-process) for semantic search.
index_codebase(directory=".")
search_code
Perform semantic search and find structural code definitions, locate where functions/classes are defined, or map out dependency references (call graphs). Uses hybrid retrieval (BM25 + vector embeddings) to find exact matches and semantic similarities.
search_code(query="parse python files", search_type="definition")
search_code(query="how do we establish the database connection", search_type="references")
search_code(query="src/auth/", search_type="file_structure")
search_docs
Understand the codebase conceptually — how things work, architectural patterns, SOPs. Searches markdown documentation, READMEs, and docstrings extracted from code.
search_docs(query="how does the authentication flow work?")
search_docs(query="installation instructions", top_k=5)
search_history
Debug regressions and understand developer intent through Git history.
search_history(query="fix login timeout", search_type="commits")
search_history(query="src/auth/login.py", search_type="file_history", target_file="src/auth/login.py")
search_history(query="server.py", search_type="blame", target_file="server.py", line_start=1, line_end=20)
Project Structure
code-memory/
├── code_memory/ # Package source
│ ├── server.py # MCP server entry point (FastMCP)
│ ├── db.py # SQLite database layer with sqlite-vec
│ ├── parser.py # Tree-sitter-based code parser
│ ├── doc_parser.py # Markdown documentation parser
│ ├── queries.py # Hybrid retrieval query layer
│ ├── git_search.py # Git history search module
│ ├── errors.py # Custom exception hierarchy
│ ├── validation.py # Input validation functions
│ ├── logging_config.py # Structured logging configuration
│ └── api_types.py # MCP response TypedDicts
├── tests/ # Test suite
├── pyproject.toml # Project metadata & dependencies
└── prompts/ # Milestone prompt engineering files
Troubleshooting
"Git repository not found" error
Make sure you're running search_history from within a git repository. The tool searches upward from the current directory to find .git.
Empty search results
Run index_codebase(directory=".") first to index your code and documentation. The index is stored locally in code_memory.db.
Slow indexing
Indexing generates embeddings using a local sentence-transformers model. The first run downloads the model (~600MB for jina-code-embeddings-0.5b). Subsequent runs are faster.
Embedding model errors
Ensure you have enough disk space and memory. The jina-code-embeddings-0.5b model requires ~1GB RAM when loaded.
Privacy & Security
Your code never leaves your machine. Unlike cloud-based code intelligence tools, code-memory runs entirely locally:
- Zero telemetry — no usage data, analytics, or tracking
- Zero external API calls — all processing happens in-process
- Zero cloud dependencies — works without internet (after initial setup)
- Your data stays local — indexes stored in local SQLite database
This makes code-memory ideal for:
- Proprietary and confidential codebases
- Security-conscious organizations
- Air-gapped development environments
- Privacy-focused developers
See COMPARISON.md for a detailed comparison with cloud-based alternatives.
Air-gapped & Offline Support
code-memory works in completely isolated environments:
Method 1: Pre-built Binary + Cached Model
-
On a connected machine, run code-memory once to cache the embedding model:
uvx code-memory # Model downloads to ~/.cache/huggingface/ -
Transfer to air-gapped machine:
- Standalone binary from GitHub Releases
- Model cache directory (
~/.cache/huggingface/hub/models--*)
-
Run on air-gapped machine — no network required.
Method 2: Offline pip Install
- Download the wheel from PyPI on a connected machine
- Transfer and install:
pip install code-memory-*.whl - Pre-cache the model as above
- Run offline
Roadmap
- Milestone 1 — Project scaffolding & MCP protocol wiring
- Milestone 2 — Implement
search_codewith AST parsing + SQLite +sqlite-vec - Milestone 3 — Implement
search_historywith Git integration - Milestone 4 — Implement
search_docswith semantic search - Milestone 5 — Production hardening & packaging
Contributing
See CONTRIBUTING.md for development setup and guidelines.
Changelog
See CHANGELOG.md for version history.
License
MIT
Related MCP servers

io.github.kapillamba4/meta-prompt-mcp
MCP server providing official Google and Anthropic prompting guides for meta-prompt generation.

humanMCP — kapoost
Personal MCP server for humans who create. Proof of authorship, license control.
Federated listings from personal humanMCP servers. Search offers, trades by humans.

io.github.kapruka/reviewguru
Query Sri Lankan businesses, doctors, and reviews from any MCP-aware AI agent.

OpenProject
88 tools for the OpenProject API v3: work packages, attachments, git activity, time, reporting
AI agent payments with human approval. Local-first, encrypted credentials, budget controls.
View repository →