Infimium MCP Server
io.github.aryankumar06/infimium
Private AI context layer for semantic code search, dependency graphs, memory, and workspaces—100% local, zero token bloat.
What is the Infimium MCP server?
Infimium is a private context layer and super brain for your codebase that gives AI agents persistent memory, deep dependency graphs, and instant code context. It runs 100% locally using semantic search, symbol expansion, and project memory to reduce token overhead by ~99.5% compared to full-text retrieval.
Infimium solves the problem of AI agents reading too much code or missing the right symbols in large repositories. It indexes your codebase locally, provides semantic code search that returns AST signatures first, maintains project memory across sessions, and builds dependency graphs—all without sending code to external services. Use it to give Claude, Cursor, or other MCP clients deep understanding of your repository with minimal token cost.
How to install Infimium
Copy-paste configuration for popular MCP clients.
SEARCH_API_KEYsecretOptional Tinyfish API key for web_search.
SEARCH_PROVIDERSearch provider. Use tinyfish for the current web_search implementation.
CODEBASE_PATHAbsolute path to the codebase Infimium should index.
LOCAL_DOCS_PATHAbsolute path to local docs Infimium should index.
CHROMADB_HOSTChromaDB HTTP endpoint.
OLLAMA_HOSTOllama HTTP endpoint.
SHELL_ALLOWLISTComma-separated commands allowed by the shell tool.
Tools & capabilities
Tools this server exposes to the agent.
hello_infimium— Confirms the MCP server is healthy.get_context— Reads saved YAML repo context, current memory and handoff; explicit refresh updates Git/index state.infimium_update— Refreshes episodic memory and handoff graphs; controls automatic memory checkpoints.semantic_code_search— Finds code by meaning and returns symbol signatures first.expand_symbol— Loads one full implementation only when needed.query_local_docs— Searches local Markdown, text, HTML, and PDF files.dep_graph— Shows imports, callers, callees, and HTTP routes for a symbol.project_memory— Keeps active scratchpad events, compact milestones, and durable project rules across agents.plan— Builds a grounded implementation plan from code and graph context.web_search— Searches the web through optional Tinyfish configuration.fetch_url— Extracts readable Markdown or text from a URL.shell— Runs allowlisted commands with timeouts and output limits.
Use cases
- Semantic code search to find relevant symbols and implementations without reading entire files
- Build implementation plans grounded in actual codebase structure and dependencies
- Maintain persistent project memory and task context across multiple agent sessions
- Explore dependency graphs to understand callers, callees, and API routes for any symbol
- Reduce token overhead by retrieving AST signatures first and expanding full code only when needed
Infimium MCP server FAQ
Infimium is a private context layer that indexes your codebase locally and gives AI agents semantic code search, dependency graphs, and persistent memory. It reduces token overhead by ~99.5% by returning AST signatures first instead of full implementations.
Yes, Infimium is free and open-source under the MIT license. It runs entirely on your machine with no external service required (except optional web search via Tinyfish).
Run `npx infimium@latest setup` from your project folder, then add the MCP configuration to your client: `{"mcpServers": {"infimium": {"command": "npx", "args": ["-y", "infimium", "serve"]}}}`. Restart your client and use commands like `Use Infimium semantic_code_search`.
Node.js 22.5+, and Ollama (optional but recommended for embeddings and plan generation). Infimium stores all code, embeddings, and memory locally in `~/.infimium/`.
No. Code, docs, embeddings, memory, graphs, and queries remain 100% local. Infimium only sends anonymous lifecycle telemetry (which can be disabled) and optionally uses Tinyfish for web search if configured.
Infimium drops initial payload cost from ~1,460 tokens per symbol to ~8 tokens by returning AST signatures first. Full implementations are loaded only on demand with `expand_symbol`.
README (reference)
Source of truth, from the repository.
Infimium
The Private Context Layer & Super Brain for Your Codebase. Give AI agents persistent memory, deep dependency graphs, and instant code context -- 100% local, zero token bloat.
Demo
Why
Large repositories make agents read too much code or miss the right symbol. Infimium retrieves compact, relevant context before the agent starts editing.
200,000 lines of code
Agent reads everything -> context blown + expensive
grep "price calculation" -> misses calcPropertyValue()
tool: semantic_code_search
query: "price calculation logic"
-> services/property/calc.ts:142 · calcPropertyValue()
-> callers: getListingPrice(), estimatePropertyTax()
Quick Start
Requires Node.js 22.5+. From your project folder:
cd /path/to/your/project
npx infimium@latest setup
Run setup from the repository you want to index, not from your home directory (~).
Infimium stops broad roots automatically so it cannot scan unrelated files.
That creates global config, starts Ollama if it is installed, pulls nomic-embed-text, indexes the current project or workspace, runs doctor, and opens Playground.
The published CLI keeps its executable entrypoint, so MCP clients can launch it directly through the configuration below.
If Ollama is not installed yet:
npx infimium@latest setup --install-deps
infimium setup creates one global config at ~/.infimium/.env. You do not need a .env in every project. Code, docs, memory, graphs, and vectors are stored locally under ~/.infimium/.
Web search is optional. Add a Tinyfish key only when you need it:
SEARCH_PROVIDER=tinyfish
SEARCH_API_KEY=your_key
Full infimium plan generation also needs a local text model:
ollama pull llama3.1
infimium plan --dry-run "your task" works without this model and shows the retrieved code context first.
Connect Your Agent
Cursor, Windsurf, Claude Desktop, and other MCP clients:
{
"mcpServers": {
"infimium": {
"command": "npx",
"args": ["-y", "infimium", "serve"]
}
}
}
Restart the client, then use:
Use Infimium hello_infimium.
Use Infimium get_context before starting.
Use Infimium semantic_code_search to explain this repository.
Infimium normally uses the MCP process working directory. If your client starts it elsewhere, pass project_path once; Infimium remembers the active project and auto-indexes it.
Tools
| Tool | What it does |
|---|---|
hello_infimium | Confirms the MCP server is healthy. |
get_context | Reads saved YAML repo context, current memory and handoff; explicit refresh updates Git/index state. |
infimium_update | Refreshes episodic memory and handoff graphs; controls automatic memory checkpoints. |
semantic_code_search | Finds code by meaning and returns symbol signatures first. |
expand_symbol | Loads one full implementation only when needed. |
query_local_docs | Searches local Markdown, text, HTML, and PDF files. |
dep_graph | Shows imports, callers, callees, and HTTP routes for a symbol. |
project_memory | Keeps active scratchpad events, compact milestones, and durable project rules across agents. |
plan | Builds a grounded implementation plan from code and graph context. |
web_search | Searches the web through optional Tinyfish configuration. |
fetch_url | Extracts readable Markdown or text from a URL. |
shell | Runs allowlisted commands with timeouts and output limits. |
CLI
| Command | Description |
|---|---|
infimium doctor | Run health checks on your dependencies and setup. |
infimium status | Show the current status of the index and memory. |
infimium --help | Shows all relevant cli commands. |
infimium playground | Launch the local web UI to explore index, graph, and memory. |
infimium index | Scan and index the current project directory (code, docs, dependencies). |
infimium watch | Run the indexer in watch mode to continuously index changes. |
infimium get-context | Output the full flattened context as YAML (layer.md). |
infimium code-search <query> | Semantically search code and return symbol signatures. |
infimium expand-symbol <symbol> | Fetch the full implementation code for a specific symbol. |
infimium docs-search <query> | Semantically search local markdown/text documentation. |
infimium dep-graph <symbol> | Show dependencies, callers, callees, and route graph. |
infimium plan --dry-run "<task>" | Draft an implementation plan based on a given prompt. |
infimium remember "<note>" | Add a milestone, progress, or decision to project memory. |
infimium resume | Show the active task and recent scratchpad memory events. |
infimium memory complete | Compact the active scratchpad into an archived milestone. |
infimium memory search "<query>" | Semantically search past project rules and memory ledger. |
Use npx infimium ... if you did not install the package globally.
Project Memory
Refresh memory yourself, or enable periodic checkpoints for the current project:
infimium update --note "Implemented login validation" --task "Finish login" --handoff "Run the auth tests next"
infimium update start --interval 300
infimium update status
infimium update stop
Replace the example notes with your own. Add --project /path/to/repo to select a project and --file src/auth.ts to attach a relevant file. infimium_update and infimium-update are CLI aliases. MCP agents use the infimium_update tool with action: refresh|start|stop|status, project_path, and optional note, task, handoff, files, or interval_seconds.
Auto-update is opt-in and runs while the foreground CLI watcher or an MCP server is open. Its per-project setting survives restarts; stop disables future checkpoints (an in-flight refresh may finish). Checkpoints link episodes, tasks, file references, and handoff notes in local SQLite. Unchanged observations are deduplicated. Automatic checkpoints record observable state, not guessed intent, and never mark a task complete.
get-context / get_context now reads saved context and current memory without rescanning the repo. It includes a bounded memory graph and guidance to answer repo-overview questions only when asked, using Infimium memory first. Missing context is reported explicitly. Run infimium update or get-context --refresh to refresh filesystem context. These are agent guidelines, not an enforcement mechanism for other clients.
Infimium keeps memory bounded across long sessions:
- Scratchpad: recent events for the active task.
- Archive: compact summaries of completed tasks.
- Ledger: durable decisions, rules, quirks, and unresolved blockers.
Record meaningful progress while working:
infimium remember "Added rate-limit middleware" --type progress --task "Rate limiting"
infimium remember "Use Redis-backed counters in production" --type decision
When the task is complete:
infimium memory complete
Infimium uses the local llama3.1 model when available and falls back to deterministic compaction when it is not. Raw compacted events remain stored locally for seven days before pruning. get_context never calls an LLM or network service.
From a source checkout, build once and run the local playground with:
npm run build
npm run playground
Local Architecture
- Ollama creates embeddings on your machine.
- Embedded SQLite stores vectors, index metadata, project memory, and graph edges. No ChromaDB or Docker service is required.
- Documents use recursive boundary-aware chunks instead of blind fixed slices.
- JavaScript, TypeScript, Python, and Dart parsers are bundled.
- Go, Rust, and Java Tree-sitter WASM grammars download on first use and cache in
~/.infimium/grammars/. .gitignore,.infimiumignore, and framework defaults exclude dependencies, build output, Flutter artifacts, caches, and binaries before indexing.semantic_code_searchreturns signatures;expand_symbolprovides full code on demand.- Project memory uses session-scoped scratchpads, compact milestone archives, and a versioned semantic ledger.
get_contextemits static anchors, dynamic repository state, and active execution as separate YAML zones.
Multiple Projects
Run the normal index command from a folder containing related projects:
infimium index
Infimium detects immediate project roots from files such as pubspec.yaml, package.json, Cargo.toml, and go.mod. It shows the detected roles and dependencies, asks once, then creates infimium.workspace.json, indexes every project, and opens Playground.
For unattended setup:
infimium index --yes --no-playground
Use --no-workspace to index only the current project. Workspace projects keep separate memory and Git state while get_context includes balanced summaries and graph relationships from related projects.
Infimium - Playground
Infimium drops the initial payload cost from approximately 1,460 tokens to 8 tokens per symbol. Semantic search returns the AST signature first; the agent requests the full implementation only when it needs it with expand_symbol.
Full implementation ~1,460 tokens
AST skeleton ~8 tokens
Initial payload reduction ~99.5%
These are Playground reference values, not a claim that every function has the same size. Inspect your own indexed repository and compare AST-first retrieval with full-text retrieval locally:
infimium playground
Open Token Economics to see the estimated token difference across your actual indexed symbols.
Privacy
Code, docs, embeddings, memory, graph data, prompts, queries, file paths, and repo names remain local.
Infimium sends privacy-safe anonymous lifecycle telemetry so we can understand setup success:
init_started,init_completeddoctor_run,doctor_passedindex_started,index_completed,setup_completedserve_started,first_tool_call,playground_opened
Telemetry includes an anonymous install ID, Infimium version, OS, Node major version, timestamp, and event name. It never includes code, file paths, repo names, prompts, search queries, memory notes, API keys, or user identity.
Disable it anytime:
infimium telemetry off
or set:
INFIMIUM_TELEMETRY=false
Troubleshooting
FAQ & Common Confusions
Where is layer.md?
When you run infimium get-context, it prints saved context directly to your terminal (stdout). Refreshes store project-scoped YAML under Infimium's local data directory. To export the saved context to a file, use terminal redirection:
infimium get-context > layer.md
Why does the Playground UI say "Awaiting first agent interaction..."?
The CURRENT TASK tracker reflects stored project context. Run infimium update --task "Your task" to refresh it; get-context reads the saved snapshot.
How do I format infimium remember?
The infimium remember command requires a message and a --type flag (valid types: note, progress, decision, blocker, index, plan). If you also want it to update the active task in the Playground, include the --task flag:
infimium remember "Added rate limiting" --type progress --task "Security Features"
Database is locked
If you see Failed to start Infimium: Database is locked, it means another instance of Infimium is actively holding a lock on the SQLite memory database. This usually happens if you try to run infimium index manually in one terminal while infimium playground or infimium watch is still running in another. Simply stop the running process (Ctrl+C) before running manual commands.
General Setup Issues
Run:
infimium doctor
Every failed check prints one copy-paste fix. If setup still fails, give this prompt to your coding agent:
Set up Infimium in this repository. Install/start Ollama, pull nomic-embed-text,
run npx infimium init, run npx infimium index, and make all six
npx infimium doctor checks pass. Do not commit secrets.
Contributing
See CONTRIBUTING.md. Adding a language starts with a parser fixture and extraction test.
Self-hosting is free forever under the MIT license.
Related MCP servers
Tells you what leaked in a .har capture, who's on the page, and what to strip. No network calls.

Marrow
Shared memory for parallel AI coding agents. Local, free, with rooms and file claims.

pgops
Safe, audited PostgreSQL operations for AI agents: queries, migrations, EXPLAIN, containers

io.github.asadullokhn/teztun
Manage TezTun tunnels, subdomains, and service tokens from any MCP-compatible AI client.

Renology Renovation Cost Data
Read-only 2026 city renovation cost ranges with sources, limitations, and citation guidance.

Leteo
Persistent memory for AI coding agents: one Rust binary, one local SQLite database.

