memo MCP Server
io.github.jagoff/memo
Persistent, searchable memory for AI agents — 100% local, no cloud, fewer tokens.
What is the memo MCP server?
The memo MCP server gives Claude Code, Cursor, Cline, and other AI agents persistent, searchable memory that runs entirely on your machine. It uses hybrid retrieval (vector + BM25), stores memories as plain Markdown files, and injects relevant context automatically while spending fewer tokens than alternatives.
memo solves AI agent amnesia by providing a local-first memory system that persists across sessions. It combines vector search (MLX on Apple Silicon, CPU on Linux) with full-text search, detects contradictions in your knowledge base, maintains historical versions, and runs autonomous nightly optimization—all without cloud APIs, vector databases, or Ollama.
How to install memo
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
save— Save a fact or decision to memory.search— Search memory by meaning using hybrid vector + BM25 retrieval.ask— Query memory with natural language and get synthesized answers.recall— Inject relevant memories automatically into agent context.as-of— Query memory as it existed on a specific past date.diff— Show what changed in memory between two dates.contradict scan— Find conflicting facts across the memory corpus.contradict triage— Resolve contradictions by fusing, keeping newer, or dismissing.dream run— Run autonomous nightly optimization pipeline: inventory, signal mining, conflict resolution, pruning, synthesis, optimization, and pre-warming.resume— Reopen any session from any agent.sync— Sync memory across machines via git remote.chat— Local chat UI over your memory.reindex— Rebuild SQLite index from Markdown source files.history— View and query memory change history.graph— Visualize knowledge graph with optional code symbol edges.export— Export memory to various formats.import— Import memories from external sources.
Use cases
- Save architectural decisions and design patterns once, have every agent session recall them automatically without re-explaining.
- Query your memory by meaning (e.g., 'what database did we pick?') and get synthesized answers grounded in your actual decisions.
- Detect when you change your mind about a decision and automatically flag stale memories so agents stop reintroducing old choices.
- Rewind memory to any past date to understand why a decision was made, not just what was decided.
- Run nightly optimization to prune stale facts, resolve contradictions, synthesize cross-cluster insights, and pre-warm embeddings for faster recall.
memo MCP server FAQ
memo is a persistent, searchable memory system for AI agents that runs 100% locally on your machine. It stores memories as Markdown files, uses hybrid vector + BM25 search, and automatically injects relevant context into agent sessions—all without cloud APIs or external databases.
Yes. memo is MIT-licensed open source. Installation is free via pip, brew, or the installer script. It requires ~8 GB of disk for bundled ML models on first install.
Run the installer (`curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.16.0/install.sh | bash`), then run `memo doctor`. The installer automatically wires memo into every MCP client it finds (Claude Desktop, Cursor, Cline, Continue). Per-client setup docs are in the reference.
No. memo runs entirely on your machine with no cloud APIs, no API keys, and no external services. All embeddings, search, and synthesis happen locally. Memory syncs only if you explicitly point `memo sync` at a git remote you own.
macOS with Apple Silicon (M1–M4) gets full support with MLX. Linux/Ubuntu works with CPU backend (`pipx install "mlx-memo[cpu]"`). Intel Macs are unsupported. Python ≥ 3.13 is required. Docker is available for cross-platform use.
memo's default MCP profile exposes 43 tools with ~9.7k schema tokens, versus 165 tools / ~30.6k tokens on full surfaces—68% less overhead per session. Ambient recall injects one relevant memory capped at ~160 tokens. Use `memo roi` and `memo tokens` to measure real savings in your setup.
README (reference)
Source of truth, from the repository.
memo
Your coding agent starts every session with amnesia. memo fixes that — 100% on your own machine.
Persistent, searchable memory for Claude Code, Codex, Cursor, Cline, Devin, and OpenCode. No cloud, no API keys, no Ollama, no vector DB to run. And it spends fewer tokens, not more.

Install
curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.16.0/install.sh | bash
<sub>Prefer a package manager? uv tool install mlx-memo · pipx install mlx-memo · brew tap jagoff/memo && brew install mlx-memo</sub>
Then:
memo doctor # self-check
memo save 'we use Postgres, not Mongo' # save a decision
memo search 'what database did we pick?' # search by meaning
That's it. Your agents pick it up over MCP automatically — the installer wires every client it finds.
<details> <summary>Installing on another Mac or handing setup to an agent?</summary>New Mac:
curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.16.0/install.sh | bash
memo sync bootstrap git@github.com:yourname/memo-sync.git
Agent-managed setup:
curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.16.0/install.sh | bash
memo doctor --strict-runtime
</details>
On Linux or just want to look around first?
docker run --rm ghcr.io/jagoff/memo:latest memo doctor
Why this saves you money
Most memory servers add context. memo is built to remove it.
| Profile | Tools | Schema tokens |
|---|---|---|
agent (default) | 43 | ~9.7k |
core / slim | 60 | ~13.2k |
full / default | 165 | ~30.6k |
The default MCP surface is 43 tools, not 165 — 74% fewer tools, and about 68% less schema context: 43 tools / ~9.7k schema tokens versus 165 tools / ~30.6k tokens on the full surface — overhead paid every session, in every client.
Ambient recall injects one relevant memory before the model answers. The bundled Claude Code hook caps that injection at ~160 tokens. memo roi reports the real grounding and re-ask counts — the estimated-savings figure it used to print was removed in 4.14.0, because multiplying those counts by hardcoded constants was a savings claim memo could not support. For measured savings, memo tokens reads the provider's own usage counters through the context-compression proxy.
memo roi # value from grounded recalls and avoided re-asks
memo tokens # usage-savings ledger
Three things nothing else does
🕰️ Time-machine — query your knowledge as it was
memo as-of ask "what was the deploy strategy?" --date 2026-02-01
memo diff --from 2026-01-01 --to 2026-03-01
Full historical reconstruction by reverse-replaying history.db. Useful when you need to know why past-you made a call, not just what past-you decided.
⚡ Contradiction radar — memory that notices when you change your mind
memo contradict scan # find conflicting facts corpus-wide
memo contradict triage # resolve: fuse / newer-wins / dismiss
Change a decision and memo flags the now-stale version, so the agent stops reintroducing what you already threw out.
🔮 Dream — it optimizes itself while you sleep
memo dream run
A 7-phase nightly pipeline: inventory → mine signals → resolve conflicts → prune stale → synthesize cross-cluster insights → optimize → pre-warm the top-100 query embeddings so tomorrow's recall stays under 200 ms. Every run writes a receipt you can audit. Zero intervention.
How it works
Hybrid retrieval. A vector leg (MLX on Apple Silicon, sentence-transformers on CPU) and a BM25 leg (FTS5, diacritic-folding for Spanish) run in parallel, fuse via Reciprocal Rank Fusion, then go through an optional MLX cross-encoder rerank.
Markdown is the source of truth. Every memory is a plain .md file you can read, grep, and version-control. SQLite is a derived index that rebuilds from the files at any time — hand-edit in Obsidian and your edit wins on the next memo reindex. Nothing is locked in a database you can't open.
Prompts and memories stay on your machine. Embedder, reranker, and LLM all run in-process. No telemetry. Memory travels only if you point memo sync at a git remote you own. Normal startup is fully offline; remote update checks and auto-update require an explicit opt-in. → Privacy and network policy
Also in the box: cross-agent memo resume (reopen any session from any agent), cross-Mac git sync, a knowledge graph with optional codegraph symbol edges, encrypted secret storage, OCR/audio ingestion, evidence packs, outcome learning, signed federation, and a local chat UI over your memory (memo chat serve). → Full feature reference
How it compares
Verified July 2026 against each project's own docs. Corrections welcome — open an issue and I'll fix the table.
| memo | mem0 | letta | cognee | basic-memory | cipher | |
|---|---|---|---|---|---|---|
| 100% local, no cloud API | ✅ | ⚠️ | ⚠️ | ⚠️ | ✅ | ⚠️ |
| Time-machine (rewind to any date) | ✅ | ❌ | ⚠️ | ❌ | ⚠️ | ⚠️ |
| Contradiction detection + resolution | ✅ | ⚠️ | ⚠️ | ❌ | ❌ | ❌ |
| Autonomous nightly maintenance | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Token-economy MCP profiles | ✅ | ❌ | ❌ | ⚠️ | ✅ | ❌ |
| Markdown / Obsidian as source of truth | ✅ | ❌ | ⚠️ | ❌ | ✅ | ❌ |
<sub>✅ first-class · ⚠️ partial, config-gated, or add-on · ❌ absent</sub>
Closest comparators are basic-memory (local-first + Obsidian + MCP — same thesis) and cipher (memory for coding agents).
Requirements
| Support | |
|---|---|
| macOS, Apple Silicon (M1–M4) | Full — MLX embedder + reranker + ask/synthesize/dream |
| Linux / Ubuntu | Standalone CPU backend — search, recall, save. pipx install "mlx-memo[cpu]" · docs/ubuntu.md |
| Intel Mac | Unsupported — current PyTorch releases do not ship Python 3.13 wheels for this platform |
| Docker | Cross-platform, CPU backend · docs/docker.md |
Python ≥ 3.13 (the installer handles this via uv if you don't have it). First install pulls ~8 GB of models, 5–15 min. Optional: an Obsidian vault — without one, memo uses ~/Documents/memo/.
Docs
| Install detail, installer knobs, new-Mac migration | reference.md › Install |
| Per-client MCP setup (Claude Desktop, Cursor, Cline, Continue) | reference.md › MCP setup |
| Ambient recall, capture, and tuning | reference.md › Ambient memory |
Full CLI reference (145 commands) + memo tui | reference.md › CLI |
All MEMO_* flags and model profiles | reference.md › Configuration |
| Architecture and design notes | reference.md › Design |
| Privacy and network policy | PRIVACY.md |
All 145 top-level CLI commands
<details> <summary>Complete command inventory (kept here so CI detects CLI/documentation drift)</summary>Core: save search ask get edit rename delete list
Recall & Hooks: recall recall-hook context briefing continuity prewarm capture-tick capture-stop interject ask-gaps guard digest
Session & History: history as-of diff record-history session chat-session resume reflect mine-history episodes chronicle
Maintenance: reindex maintain review dream consolidate synthesize dedupe cross-dedup retier contradict coordinate terminal invalidate temporal compress-context ops
Analysis & Quality: health stats doctor journey-check lint drift analytics eval roi tokens token-savings usefulness gaps outcome profile confidence graduation hype definitive evidence
Knowledge Graph: graph entities entity extract-entities links version related
Advanced Search: embed rerank contextual retrieve context-pack chat chat-ask repo
Import / Export / Sync: import export backup restore sync ingest federation
Visualization: tui dashboard map logs hook-log
Setup & Config: init setup config install-mcp install-watcher uninstall-watcher install-slash install-statusline install-recall-hook install-shell-wrapper install-shims startup-banner migrate migrate-vault migrate-independence update upgrade self-update watch release onboard
Daemons: daemons recall-daemon ingest-daemon maint-daemon embed-daemon idle-daemon
Other: backend-native collaborative events feedback query mandate drift sleep-cycle operational ocr-image provenance secret verbatim mcp-command codex-badge debug-recall http-api proxy mine-git token-gate fix undo code-facts code-nudge code-health
Contributing
git clone https://github.com/jagoff/memo && cd memo
uv pip install -e '.[dev]'
Issues and PRs welcome — see CONTRIBUTING.md. If memo is useful to you, a ⭐ genuinely helps other people find it.
MIT licensed. Built on Apple MLX, sqlite-vec, and codegraph.
Related MCP servers

SEC Intelligence
SEC 10-K/10-Q/8-K search, RAG Q&A, comparison, and anomaly detection with citations.

Vocab Voyage
20 MCP tools + 17 widgets for SAT/ISEE/SSAT/GRE/GMAT/LSAT prep. Flashcards, quizzes & games. Hosted.

MCP server fronting a self-hosted coordination-bus HTTP API: post/read/claim/release/heartbeat.

Windows desktop-control MCP server: screenshot, window mgmt, mouse/keyboard input, recording.

MCP server for the Discord REST API: 5 read tools always on, 7 write tools env-gated, default off.

MCP server for the GitHub REST API: issues, PRs, repos; read always on, write env-gated.
