Total Agent Memory MCP Server
io.github.vbcherepanov/total-agent-memory
Persistent local memory for AI coding agents with temporal knowledge graphs and semantic recall.
What is the Total Agent Memory MCP server?
Total Agent Memory (TAM) is an open-source MCP memory server that gives coding agents persistent, local recall across sessions. It stores decisions, solutions, facts, and errors in a SQLite database and retrieves them using BM25 full-text search, dense embeddings, fuzzy matching, and knowledge graphs. TAM works with Claude Code, Cursor, Codex, and any MCP client without requiring API keys or LLM calls for retrieval.
TAM solves the problem of coding agents starting each session without memory of earlier decisions, fixes, and conventions. It maintains a local SQLite store of project knowledge—decisions, errors, facts, and session summaries—and returns relevant results through semantic search combining full-text, embedding, fuzzy, and graph-based ranking. Developers using coding agents daily can inspect and modify the memory, and teams can run a shared server with role-based access and audit trails.
How to install Total Agent Memory
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
memory_save— Save a memory record (decision, solution, fact, error, or session summary) with optional context and project scope.memory_recall— Retrieve relevant memories using semantic search combining BM25, embeddings, fuzzy matching, and knowledge graphs.kg_add_fact— Add a fact to the knowledge graph with optional validity intervals and single-valued retirement logic.kg_at— Query knowledge graph facts at a specific time or validity interval.
Use cases
- Store architectural decisions and project conventions so agents recall them in future sessions without re-explanation.
- Build a searchable knowledge base of errors, fixes, and solutions encountered during development.
- Maintain a temporal knowledge graph of facts about your codebase that agents can query and update.
- Run a team memory server with role-based access, audit trails, and separate personal, department, and company stores.
- Benchmark and evaluate memory retrieval offline without LLM calls, enabling reproducible results.
Total Agent Memory MCP server FAQ
TAM is an open-source MCP server that gives coding agents persistent local memory across sessions. It stores decisions, solutions, errors, and facts in a SQLite database and retrieves them using semantic search (BM25, embeddings, fuzzy matching, and knowledge graphs).
Yes, TAM is open-source under the MIT license. It runs entirely on your machine with no API keys or external services required for memory operations.
Install via PyPI (`pip install total-agent-memory` or `pipx install total-agent-memory`), then run `tam setup` to auto-detect and register with your clients. Alternatively, manually add it to your client's MCP configuration with the command `total-agent-memory`.
No. TAM runs locally with no external dependencies. The default profile makes no LLM calls on write or search, so retrieval is fully offline and reproducible. Optional features like cross-encoder reranking require PyTorch but no API keys.
Yes. TAM can run as a team server with PostgreSQL backend, role-based access control, separate personal/department/company stores, authorship tracking, and audit trails.
TAM requires Python 3.11 or newer. CI tests Python 3.11, 3.12, and 3.13 on Ubuntu, Windows, and macOS.
README (reference)
Source of truth, from the repository.
Persistent, local memory for AI coding agents: Claude Code, Codex CLI, Cursor and any MCP client.
total-agent-memory (TAM) is an open-source memory server for AI coding agents. Coding agents start every session without memory of earlier ones, so decisions, fixes and project conventions have to be explained again. TAM stores decisions, solutions, facts, errors and session summaries on your machine and returns them through the Model Context Protocol (MCP), so any MCP client can use it without code changes.
Each store is one directory built around a SQLite database. Recall combines
full-text BM25, dense embeddings computed locally, fuzzy matching and a
knowledge graph, and fuses the ranked lists with reciprocal rank fusion; an
optional cross-encoder can rerank the result. The default profile makes no LLM
call on write or search, so retrieval can be measured offline and a rerun gives
the same result. Facts can carry validity intervals (kg_add_fact, kg_at),
and a newer value of a single-valued fact can retire the older one.
TAM is for developers who use coding agents daily and for researchers who need a memory baseline they can run locally, inspect and change. The same package can also run as a team server: personal, department and company areas are separate stores behind one gateway, with roles, authorship and an audit trail (team server).
Installation
TAM needs Python 3.11 or newer. CI tests Python 3.11, 3.12 and 3.13 on Ubuntu, Windows and macOS.
From PyPI (use a virtual environment, or pipx / uvx):
pip install total-agent-memory # or: pipx install total-agent-memory
uvx total-agent-memory # run once without installing
This installs the total-agent-memory (alias tam) MCP server and the
tam-team, tam-remote and lookup-memory commands. The optional
cross-encoder reranker pulls in PyTorch and is an extra:
pip install "total-agent-memory[rerank]". The PostgreSQL backend of the team
server is the [postgres] extra.
Register it with your clients. Run tam setup at a terminal. The wizard
asks "Just me" or "Company server", detects Claude Code, Claude Desktop, Codex,
Cursor, Windsurf, Gemini CLI, Cline and OpenCode, and registers the server with
the ones you pick (setup wizard).
Other channels:
| Channel | Command |
|---|---|
| Docker (linux/amd64, linux/arm64) | docker run -p 3737:3737 -p 37737:37737 -v ~/.tam:/data ghcr.io/vbcherepanov/total-agent-memory:14.7.0 — MCP over HTTP on :3737/mcp, dashboard on :37737 |
| npx connector | npx -y total-agent-memory connect claude-code (or codex, cursor, cline, continue, aider, windsurf, gemini-cli, opencode) |
| Claude Code plugin (server, skill and capture hooks) | /plugin marketplace add vbcherepanov/total-agent-memory<br>/plugin install total-agent-memory@vbcherepanov |
| Plugin for Claude Code and Cowork, from the plugin repository (server and skill, no hooks; needs uv) | /plugin marketplace add vbcherepanov/total-agent-memory-plugin<br>/plugin install total-agent-memory@vbcherepanov |
| The same plugin for Codex CLI | codex plugin marketplace add vbcherepanov/total-agent-memory-plugin<br>codex plugin add total-agent-memory@vbcherepanov |
| Source checkout with IDE hooks and background services | git clone https://github.com/vbcherepanov/total-agent-memory.git && cd total-agent-memory && ./install.sh --ide claude-code (Windows: install.ps1 -Ide claude-code) |
Use one of the two Claude Code plugins, not both. Per-platform details, WSL2, the IDE matrix, uninstalling and troubleshooting are in the installation guide.
Quick start
-
Install the package and run
tam setup, or add the server to your client by hand. For Claude Code:claude mcp add memory -- total-agent-memoryFor clients that use an
mcpServersfile:{ "mcpServers": { "memory": { "command": "total-agent-memory" } } }If you installed into a virtual environment, use the full path to its
total-agent-memoryexecutable. -
Restart the client. Memory is stored in
~/.tam/unlessTAM_MEMORY_DIRpoints elsewhere. -
Ask the agent to remember something ("remember that we chose PostgreSQL for billing because of row-level security"). It calls
memory_save. In a later session, ask "which database did we choose for billing?"; it callsmemory_recalland gets the record back.
To try the server without an agent, this script starts it over stdio with the
MCP Python SDK (installed as a dependency), saves one record and recalls it.
The SDK passes only a few variables to the server by default, so the script
hands over the full environment, including TAM_MEMORY_DIR:
import asyncio
import os
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
async def main() -> None:
server = StdioServerParameters(command="total-agent-memory", env=dict(os.environ))
async with stdio_client(server) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
await session.call_tool("memory_save", {
"type": "decision",
"content": "Chose PostgreSQL over MySQL for the billing service",
"context": "WHY: row-level security per tenant",
"project": "demo",
})
result = await session.call_tool("memory_recall", {
"query": "which database for billing", "project": "demo", "limit": 3,
})
print(result.content[0].text)
asyncio.run(main())
TAM_MEMORY_DIR="$(mktemp -d)" python quickstart.py
The first run downloads the default embedding model. The output
is JSON with the saved decision under results.decision. More examples:
tools and interfaces.
Running the tests
The tests run from a source checkout. These commands mirror the CI workflows in .github/workflows:
git clone https://github.com/vbcherepanov/total-agent-memory.git
cd total-agent-memory
python3 -m venv .venv
.venv/bin/pip install -e . -r requirements-dev.txt
export FASTEMBED_CACHE_PATH="$PWD/.tam-models" TAM_MEMORY_DIR="$PWD/.tam-test-memory"
.venv/bin/python tests/smoke/prewarm_models.py # downloads the text and code embedding models (~850 MB) once
.venv/bin/python -m pytest tests -q
The tests need no API keys and no LLM; the full suite takes 15 to 20 minutes
on a laptop. Do not set MEMORY_LLM_ENABLED=false for the full suite: the
configuration tests check LLM auto-detection. Tests that need services you do not
have are skipped:
- PostgreSQL (team server backend): install the extra with
.venv/bin/pip install -e ".[postgres]", then run.venv/bin/python -m pytest tests -m postgres --backend=postgresor--backend=both. WithoutTAM_TEST_PG_URLthe fixtures start a pgvector container through Docker. See CONTRIBUTING.md and .github/workflows/postgres.yml. - Browser tests (
tests/browser, Playwright with Chromium, Firefox and WebKit) run in the image built from docker/Dockerfile.browser; see thebrowserjob in .github/workflows/smoke.yml. - Installed-package smoke test:
python tests/smoke/installed_runtime.pychecks an installed wheel over local and remote MCP.
Reproducing the benchmarks
The retrieval benchmarks (LoCoMo, LongMemEval, BEAM) need no API key and run with scripts in benchmarks/ once the public datasets are downloaded. Results with the default profile, v13.0.0:
| Benchmark | Metric | Result |
|---|---|---|
| LongMemEval (470 questions) | R@5 (recall_any) | 95.1% |
| LoCoMo (1,536 questions) | R@5 | 0.607 |
| BEAM, 1M-token scale (625 probes) | R@5 | 0.448 |
Dataset locations, commands, end-to-end (LLM-judged) accuracy, negative controls, latency and the 14.x studies are in docs/benchmarks.md.
Documentation
- Installation guide: all channels, per-platform setup, WSL2, IDE matrix, troubleshooting
- Setup wizard:
tam setupand the team server's web wizard - Tools and interfaces: the 77 MCP tools, CLI, TypeScript SDK, dashboard
- Configuration: environment variables, LLM providers, performance tuning
- Local settings page and internal LLM settings
- Architecture
- Team server quick start, with dashboard, PostgreSQL, backup, onboarding and reports
- Benchmarks and comparison with other systems (April 2026 snapshot)
- Updating and upgrading
- What is new in 14.x, roadmap and v8–v13 history, CHANGELOG
- Security policy
Citation
A paper describing TAM is under review at the Journal of Open Source Software (paper/paper.md). Until it is published, please cite the software and the preprint:
Cherepanov, V. total-agent-memory (version 14.7.0) [software]. https://github.com/vbcherepanov/total-agent-memory
Preprint: doi:10.5281/zenodo.23011523
Citation metadata for the software is in CITATION.cff.
Contributing
Issues, pull requests and benchmark reproductions are welcome. See CONTRIBUTING.md for the development setup, the rules for a pull request and the commit convention. Report security issues privately as described in SECURITY.md. Donations to support development: PayPal.
License
MIT. See LICENSE. Third-party licenses are listed in THIRD-PARTY-LICENSES.md.
Related MCP servers

Lviv Public Transport
Lviv public transport MCP: stops, timetables, routes, and live vehicle positions. No API key.

Dokploy MCP
Dokploy MCP server for Codex, Cursor, and Claude with compact stdio and hosted HTTP modes
CISSP exam prep: 3,600+ practice questions, study stats, and readiness scoring
View repository →
io.github.vdalhambra/financekit-mcp
Stock quotes, technical analysis, crypto prices, and portfolio insights for AI agents

io.github.vdalhambra/siteaudit-mcp
SEO, performance, and security audits for any URL — no API keys required

io.github.vdappdev2/address
MCP server for Verus addresses — generate, validate, list transparent and shielded addresses
