PluginBench
MCP Server
Active
MIT

io.github.alex-feel/mcp-context-server MCP Server

io.github.alex-feel/mcp-context-server

Persistent multimodal context storage for LLM agents with full-text, semantic, and hybrid search.

What is the io.github.alex-feel/mcp-context-server MCP server?

The MCP Context Server is a high-performance Model Context Protocol server that provides persistent multimodal context storage for LLM agents. It enables seamless context sharing across multiple agents working on the same task through thread-based scoping, with support for text and images, flexible metadata filtering, and advanced search capabilities including full-text, semantic, and hybrid search.

Store and retrieve context entries with rich metadata across multiple agents and threads. The server supports multimodal content (text and images), offers powerful search options (full-text, semantic, hybrid, and grep-style pattern matching), automatic LLM-based summarization, partial reads via character/line/outline ranges, and record navigation with table-of-contents generation. Choose between SQLite (zero-config) or PostgreSQL (high-concurrency) backends.

How to install io.github.alex-feel/mcp-context-server

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • LOG_LEVEL

    Log level

  • STORAGE_BACKEND

    Storage backend type: sqlite (default) or postgresql

  • MAX_IMAGE_SIZE_MB

    Maximum individual image size in megabytes

  • MAX_TOTAL_SIZE_MB

    Maximum total request size in megabytes

  • DB_PATH

    Custom database file location path

  • POOL_MAX_READERS

    Maximum number of concurrent read connections in the pool

  • POOL_MAX_WRITERS

    Maximum number of concurrent write connections in the pool

  • POOL_CONNECTION_TIMEOUT_S

    Connection timeout in seconds

  • POOL_IDLE_TIMEOUT_S

    Idle connection timeout in seconds

  • POOL_HEALTH_CHECK_INTERVAL_S

    Connection health check interval in seconds

  • RETRY_MAX_RETRIES

    Maximum number of retry attempts for failed operations

  • RETRY_BASE_DELAY_S

    Base delay in seconds between retry attempts

  • RETRY_MAX_DELAY_S

    Maximum delay in seconds between retry attempts

  • RETRY_JITTER

    Enable random jitter in retry delays

  • RETRY_BACKOFF_FACTOR

    Exponential backoff multiplication factor for retries

  • SQLITE_FOREIGN_KEYS

    Enable SQLite foreign key constraints

  • SQLITE_JOURNAL_MODE

    SQLite journal mode (e.g., WAL, DELETE)

  • SQLITE_SYNCHRONOUS

    SQLite synchronous mode (e.g., NORMAL, FULL, OFF)

  • SQLITE_TEMP_STORE

    SQLite temporary storage location (e.g., MEMORY, FILE)

  • SQLITE_MMAP_SIZE

    SQLite memory-mapped I/O size in bytes

  • SQLITE_CACHE_SIZE

    SQLite cache size (negative value for KB, positive for pages)

  • SQLITE_PAGE_SIZE

    SQLite page size in bytes

  • SQLITE_WAL_AUTOCHECKPOINT

    SQLite WAL autocheckpoint threshold in pages

  • SQLITE_BUSY_TIMEOUT_MS

    SQLite busy timeout in milliseconds

  • SQLITE_WAL_CHECKPOINT

    SQLite WAL checkpoint mode (e.g., PASSIVE, FULL, RESTART)

  • SHUTDOWN_TIMEOUT_S

    Server shutdown timeout in seconds

  • SHUTDOWN_TIMEOUT_TEST_S

    Test mode shutdown timeout in seconds

  • QUEUE_TIMEOUT_S

    Queue operation timeout in seconds

  • QUEUE_TIMEOUT_TEST_S

    Test mode queue timeout in seconds

  • CIRCUIT_BREAKER_FAILURE_THRESHOLD

    Circuit breaker failure threshold before opening

  • CIRCUIT_BREAKER_RECOVERY_TIMEOUT_S

    Circuit breaker recovery timeout in seconds

  • CIRCUIT_BREAKER_HALF_OPEN_MAX_CALLS

    Maximum calls allowed in circuit breaker half-open state

  • POSTGRESQL_CONNECTION_STRING
    secret

    Complete PostgreSQL connection string (overrides individual settings if provided)

  • POSTGRESQL_HOST

    PostgreSQL server host address

  • POSTGRESQL_PORT

    PostgreSQL server port number

  • POSTGRESQL_USER

    PostgreSQL database username

  • POSTGRESQL_PASSWORD
    secret

    PostgreSQL database password

  • POSTGRESQL_DATABASE

    PostgreSQL database name

  • POSTGRESQL_POOL_MIN

    PostgreSQL connection pool minimum size

  • POSTGRESQL_POOL_MAX

    PostgreSQL connection pool maximum size

  • POSTGRESQL_POOL_TIMEOUT_S

    PostgreSQL connection pool timeout in seconds

  • POSTGRESQL_COMMAND_TIMEOUT_S

    PostgreSQL command execution timeout in seconds

  • POSTGRESQL_MIGRATION_TIMEOUT_S

    Timeout in seconds for PostgreSQL migration operations (default: 300)

  • POSTGRESQL_MAX_INACTIVE_LIFETIME_S

    Close idle PostgreSQL connections after this many seconds (0 to disable, default: 300)

  • POSTGRESQL_MAX_QUERIES

    Recycle PostgreSQL connections after this many queries (0 to disable, default: 10000)

  • POSTGRESQL_TCP_KEEPALIVES_IDLE_S

    Seconds of idle time before sending first TCP keepalive probe (0 to disable, default: 15)

  • POSTGRESQL_TCP_KEEPALIVES_INTERVAL_S

    Seconds between subsequent TCP keepalive probes (0 to disable, default: 5)

  • POSTGRESQL_TCP_KEEPALIVES_COUNT

    Number of failed TCP keepalive probes before connection is considered dead (0 to disable, default: 3)

  • POSTGRESQL_STATEMENT_CACHE_SIZE

    asyncpg prepared statement cache size. Set to 0 for external pooler compatibility (PgBouncer transaction mode, Pgpool-II, etc.). Default: 100

  • POSTGRESQL_MAX_CACHED_STATEMENT_LIFETIME_S

    Maximum lifetime of cached prepared statements in seconds (default: 300). Has no effect when statement_cache_size=0

  • POSTGRESQL_MAX_CACHEABLE_STATEMENT_SIZE

    Maximum size of statement to cache in bytes (default: 15360). Has no effect when statement_cache_size=0

  • POSTGRESQL_SSL_MODE

    PostgreSQL SSL mode (disable, allow, prefer, require, verify-ca, verify-full)

  • POSTGRESQL_SCHEMA

    PostgreSQL schema name for table and index operations (default: public)

  • ENABLE_SEMANTIC_SEARCH

    Enable semantic search functionality

  • ENABLE_EMBEDDING_GENERATION

    Enable embedding generation for stored context. Default true - server fails if dependencies not met. Set false to disable embeddings.

  • OLLAMA_HOST

    Ollama API host URL for embedding generation

  • OLLAMA_AUTO_PULL

    Automatically pull missing Ollama models on startup (default: true)

  • OLLAMA_PULL_TIMEOUT_S

    Timeout in seconds for pulling Ollama models (default: 900, range: 30-3600)

  • EMBEDDING_OLLAMA_TRUNCATE

    Ollama embedding truncation mode: false (default) returns error when context exceeded, true enables silent truncation

  • EMBEDDING_OLLAMA_NUM_CTX

    Ollama embedding context window size in tokens (default: 4096, range: 512-2097152)

  • EMBEDDING_MODEL

    Embedding model name for semantic search

  • EMBEDDING_DIM

    Embedding vector dimensions

  • EMBEDDING_TIMEOUT_S

    Timeout in seconds for embedding generation API calls

  • EMBEDDING_RETRY_MAX_ATTEMPTS

    Maximum number of retry attempts for embedding generation

  • EMBEDDING_RETRY_BASE_DELAY_S

    Base delay in seconds between retry attempts (with exponential backoff)

  • EMBEDDING_MAX_CONCURRENT

    Maximum concurrent embedding generation operations (default: 3, range: 1-20)

  • ENABLE_SUMMARY_GENERATION

    Enable summary generation for stored context. Default true - server fails if dependencies not met. Set false to disable summaries.

  • SUMMARY_PROVIDER

    Summary provider: ollama (default), openai, or anthropic

  • SUMMARY_MODEL

    Summary generation model name (default: qwen3:0.6b)

  • SUMMARY_MAX_TOKENS

    Maximum output tokens for summary generation (default: 4000, range: 50-16384). Increase if summaries are truncated by reasoning models

  • SUMMARY_TIMEOUT_S

    Timeout in seconds for summary generation API calls

  • SUMMARY_RETRY_MAX_ATTEMPTS

    Maximum number of retry attempts for summary generation

  • SUMMARY_RETRY_BASE_DELAY_S

    Base delay in seconds between retry attempts (with exponential backoff)

  • SUMMARY_MAX_CONCURRENT

    Maximum concurrent summary generation operations (default: 3, range: 1-20)

  • SUMMARY_PROMPT

    Custom summarization prompt. Overrides the built-in default. Used as system message for the LLM.

  • SUMMARY_MIN_CONTENT_LENGTH

    Minimum text content length in characters to trigger summary generation (default: 500, range: 0-10000). Set to 0 to always generate.

  • SUMMARY_OLLAMA_NUM_CTX

    Ollama summary context window size in tokens (default: 32768, range: 512-2097152)

  • SUMMARY_OLLAMA_TRUNCATE

    Ollama summary truncation mode: false (default) returns error when context exceeded, true enables silent truncation

  • SUMMARY_OPENAI_REASONING_EFFORT

    Reasoning effort level for OpenAI reasoning models (default: low). Valid values vary by generation: gpt-5: low, medium, high; gpt-5.1+: none, low, medium, high, xhigh. Default low is universally valid across all generations

  • SUMMARY_ANTHROPIC_EFFORT

    Effort level for Anthropic Claude models (default: none). Valid values: max, high, medium, low. Controls inference effort (adaptive thinking)

  • ANTHROPIC_API_KEY
    secret

    Anthropic API key for summary generation

  • ENABLE_FTS

    Enable full-text search functionality

  • FTS_LANGUAGE

    Language for FTS stemming (e.g., english, german, french)

  • FTS_RERANK_WINDOW_SIZE

    Characters of context around each FTS match for reranking passage extraction (default: 750)

  • FTS_RERANK_GAP_MERGE

    Merge FTS match regions within this character distance (default: 100)

  • ENABLE_HYBRID_SEARCH

    Enable hybrid search combining FTS and semantic search with RRF fusion

  • HYBRID_RRF_K

    RRF smoothing constant for hybrid search (default 60)

  • HYBRID_RRF_OVERFETCH

    Multiplier for over-fetching results before RRF fusion (default: 2)

  • HYBRID_FTS_OR_THRESHOLD

    Minimum significant query terms to switch hybrid FTS from AND to OR logic (default: 4)

  • SEARCH_DEFAULT_SORT_BY

    Default sort order for search results: relevance (only 'relevance' supported in current version)

  • SEARCH_TRUNCATION_LENGTH

    Maximum character length for truncated text_content in search results (default: 300, range: 50-1000)

  • ENABLE_CHUNKING

    Enable text chunking for embedding generation (default: true)

  • CHUNK_SIZE

    Target chunk size in characters (default: 1500)

  • CHUNK_OVERLAP

    Overlap between chunks in characters (default: 150)

  • CHUNK_AGGREGATION

    Chunk score aggregation method: max (only 'max' supported in current version)

  • CHUNK_DEDUP_OVERFETCH

    Multiplier for over-fetching chunks before deduplication (default: 5)

  • ENABLE_RERANKING

    Enable cross-encoder reranking of search results (default: true)

  • RERANKING_PROVIDER

    Reranking provider (default: flashrank)

  • RERANKING_MODEL

    Reranking model name (default: ms-marco-MiniLM-L-12-v2)

  • RERANKING_MAX_LENGTH

    Maximum input length for reranking in tokens (default: 512)

  • RERANKING_OVERFETCH

    Multiplier for over-fetching results before reranking (default: 4)

  • RERANKING_CACHE_DIR

    Directory for caching reranking models

  • RERANKING_CHARS_PER_TOKEN

    Estimated characters per token for passage size validation (default: 4.0, range: 2.0-8.0)

  • RERANKING_INTRA_OP_THREADS

    ONNX Runtime intra-operation parallelism threads for reranking (default: 0 = auto-detect)

  • RERANKING_CPU_MEM_ARENA

    Enable ONNX Runtime CPU memory arena for reranking (default: false)

  • RERANKING_BATCH_SIZE

    Maximum passages per ONNX Runtime inference batch during reranking (default: 32)

  • EMBEDDING_PROVIDER

    Embedding provider: ollama (default), openai, azure, huggingface, or voyage

  • OPENAI_API_KEY
    secret

    OpenAI API key for OpenAI embedding provider

  • OPENAI_API_BASE

    Custom base URL for OpenAI-compatible APIs

  • OPENAI_ORGANIZATION

    OpenAI organization ID

  • AZURE_OPENAI_API_KEY
    secret

    Azure OpenAI API key

  • AZURE_OPENAI_ENDPOINT

    Azure OpenAI endpoint URL

  • AZURE_OPENAI_EMBEDDING_DEPLOYMENT_NAME

    Azure OpenAI embedding deployment name

  • AZURE_OPENAI_API_VERSION

    Azure OpenAI API version (default: 2024-02-01)

  • HUGGINGFACEHUB_API_TOKEN
    secret

    HuggingFace Hub API token for HuggingFace embedding provider

  • VOYAGE_API_KEY
    secret

    Voyage AI API key for Voyage embedding provider

  • VOYAGE_TRUNCATION

    Voyage AI truncation mode: false (default) returns error when context exceeded, true enables silent truncation

  • VOYAGE_BATCH_SIZE

    Voyage AI batch size for embedding requests

  • LANGSMITH_TRACING

    Enable LangSmith tracing

  • LANGSMITH_API_KEY
    secret

    LangSmith API key

  • LANGSMITH_PROJECT

    LangSmith project name

  • LANGSMITH_ENDPOINT

    LangSmith API endpoint URL

  • METADATA_INDEXED_FIELDS

    Comma-separated list of metadata fields to index (field:type format)

  • METADATA_INDEX_SYNC_MODE

    Index sync mode: strict (fail), auto (sync), warn (log), additive (default, add missing only)

  • MCP_TRANSPORT

    Transport mode: stdio for local, http for Docker/remote

  • FASTMCP_HOST

    HTTP bind address (use 0.0.0.0 for Docker)

  • FASTMCP_PORT

    HTTP port number

  • FASTMCP_STATELESS_HTTP

    Enable stateless HTTP mode for horizontal scaling. Enabled by default as the server has no stateful MCP features. Set to false only if you need server-side MCP session tracking.

  • DISABLED_TOOLS

    Comma-separated list of tools to disable (e.g., delete_context,update_context)

  • MCP_AUTH_TOKEN
    secret

    Bearer token for HTTP authentication (required when using SimpleTokenVerifier)

  • MCP_AUTH_CLIENT_ID

    Client ID to assign to authenticated requests

  • MCP_AUTH_PROVIDER

    Authentication provider: none (default), simple_token

  • MCP_SERVER_INSTRUCTIONS

    Custom server instructions text. Overrides built-in default. Set to empty string to disable.

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "mcp-context-server": {
      "command": "uvx",
      "args": [
        "mcp-context-server"
      ],
      "env": {
        "LOG_LEVEL": "<YOUR_LOG_LEVEL>",
        "STORAGE_BACKEND": "<YOUR_STORAGE_BACKEND>",
        "MAX_IMAGE_SIZE_MB": "<YOUR_MAX_IMAGE_SIZE_MB>",
        "MAX_TOTAL_SIZE_MB": "<YOUR_MAX_TOTAL_SIZE_MB>",
        "DB_PATH": "<YOUR_DB_PATH>",
        "POOL_MAX_READERS": "<YOUR_POOL_MAX_READERS>",
        "POOL_MAX_WRITERS": "<YOUR_POOL_MAX_WRITERS>",
        "POOL_CONNECTION_TIMEOUT_S": "<YOUR_POOL_CONNECTION_TIMEOUT_S>",
        "POOL_IDLE_TIMEOUT_S": "<YOUR_POOL_IDLE_TIMEOUT_S>",
        "POOL_HEALTH_CHECK_INTERVAL_S": "<YOUR_POOL_HEALTH_CHECK_INTERVAL_S>",
        "RETRY_MAX_RETRIES": "<YOUR_RETRY_MAX_RETRIES>",
        "RETRY_BASE_DELAY_S": "<YOUR_RETRY_BASE_DELAY_S>",
        "RETRY_MAX_DELAY_S": "<YOUR_RETRY_MAX_DELAY_S>",
        "RETRY_JITTER": "<YOUR_RETRY_JITTER>",
        "RETRY_BACKOFF_FACTOR": "<YOUR_RETRY_BACKOFF_FACTOR>",
        "SQLITE_FOREIGN_KEYS": "<YOUR_SQLITE_FOREIGN_KEYS>",
        "SQLITE_JOURNAL_MODE": "<YOUR_SQLITE_JOURNAL_MODE>",
        "SQLITE_SYNCHRONOUS": "<YOUR_SQLITE_SYNCHRONOUS>",
        "SQLITE_TEMP_STORE": "<YOUR_SQLITE_TEMP_STORE>",
        "SQLITE_MMAP_SIZE": "<YOUR_SQLITE_MMAP_SIZE>",
        "SQLITE_CACHE_SIZE": "<YOUR_SQLITE_CACHE_SIZE>",
        "SQLITE_PAGE_SIZE": "<YOUR_SQLITE_PAGE_SIZE>",
        "SQLITE_WAL_AUTOCHECKPOINT": "<YOUR_SQLITE_WAL_AUTOCHECKPOINT>",
        "SQLITE_BUSY_TIMEOUT_MS": "<YOUR_SQLITE_BUSY_TIMEOUT_MS>",
        "SQLITE_WAL_CHECKPOINT": "<YOUR_SQLITE_WAL_CHECKPOINT>",
        "SHUTDOWN_TIMEOUT_S": "<YOUR_SHUTDOWN_TIMEOUT_S>",
        "SHUTDOWN_TIMEOUT_TEST_S": "<YOUR_SHUTDOWN_TIMEOUT_TEST_S>",
        "QUEUE_TIMEOUT_S": "<YOUR_QUEUE_TIMEOUT_S>",
        "QUEUE_TIMEOUT_TEST_S": "<YOUR_QUEUE_TIMEOUT_TEST_S>",
        "CIRCUIT_BREAKER_FAILURE_THRESHOLD": "<YOUR_CIRCUIT_BREAKER_FAILURE_THRESHOLD>",
        "CIRCUIT_BREAKER_RECOVERY_TIMEOUT_S": "<YOUR_CIRCUIT_BREAKER_RECOVERY_TIMEOUT_S>",
        "CIRCUIT_BREAKER_HALF_OPEN_MAX_CALLS": "<YOUR_CIRCUIT_BREAKER_HALF_OPEN_MAX_CALLS>",
        "POSTGRESQL_CONNECTION_STRING": "<YOUR_POSTGRESQL_CONNECTION_STRING>",
        "POSTGRESQL_HOST": "<YOUR_POSTGRESQL_HOST>",
        "POSTGRESQL_PORT": "<YOUR_POSTGRESQL_PORT>",
        "POSTGRESQL_USER": "<YOUR_POSTGRESQL_USER>",
        "POSTGRESQL_PASSWORD": "<YOUR_POSTGRESQL_PASSWORD>",
        "POSTGRESQL_DATABASE": "<YOUR_POSTGRESQL_DATABASE>",
        "POSTGRESQL_POOL_MIN": "<YOUR_POSTGRESQL_POOL_MIN>",
        "POSTGRESQL_POOL_MAX": "<YOUR_POSTGRESQL_POOL_MAX>",
        "POSTGRESQL_POOL_TIMEOUT_S": "<YOUR_POSTGRESQL_POOL_TIMEOUT_S>",
        "POSTGRESQL_COMMAND_TIMEOUT_S": "<YOUR_POSTGRESQL_COMMAND_TIMEOUT_S>",
        "POSTGRESQL_MIGRATION_TIMEOUT_S": "<YOUR_POSTGRESQL_MIGRATION_TIMEOUT_S>",
        "POSTGRESQL_MAX_INACTIVE_LIFETIME_S": "<YOUR_POSTGRESQL_MAX_INACTIVE_LIFETIME_S>",
        "POSTGRESQL_MAX_QUERIES": "<YOUR_POSTGRESQL_MAX_QUERIES>",
        "POSTGRESQL_TCP_KEEPALIVES_IDLE_S": "<YOUR_POSTGRESQL_TCP_KEEPALIVES_IDLE_S>",
        "POSTGRESQL_TCP_KEEPALIVES_INTERVAL_S": "<YOUR_POSTGRESQL_TCP_KEEPALIVES_INTERVAL_S>",
        "POSTGRESQL_TCP_KEEPALIVES_COUNT": "<YOUR_POSTGRESQL_TCP_KEEPALIVES_COUNT>",
        "POSTGRESQL_STATEMENT_CACHE_SIZE": "<YOUR_POSTGRESQL_STATEMENT_CACHE_SIZE>",
        "POSTGRESQL_MAX_CACHED_STATEMENT_LIFETIME_S": "<YOUR_POSTGRESQL_MAX_CACHED_STATEMENT_LIFETIME_S>",
        "POSTGRESQL_MAX_CACHEABLE_STATEMENT_SIZE": "<YOUR_POSTGRESQL_MAX_CACHEABLE_STATEMENT_SIZE>",
        "POSTGRESQL_SSL_MODE": "<YOUR_POSTGRESQL_SSL_MODE>",
        "POSTGRESQL_SCHEMA": "<YOUR_POSTGRESQL_SCHEMA>",
        "ENABLE_SEMANTIC_SEARCH": "<YOUR_ENABLE_SEMANTIC_SEARCH>",
        "ENABLE_EMBEDDING_GENERATION": "<YOUR_ENABLE_EMBEDDING_GENERATION>",
        "OLLAMA_HOST": "<YOUR_OLLAMA_HOST>",
        "OLLAMA_AUTO_PULL": "<YOUR_OLLAMA_AUTO_PULL>",
        "OLLAMA_PULL_TIMEOUT_S": "<YOUR_OLLAMA_PULL_TIMEOUT_S>",
        "EMBEDDING_OLLAMA_TRUNCATE": "<YOUR_EMBEDDING_OLLAMA_TRUNCATE>",
        "EMBEDDING_OLLAMA_NUM_CTX": "<YOUR_EMBEDDING_OLLAMA_NUM_CTX>",
        "EMBEDDING_MODEL": "<YOUR_EMBEDDING_MODEL>",
        "EMBEDDING_DIM": "<YOUR_EMBEDDING_DIM>",
        "EMBEDDING_TIMEOUT_S": "<YOUR_EMBEDDING_TIMEOUT_S>",
        "EMBEDDING_RETRY_MAX_ATTEMPTS": "<YOUR_EMBEDDING_RETRY_MAX_ATTEMPTS>",
        "EMBEDDING_RETRY_BASE_DELAY_S": "<YOUR_EMBEDDING_RETRY_BASE_DELAY_S>",
        "EMBEDDING_MAX_CONCURRENT": "<YOUR_EMBEDDING_MAX_CONCURRENT>",
        "ENABLE_SUMMARY_GENERATION": "<YOUR_ENABLE_SUMMARY_GENERATION>",
        "SUMMARY_PROVIDER": "<YOUR_SUMMARY_PROVIDER>",
        "SUMMARY_MODEL": "<YOUR_SUMMARY_MODEL>",
        "SUMMARY_MAX_TOKENS": "<YOUR_SUMMARY_MAX_TOKENS>",
        "SUMMARY_TIMEOUT_S": "<YOUR_SUMMARY_TIMEOUT_S>",
        "SUMMARY_RETRY_MAX_ATTEMPTS": "<YOUR_SUMMARY_RETRY_MAX_ATTEMPTS>",
        "SUMMARY_RETRY_BASE_DELAY_S": "<YOUR_SUMMARY_RETRY_BASE_DELAY_S>",
        "SUMMARY_MAX_CONCURRENT": "<YOUR_SUMMARY_MAX_CONCURRENT>",
        "SUMMARY_PROMPT": "<YOUR_SUMMARY_PROMPT>",
        "SUMMARY_MIN_CONTENT_LENGTH": "<YOUR_SUMMARY_MIN_CONTENT_LENGTH>",
        "SUMMARY_OLLAMA_NUM_CTX": "<YOUR_SUMMARY_OLLAMA_NUM_CTX>",
        "SUMMARY_OLLAMA_TRUNCATE": "<YOUR_SUMMARY_OLLAMA_TRUNCATE>",
        "SUMMARY_OPENAI_REASONING_EFFORT": "<YOUR_SUMMARY_OPENAI_REASONING_EFFORT>",
        "SUMMARY_ANTHROPIC_EFFORT": "<YOUR_SUMMARY_ANTHROPIC_EFFORT>",
        "ANTHROPIC_API_KEY": "<YOUR_ANTHROPIC_API_KEY>",
        "ENABLE_FTS": "<YOUR_ENABLE_FTS>",
        "FTS_LANGUAGE": "<YOUR_FTS_LANGUAGE>",
        "FTS_RERANK_WINDOW_SIZE": "<YOUR_FTS_RERANK_WINDOW_SIZE>",
        "FTS_RERANK_GAP_MERGE": "<YOUR_FTS_RERANK_GAP_MERGE>",
        "ENABLE_HYBRID_SEARCH": "<YOUR_ENABLE_HYBRID_SEARCH>",
        "HYBRID_RRF_K": "<YOUR_HYBRID_RRF_K>",
        "HYBRID_RRF_OVERFETCH": "<YOUR_HYBRID_RRF_OVERFETCH>",
        "HYBRID_FTS_OR_THRESHOLD": "<YOUR_HYBRID_FTS_OR_THRESHOLD>",
        "SEARCH_DEFAULT_SORT_BY": "<YOUR_SEARCH_DEFAULT_SORT_BY>",
        "SEARCH_TRUNCATION_LENGTH": "<YOUR_SEARCH_TRUNCATION_LENGTH>",
        "ENABLE_CHUNKING": "<YOUR_ENABLE_CHUNKING>",
        "CHUNK_SIZE": "<YOUR_CHUNK_SIZE>",
        "CHUNK_OVERLAP": "<YOUR_CHUNK_OVERLAP>",
        "CHUNK_AGGREGATION": "<YOUR_CHUNK_AGGREGATION>",
        "CHUNK_DEDUP_OVERFETCH": "<YOUR_CHUNK_DEDUP_OVERFETCH>",
        "ENABLE_RERANKING": "<YOUR_ENABLE_RERANKING>",
        "RERANKING_PROVIDER": "<YOUR_RERANKING_PROVIDER>",
        "RERANKING_MODEL": "<YOUR_RERANKING_MODEL>",
        "RERANKING_MAX_LENGTH": "<YOUR_RERANKING_MAX_LENGTH>",
        "RERANKING_OVERFETCH": "<YOUR_RERANKING_OVERFETCH>",
        "RERANKING_CACHE_DIR": "<YOUR_RERANKING_CACHE_DIR>",
        "RERANKING_CHARS_PER_TOKEN": "<YOUR_RERANKING_CHARS_PER_TOKEN>",
        "RERANKING_INTRA_OP_THREADS": "<YOUR_RERANKING_INTRA_OP_THREADS>",
        "RERANKING_CPU_MEM_ARENA": "<YOUR_RERANKING_CPU_MEM_ARENA>",
        "RERANKING_BATCH_SIZE": "<YOUR_RERANKING_BATCH_SIZE>",
        "EMBEDDING_PROVIDER": "<YOUR_EMBEDDING_PROVIDER>",
        "OPENAI_API_KEY": "<YOUR_OPENAI_API_KEY>",
        "OPENAI_API_BASE": "<YOUR_OPENAI_API_BASE>",
        "OPENAI_ORGANIZATION": "<YOUR_OPENAI_ORGANIZATION>",
        "AZURE_OPENAI_API_KEY": "<YOUR_AZURE_OPENAI_API_KEY>",
        "AZURE_OPENAI_ENDPOINT": "<YOUR_AZURE_OPENAI_ENDPOINT>",
        "AZURE_OPENAI_EMBEDDING_DEPLOYMENT_NAME": "<YOUR_AZURE_OPENAI_EMBEDDING_DEPLOYMENT_NAME>",
        "AZURE_OPENAI_API_VERSION": "<YOUR_AZURE_OPENAI_API_VERSION>",
        "HUGGINGFACEHUB_API_TOKEN": "<YOUR_HUGGINGFACEHUB_API_TOKEN>",
        "VOYAGE_API_KEY": "<YOUR_VOYAGE_API_KEY>",
        "VOYAGE_TRUNCATION": "<YOUR_VOYAGE_TRUNCATION>",
        "VOYAGE_BATCH_SIZE": "<YOUR_VOYAGE_BATCH_SIZE>",
        "LANGSMITH_TRACING": "<YOUR_LANGSMITH_TRACING>",
        "LANGSMITH_API_KEY": "<YOUR_LANGSMITH_API_KEY>",
        "LANGSMITH_PROJECT": "<YOUR_LANGSMITH_PROJECT>",
        "LANGSMITH_ENDPOINT": "<YOUR_LANGSMITH_ENDPOINT>",
        "METADATA_INDEXED_FIELDS": "<YOUR_METADATA_INDEXED_FIELDS>",
        "METADATA_INDEX_SYNC_MODE": "<YOUR_METADATA_INDEX_SYNC_MODE>",
        "MCP_TRANSPORT": "<YOUR_MCP_TRANSPORT>",
        "FASTMCP_HOST": "<YOUR_FASTMCP_HOST>",
        "FASTMCP_PORT": "<YOUR_FASTMCP_PORT>",
        "FASTMCP_STATELESS_HTTP": "<YOUR_FASTMCP_STATELESS_HTTP>",
        "DISABLED_TOOLS": "<YOUR_DISABLED_TOOLS>",
        "MCP_AUTH_TOKEN": "<YOUR_MCP_AUTH_TOKEN>",
        "MCP_AUTH_CLIENT_ID": "<YOUR_MCP_AUTH_CLIENT_ID>",
        "MCP_AUTH_PROVIDER": "<YOUR_MCP_AUTH_PROVIDER>",
        "MCP_SERVER_INSTRUCTIONS": "<YOUR_MCP_SERVER_INSTRUCTIONS>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • store_context — Store a new context entry with text, images, metadata, and tags
  • search_context — Search context entries with optional metadata filtering and date range filtering
  • semantic_search_context — Vector similarity search for meaning-based context retrieval with cross-encoder reranking
  • fts_search_context — Full-text search with stemming, ranking, and boolean queries
  • hybrid_search_context — Combined full-text and semantic search using Reciprocal Rank Fusion
  • grep_context — Literal/regex pattern matching over stored records with ripgrep-style output
  • get_context_by_ids — Retrieve context entries by their UUIDv7 identifiers
  • delete_context — Delete context entries by ID
  • update_context — Update an existing context entry
  • navigate_context — Build an on-demand Markdown table of contents per record with optional LLM summaries
  • read_context_range — Extract a slice of a record by character range, line range, or outline node_id
  • list_threads — List all thread IDs in the database
  • get_statistics — Retrieve database statistics
  • store_context_batch — Batch store multiple context entries
  • update_context_batch — Batch update multiple context entries
  • delete_context_batch — Batch delete multiple context entries

Use cases

  • Build multi-agent systems that share persistent context across tasks and threads
  • Search through large document collections using full-text, semantic, or hybrid search to find relevant information
  • Automatically summarize stored context entries to help agents determine relevance without fetching full documents
  • Extract specific sections from long records using partial reads (character/line/outline ranges) instead of loading entire entries
  • Organize and filter context using custom metadata with 16 powerful operators and tag-based retrieval

io.github.alex-feel/mcp-context-server MCP server FAQ

What is the MCP Context Server?

It's an MCP server that provides persistent multimodal context storage for LLM agents, enabling context sharing across multiple agents working on the same task. It supports text and images, advanced search (full-text, semantic, hybrid), automatic summarization, and flexible metadata filtering.

Is it free?

The server is licensed under Elastic License 2.0 (ELv2). You can use, copy, modify, distribute, and run it freely for personal projects and inside companies of any size. The restriction applies only to providing it as a hosted/managed service to third parties without a commercial agreement.

How do I install it in Claude or Cursor?

Install via PyPI (`pip install mcp-context-server`) or use the Docker image (`ghcr.io/alex-feel/mcp-context-server:2.2.2`). For step-by-step instructions and the fastest one-command Docker bootstrap, see the Connecting to Your AI Assistant Guide in the repository documentation.

What databases does it support?

It supports SQLite (default, zero-config) and PostgreSQL (high-concurrency, production-grade). Choose via the `STORAGE_BACKEND` environment variable.

Does it require authentication?

Authentication is optional and only needed for HTTP transport deployments. For local use, no authentication is required. Bearer tokens and IdP-issued JWTs are supported when needed.

What search methods are available?

Full-text search (FTS5/tsvector with stemming and boolean queries), semantic search (vector similarity with embedding providers), hybrid search (combined FTS + semantic using Reciprocal Rank Fusion), and grep-style pattern matching (literal/regex). All support cross-encoder reranking.

README (reference)

Source of truth, from the repository.

MCP Context Server

<p align="center"> <img src=".github/images/banner.png" alt="MCP Context Server - MCP-based server providing persistent multimodal context storage for LLM agents" width="100%"> </p>

PyPI MCP Registry License: Elastic License 2.0 Ask DeepWiki

A high-performance Model Context Protocol (MCP) server providing persistent multimodal context storage for LLM agents. Built with FastMCP, this server enables seamless context sharing across multiple agents working on the same task through thread-based scoping.

[!WARNING] Upgrading from v2.x? Version 3.x.x uses a new database schema with UUIDv7 primary keys. Existing v2.x databases require a one-time data migration before they can be used with v3.x.x. The opt-in CLI mcp-context-server-migrate ships with the server.

See the Migration Guide before upgrading. Fresh installations are unaffected.

Key Features

  • Multimodal Context Storage: Store and retrieve both text and images
  • UUIDv7 Context Identifiers: Every context entry is identified by a 32-character lowercase hex UUIDv7 value, providing time-ordered, globally unique IDs with a stable lex-string ordering
  • Thread-Based Scoping: Agents working on the same task share context through thread IDs
  • Flexible Metadata Filtering: Store custom structured data with any JSON-serializable fields and filter using 16 powerful operators
  • Date Range Filtering: Filter context entries by creation timestamp using ISO 8601 format
  • Tag-Based Organization: Efficient context retrieval with normalized, indexed tags
  • Summary Generation: Optional automatic LLM-based summarization returned alongside truncated text_content in all search tool results for better agent context efficiency (enabled by default with Ollama)
  • Full-Text Search: Linguistic search with stemming, ranking, boolean queries (FTS5/tsvector), and cross-encoder reranking. Auto-enabled by default (ENABLE_FTS=auto); needs no extra dependencies
  • Semantic Search: Vector similarity search for meaning-based retrieval with cross-encoder reranking. Auto-enabled by default (ENABLE_SEMANTIC_SEARCH=auto) whenever an embedding provider is available (embedding generation is on by default)
  • Hybrid Search: Combined FTS + semantic search using Reciprocal Rank Fusion (RRF) with cross-encoder reranking. Auto-enabled by default (ENABLE_HYBRID_SEARCH=auto) whenever at least one of full-text or semantic search is available
  • Server-Side Grep: Literal/regex, line-oriented, unranked pattern matching over stored records (grep_context) — the precise-locate complement to full-text/semantic search, with ripgrep-style output modes and bounded results. Auto-enabled by default (ENABLE_GREP_CONTEXT=auto), pure-Python so it behaves identically on SQLite and PostgreSQL
  • Record Navigation (index_tree): navigate_context builds an on-demand Markdown-heading table of contents per record, with the entry summary as the root node; optional per-node LLM summaries (on by default) enrich each section. Pair with read_context_range to extract any section
  • Partial Reads: read_context_range returns a slice of one record by character range, line range, or outline node_id — so an agent can read only the relevant span of a long record instead of the whole thing
  • Cross-Encoder Reranking: Automatic result refinement using FlashRank cross-encoder models for improved search precision (enabled by default)
  • Embedding Compression (default ON): Reduces embedding storage by approximately 8x out of the box in v3.0.0. Bit-packed compressed vectors keep semantic and hybrid search working without changes to the tool surface, and the read path bypasses the pgvector >2000-dimension HNSW limit. Set ENABLE_EMBEDDING_COMPRESSION=false to opt out and keep fp32 storage. See the Embedding Compression Guide
  • Multiple Database Backends: Choose between SQLite (default, zero-config) or PostgreSQL (high-concurrency, production-grade)
  • High Performance: WAL mode (SQLite) / MVCC (PostgreSQL), strategic indexing, and async operations
  • MCP Standard Compliance: Works with Claude Code, LangGraph, and any MCP-compatible client
  • Production Ready: Comprehensive test coverage, type safety, and robust error handling

Connecting to Your AI Assistant

The fastest way to connect the MCP Context Server to Claude Code is the one-command Docker bootstrap.

For step-by-step instructions, prerequisites, troubleshooting, and update/uninstall commands, see the Connecting to Your AI Assistant Guide.

Environment Configuration

The server is fully configured via environment variables, supporting core settings, transport, authentication, embedding providers, summary generation, search features, database tuning, and more. Variables can be set in your MCP client configuration, in a .env file, or directly in the shell.

For the complete reference of all environment variables with types, defaults, constraints, and descriptions, see the Environment Variables Reference.

Summary Generation

Summary generation automatically creates concise LLM-based summaries for each stored context entry. Summaries are returned in the summary field of all search tool results alongside truncated text_content, providing dense, informative summaries that help agents determine relevance without fetching full entries.

For detailed instructions including all providers (Ollama, OpenAI, Anthropic), model selection, and custom prompt configuration, see the Summary Generation Guide.

Semantic Search

Semantic search is auto-enabled by default (ENABLE_SEMANTIC_SEARCH=auto): the semantic_search_context tool registers automatically whenever an embedding provider is available (embedding generation is on by default), and skips quietly otherwise. For detailed instructions on the multiple embedding providers (Ollama, OpenAI, Azure, HuggingFace, Voyage) and how to control the toggle explicitly, see the Semantic Search Guide.

Full-Text Search

Full-text search is auto-enabled by default (ENABLE_FTS=auto) and needs no extra dependencies, using the built-in database FTS engine (FTS5 on SQLite, tsvector on PostgreSQL). For linguistic processing, stemming, ranking, and boolean queries, see the Full-Text Search Guide.

Hybrid Search

Hybrid search is auto-enabled by default (ENABLE_HYBRID_SEARCH=auto): the hybrid_search_context tool registers automatically whenever at least one of full-text or semantic search is available. For combined FTS + semantic search using Reciprocal Rank Fusion (RRF), see the Hybrid Search Guide.

Metadata Filtering

For comprehensive metadata filtering including 16 operators, nested JSON paths, and performance optimization, see the Metadata Guide.

Database Backends

The server supports multiple database backends, selectable via the STORAGE_BACKEND environment variable. SQLite (default) provides zero-configuration local storage perfect for single-user deployments. PostgreSQL offers high-performance capabilities with 10x+ write throughput for multi-user and high-traffic deployments.

For detailed configuration instructions including PostgreSQL setup with Docker, Supabase integration, connection methods, and troubleshooting, see the Database Backends Guide.

API Reference

The MCP Context Server exposes 16 MCP tools for context management:

Core Operations: store_context, search_context, get_context_by_ids, delete_context, update_context, list_threads, get_statistics

Search Tools: semantic_search_context, fts_search_context, hybrid_search_context

Navigation Tools (locate / navigate / extract): grep_context, navigate_context, read_context_range

Batch Operations: store_context_batch, update_context_batch, delete_context_batch

For complete tool documentation including parameters, return values, filtering options, and examples, see the API Reference. For when to use grep vs full-text vs semantic search, the index_tree, and partial reads, see Grep, Navigation & Partial Reads.

Docker Deployment

For production deployments with HTTP transport and container orchestration, Docker Compose configurations are available for SQLite, PostgreSQL, and external PostgreSQL (Supabase). See the Docker Deployment Guide for setup instructions and client connection details.

Kubernetes Deployment

For Kubernetes deployments, a Helm chart is provided with configurable values for different environments. See the Helm Deployment Guide for installation instructions, or the Kubernetes Deployment Guide for general Kubernetes concepts.

Authentication

For HTTP transport deployments requiring authentication, see the Authentication Guide for bearer token and IdP-issued JWT configuration.

Getting Help

License

MCP Context Server is licensed under the Elastic License 2.0 (ELv2).

In short: you may use, copy, modify, distribute, and run the software freely and at no cost — for personal projects, inside companies of any size, and as part of commercial work. The one thing you may not do without a commercial agreement is provide the software to third parties as a hosted or managed service that gives users access to any substantial set of its features or functionality (for example, a cloud "memory for agents" offering built on it).

See Commercial Licensing for plain-language examples of what is and is not permitted, and contact alexfeel@protonmail.com for commercial licensing, including hosted or managed service rights.

Releases up to and including v2.2.2 were published under the MIT License and remain available under it; the Elastic License 2.0 applies from v3.0.0 onward.

<!-- mcp-name: io.github.alex-feel/mcp-context-server -->

Related MCP servers

THThe Game Crafter logo

Design, manage, and price tabletop games on The Game Crafter via MCP.

2
TypeScript
MIT
View repository →

Inspect, drive, and debug Ren'Py games with live editing, screenshots, and state control.

9
Python
MIT
View repository →

Personal finance by conversation: expenses, receipts, statement import, budgets, net worth.

1
TypeScript
MIT
View repository →
WEWebReaper logo

WebReaper

Active

AI-native web scraper: scrape, crawl, and map any site to clean markdown via CLI or MCP server.

140
C#
MIT
View repository →

Unofficial Plan to Eat client: recipes, meal planner, leftovers, freezer, and shopping list.

0
TypeScript
MIT
View repository →
OPOpenProject logo

OpenProject

Maintained

Manage OpenProject projects, work packages, files, relations, boards, users, and notifications.

0
TypeScript
MIT
View repository →