PluginBench
MCP Server
Active

LatentGate MCP Server

io.github.KathanModh259/latent-gate

Compress images, text, and docs locally via Ollama, send compact payloads to any LLM API—save 66% tokens and costs.

What is the LatentGate MCP server?

LatentGate is an MCP server that compresses images, text, conversations, and RAG documents locally using Ollama, then sends only compact, fact-preserving payloads to cloud LLMs. It reduces token usage by ~66% on realistic workloads while working offline with no API keys required for compression.

LatentGate optimizes LLM API costs by preprocessing large inputs locally before sending them to Claude, GPT-4o, or other cloud models. It uses a VL-JEPA-inspired pipeline to compress images and text deterministically, measures savings with a real tokenizer, and verifies no facts are lost. Works as an MCP server in Claude Desktop, Cursor, and other AI tools, or as a REST API and Python library for custom integrations.

How to install LatentGate

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "latent-gate": {
      "command": "uvx",
      "args": [
        "latent-gate",
        "--with",
        "mcp>=1.0,<2",
        "--with",
        "tiktoken>=0.7.0",
        "mcp"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • read_file_optimized — Read a text file and return its optimized, compressed form (for files you read, not edit)
  • optimize_text — Optimize text you already have, optionally toward a question or max_tokens budget
  • count_tokens — Count tokens using tiktoken o200k_base
  • compress_image — Describe an image locally as a ~150-token scene payload (requires Ollama)
  • compress_text — Compress long prompts locally with fact-checking (requires Ollama)
  • compress_conversation — Compress conversation history locally (requires Ollama)
  • compress_documents — Compress RAG documents locally (requires Ollama)
  • get_stats — Retrieve session statistics and cost tracking

Use cases

  • Reduce token costs on large image uploads by 66% by compressing locally before sending to GPT-4o or Claude
  • Summarize 400-line error logs from ~14,400 tokens down to ~100 tokens for faster context window usage
  • Compress long conversation histories and RAG document sets before sending to cloud LLMs
  • Process video files with selective decoding to skip redundant frames and reduce API calls by ~2.85x
  • Track and analyze LLM API costs over time with persistent SQLite analytics and exportable reports

LatentGate MCP server FAQ

What is LatentGate?

LatentGate is a token optimizer that compresses images, text, conversations, and documents locally via Ollama before sending them to cloud LLMs like GPT-4o or Claude. It reduces token usage by ~66% while preserving facts, working offline with no API keys required for compression.

Is LatentGate free?

Yes, the core compression is free and runs locally. You only pay for cloud LLM API calls (OpenAI, Anthropic, Google, etc.), and LatentGate reduces those costs by ~66% on average.

How do I install LatentGate in Claude Desktop?

Add this to your claude_desktop_config.json and restart: {"mcpServers": {"latent-gate": {"command": "uvx", "args": ["--from", "latent-gate[mcp,tokens]", "latent-gate-mcp"]}}}. Requires uv and Ollama for full features.

Do I need Ollama?

For text-only optimization (read_file_optimized, optimize_text, count_tokens), no. For image compression and advanced text compression, yes—Ollama runs locally and is free.

What LLM providers does LatentGate support?

OpenAI, Anthropic, Google, Groq, DeepSeek, Together, Azure, AWS Bedrock, Ollama, and any OpenAI-compatible endpoint.

How do I measure token savings?

Run `latent-gate --optimizer-benchmark` to see real token counts before/after compression on a dev corpus. The API and Python library also track costs persistently in SQLite.

README (reference)

Source of truth, from the repository.

<div align="center">

LatentGate

Process Locally. Send Smart. Pay Less.

A VL-JEPA-inspired pipeline that compresses images, text, conversations, and RAG documents locally via Ollama, then sends only compact payloads to any LLM API — every saving measured with a real tokenizer and checked for lost facts.

Python 3.10+ License: Proprietary Version PRs Welcome Ollama MCP Prometheus Grafana k6 Vercel CI Downloads

Use in Claude | Quick Start | Python API | REST API | Monitoring | Load Testing | Deployment | AI Tools | Benchmarks | Contributing

</div> <!-- mcp-name: io.github.KathanModh259/latent-gate -->

The Problem

Every time you send an image or long prompt to GPT-4o / Claude / Gemini, you burn 1,000+ tokens on processing that could happen locally for free.

Traditional:  Image -> Cloud LLM (1,200 tokens) -> Answer
LatentGate:   Image -> Local Ollama (FREE) -> Cloud LLM (200 tokens) -> Answer

Features

FeatureDescription
Local-FirstVision and text compression runs on Ollama (free, no API key needed)
Token OptimizerDeterministic, fact-preserving compression: 66% fewer tokens on a realistic dev corpus in ~1ms (benchmark)
MCP ServerWorks with Claude Desktop, Cursor, Cline, Continue, Zed
Selective DecodingFor video, only call API when scene changes (~2.85x fewer calls) with cosine similarity
Text CompressionLong prompts, conversations, RAG docs compressed locally
Speed OptimizedConnection pooling, model preloading, parallel processing
Multi-ProviderOpenAI, Anthropic, Google, Groq, DeepSeek, Together, Azure, AWS Bedrock, Ollama, or any OpenAI-compatible endpoint
REST APIFastAPI server for web application integration
Video ProcessingDirect video file input with automatic frame extraction
Cost TrackingPersistent cost tracking with SQLite analytics and exportable reports
Async SupportNon-blocking async methods for FastAPI, aiohttp, etc.
Streaming ResponsesStream responses from remote LLMs
Config PersistenceYAML/TOML config files with environment variable overrides
Structured LoggingJSON-formatted logging with rotation and correlation IDs
Docker SupportDockerfile and docker-compose for easy deployment
Plugin SystemCustom processors for domain-specific compression
Multi-LanguageSupport for 30+ languages with automatic detection

Use it in Claude

LatentGate plugs into Claude as an MCP server. Its optimizer tools work offline, with no Ollama and no API key — Claude reads big logs, JSON dumps and docs through it and spends a fraction of the context. A 400-line error log goes from 14,400 to ~100 tokens.

Requires uv (uvx fetches LatentGate from PyPI on first run). The first run downloads dependencies, which can exceed Claude's 30-second MCP startup limit on a slow connection; run this once beforehand (later starts take ~2s):

uvx --from "latent-gate[mcp,tokens]" latent-gate --optimizer-benchmark

Claude Code — plugin (MCP server + a skill that tells Claude when to use it):

/plugin marketplace add KathanModh259/latent-gate
/plugin install latent-gate@latent-gate

Claude Code — MCP server only:

claude mcp add latent-gate -- uvx --from "latent-gate[mcp,tokens]" latent-gate-mcp

Claude Desktop — add to claude_desktop_config.json and restart:

{
  "mcpServers": {
    "latent-gate": {
      "command": "uvx",
      "args": ["--from", "latent-gate[mcp,tokens]", "latent-gate-mcp"]
    }
  }
}

Then just ask: "Read logs/app.log with latent-gate and tell me why orders fail."

ToolNeeds OllamaWhat it does
read_file_optimizednoRead a text file and return its optimized form (use for files you read, not files you edit)
optimize_textnoOptimize text you already have, optionally toward a question / max_tokens budget
count_tokensnoCount tokens (tiktoken o200k_base)
compress_imageyesDescribe an image locally as a ~150-token scene payload
compress_text / compress_conversation / compress_documentsyesOptimizer + fact-checked local-LLM compression
get_statsyesSession statistics

Quick Start

Install

# Core install
pip install latent-gate

# With MCP server (for Claude Desktop, Cursor, Cline, etc.)
pip install latent-gate[mcp]

# With API server (for web applications)
pip install latent-gate[api]

# With video processing
pip install latent-gate[video]

# With embedding-based similarity (more accurate selective decoding)
pip install latent-gate[embeddings]

# With LangChain integration
pip install latent-gate[langchain]

# With AWS Bedrock support
pip install latent-gate[bedrock]

# Exact token counts via tiktoken (otherwise a calibrated ~6%-error estimate)
pip install latent-gate[tokens]

# With all features
pip install latent-gate[all]

Pull Ollama Models

ollama pull llava:7b      # Vision model (required for image queries)
ollama pull llama3:8b     # Text model (required for text compression & prediction)

One-Command Quickstart

chmod +x scripts/quickstart.sh
./scripts/quickstart.sh

This starts everything: Ollama → pulls models → API server → website. See scripts/quickstart.sh for options like --no-pull, --no-website, --port 9000.

CLI Usage

# Image query
latent-gate photo.jpg "What is in this image?" --provider ollama -v

# Text compression
latent-gate --text "Your long prompt here..." --provider ollama -v

# Text from file
latent-gate --text-file prompt.txt --provider openai -v

# Image + Text combined
latent-gate photo.jpg "Analyze" --text "Extra context..." -v

# Full JSON output
latent-gate photo.jpg "Describe" --json -v

# Compress only, deterministically (no LLM call, ~1ms, reproducible)
cat prompt.txt | latent-gate --text-file - --compress-only --deterministic --level balanced

# Reproduce the token-savings benchmark (no Ollama needed)
latent-gate --optimizer-benchmark

# Production benchmark
latent-gate --benchmark --benchmark-output reports/benchmark.json

# Start API server (requires: pip install latent-gate[api])
latent-gate-api

Production Hardening

For API deployments, restrict direct image-path reads to trusted directories:

set LATENTGATE_ALLOWED_IMAGE_ROOTS=C:\safe-images;D:\datasets
latent-gate-api

Benchmark before releases so speed and savings are measured, not guessed:

latent-gate --benchmark --json

Python API

Image Query

from latent_gate import LatentGatePipeline, PipelineConfig

config = PipelineConfig(
    vision_model="llava:7b",
    predictor_model="llama3:8b",
    remote_provider="openai",
    remote_model="gpt-4o-mini",
)

with LatentGatePipeline(config) as pipeline:
    result = pipeline.query("photo.jpg", "What is in this image?")

    print(result["answer"])
    print(f"Tokens sent: ~{result['tokens_estimated']}")
    print(f"Timing: {result['timing']}")

Text Compression

# Long prompt compression
result = pipeline.query_text("Your 500-word prompt here...", mode="auto")

# Conversation history compression
messages = [
    {"role": "user", "content": "Help me with Kubernetes setup"},
    {"role": "assistant", "content": "Sure! What's your target configuration?"},
    {"role": "user", "content": "3 nodes, t3.large, us-east-1 with autoscaling"},
]
result = pipeline.query_conversation(messages, "Now give me the setup commands")

# RAG document compression
documents = ["doc1 text...", "doc2 text...", "doc3 text..."]
result = pipeline.query_documents(documents, "How do I implement JWT refresh?")

# Universal (auto-detect input type)
result = pipeline.query_universal(text="Explain this code...", image="screenshot.png")

Batch Processing

# Sequential with selective decoding (skips redundant API calls)
results = pipeline.query_batch(image_paths, "Describe each scene")

# Parallel processing
results = pipeline.query_batch(image_paths, "Describe each scene", parallel=True, max_workers=4)

# Text batch
results = pipeline.query_batch_texts(text_list, question="Summarize each")

Streaming

# Stream image query
for token in pipeline.query_stream("photo.jpg", "Describe this"):
    print(token, end="", flush=True)

# Stream text query
for token in pipeline.query_text_stream("Long prompt...", mode="compress"):
    print(token, end="", flush=True)

REST API

Start Server

# Default (0.0.0.0:8000)
latent-gate-api

# Custom host/port
# Linux/macOS:
LATENTGATE_HOST=127.0.0.1 LATENTGATE_PORT=9000 latent-gate-api

# Windows PowerShell:
$env:LATENTGATE_HOST="127.0.0.1"; $env:LATENTGATE_PORT="9000"; latent-gate-api

# Windows CMD:
set LATENTGATE_HOST=127.0.0.1 && set LATENTGATE_PORT=9000 && latent-gate-api

Endpoints

MethodEndpointDescription
GET/healthHealth check (Ollama connection status)
GET/statsSession usage statistics
POST/query/imageImage query
POST/query/textText compression
POST/query/conversationConversation compression
POST/query/documentsRAG document compression
POST/query/universalAuto-detect input type
POST/query/image/uploadUpload image for query

Example Requests

import requests

# Image query
response = requests.post("http://localhost:8000/query/image", json={
    "image_path": "photo.jpg",
    "question": "What is in this image?"
})

# Text query
response = requests.post("http://localhost:8000/query/text", json={
    "text": "Your long prompt here...",
    "question": "Summarize this",
    "mode": "auto"  # auto | compress | summarize | condense | code
})

# Health check
response = requests.get("http://localhost:8000/health")
print(response.json())  # {"status": "healthy", "ollama_connected": true, ...}

Async Support

import asyncio
from latent_gate import AsyncLatentGatePipeline, PipelineConfig

async def main():
    async with AsyncLatentGatePipeline() as pipeline:
        # Single queries
        result = await pipeline.query("photo.jpg", "What is this?")
        result = await pipeline.query_text("Long prompt...")

        # Concurrent batch processing
        results = await pipeline.query_many_images(
            ["img1.jpg", "img2.jpg", "img3.jpg"],
            "Describe each image",
            max_concurrent=3,
        )

asyncio.run(main())

Video Processing

from latent_gate import LatentGatePipeline, PipelineConfig, VideoProcessor, VideoConfig

config = PipelineConfig(
    vision_model="llava:7b",
    remote_provider="ollama",
    remote_model="llama3:8b",
)

video_config = VideoConfig(
    fps=1.0,            # Extract 1 frame per second
    max_frames=100,     # Max frames to process
    quality=95,         # JPEG quality
    resize_width=640,   # Resize frames (saves processing time)
)

with VideoProcessor(config, video_config) as processor:
    result = processor.process_video("video.mp4", "Describe the action")

    print(f"Frames processed: {result['total_frames']}")
    print(f"Unique scenes: {result['statistics']['unique_scenes']}")
    print(f"Skip rate: {result['statistics']['skip_rate']}")

Configuration

Config File

# latentgate.yaml
ollama_base_url: http://localhost:11434
vision_model: llava:7b
predictor_model: llama3:8b
remote_provider: openai
remote_model: gpt-4o-mini
selective_decoding: true
similarity_threshold: 0.85
use_embeddings: true
enable_caching: true
temperature: 0.1
request_timeout: 120
track_costs: true
cost_db_path: "latentgate_costs.db"
from latent_gate import get_config, LatentGatePipeline

config = get_config("latentgate.yaml")
with LatentGatePipeline(config) as pipeline:
    result = pipeline.query("photo.jpg", "Describe this")

Environment Variables

VariableDescriptionDefault
OPENAI_API_KEYOpenAI API key-
ANTHROPIC_API_KEYAnthropic API key-
GOOGLE_API_KEYGoogle API key-
LATENTGATE_REMOTE_PROVIDEROverride remote provideropenai
LATENTGATE_REMOTE_MODELOverride remote modelprovider default (e.g. gpt-4o-mini, claude-sonnet-5)
LATENTGATE_VISION_MODELOverride vision modelllava:7b
LATENTGATE_LOG_LEVELLog levelINFO
LATENTGATE_LOG_FILELog file path-
LATENTGATE_LOG_JSONJSON log formatfalse
LATENTGATE_TRACK_COSTSEnable cost analyticsfalse
LATENTGATE_COST_DB_PATHPath to SQLite DB.latentgate_costs.db
LATENTGATE_API_KEYRequire Authorization: Bearer <key> on the API (incl. /compress; WebSocket clients may pass ?api_key=)unset (open)
LATENTGATE_COMPRESSION_LEVELToken optimizer level: lossless, balanced, aggressivebalanced
LATENTGATE_COMPRESSION_STRATEGYauto (optimizer + fact-checked local LLM) or deterministicauto
LATENTGATE_TARGET_TOKEN_BUDGETMax tokens for a compressed prompt (0 = no budget)0
LATENTGATE_MAX_OUTPUT_TOKENSMax tokens the cloud model may generate per answer4096
LATENTGATE_PRELOADWarm Ollama models in the background at API startuptrue
LATENTGATE_MAX_CONCURRENT_REQUESTSMax concurrent pipeline calls in the API server3
LATENTGATE_CORS_ORIGINSComma-separated allowed CORS originshttp://localhost:5173

Save Config

from latent_gate import PipelineConfig, save_config

config = PipelineConfig(remote_provider="anthropic", remote_model="claude-sonnet-5")
save_config(config, "my_config.yaml")

Docker

# Start full stack (API + Ollama)
docker compose up -d

# Pull models first (one-time setup)
docker compose --profile setup up ollama-init

# Start with monitoring (Prometheus + Grafana)
docker compose --profile monitoring up -d

# Build and run manually
docker build -t latent-gate .
docker run -p 8000:8000 latent-gate

The docker-compose setup includes:

  • latent-gate API server (port 8000)
  • Ollama local LLM server (port 11434)
  • ollama-init container that auto-pulls required models (profile: setup)
  • Prometheus metrics collector on port 9090 (profile: monitoring)
  • Grafana dashboard on port 3000 (profile: monitoring, credentials: admin/latentgate)

Monitoring Stack

Start with monitoring:

docker compose --profile monitoring up -d

Access:

The LatentGate dashboard auto-loads in Grafana with 9 panels covering request rate, latency percentiles (p50/p95/p99), token savings, error rates, pipeline health, and endpoint breakdown.

Customize via environment variables:

  • GRAFANA_ADMIN — Grafana admin username (default: admin)
  • GRAFANA_PASSWORD — Grafana password (default: latentgate)
  • GRAFANA_ANONYMOUS — Enable anonymous access (default: true)
  • LATENTGATE_ENABLE_METRICS — Enable Prometheus metrics (default: true)

AI Coding Tool Integration (MCP)

LatentGate works as a Model Context Protocol (MCP) server with every major AI coding tool. Your AI assistant automatically compresses images, long prompts, and documents before they reach the cloud model.

Supported Tools

ToolStatusSetup
VS Code / CopilotSupportedExtension
Claude DesktopSupportedMCP Config
Claude Code (CLI)SupportedSkill
CursorSupportedMCP Config
Cline (VS Code)SupportedMCP Config
Continue.devSupportedMCP Config
Zed EditorSupportedMCP Config

VS Code Extension

code --install-extension KathanModh259.latent-gate-vscode

Features:

  • Right-click any image to compress with LatentGate
  • Select text and press Ctrl+Shift+Alt+C to compress
  • Cost dashboard in activity bar
  • Auto-configures MCP for Copilot Chat
  • Status bar showing token savings

MCP Setup

For Claude, see Use it in Claude. For Cursor, Cline, Continue, Zed and other MCP clients, use the same server command:

{
  "mcpServers": {
    "latent-gate": {
      "command": "uvx",
      "args": ["--from", "latent-gate[mcp,tokens]", "latent-gate-mcp"]
    }
  }
}

Or install it into your environment (pip install "latent-gate[mcp,tokens]") and use "command": "latent-gate-mcp". Note the command is latent-gate-mcp — plain latent-gate is the CLI and will not speak MCP. The image and compress_* tools additionally need ollama pull llava:7b and ollama pull phi3:mini.

See integrations/ folder for detailed setup guides per tool.


Speed Optimizations

OptimizationWhat It DoesImpact
Connection PoolingReuses HTTP connections via requests.Session~30-50% faster per call
Model PreloadingWarms up Ollama models on init (keep_alive)Eliminates 5-15s cold start
Shorter PromptsOptimized extraction prompts produce fewer output tokens~20% faster generation
3-Tier JSON ParsingFast parse, extract from text, LLM fallbackAvoids slow LLM call 90% of time
Parallel ProcessingImage and text processed simultaneously via ThreadPool~40% faster combined queries
Content-Hash CachingDisk cache for repeated imagesInstant on cache hit
Selective DecodingCosine similarity skips redundant API calls~2.85x fewer calls

Cost Benchmarks

Image Queries (by provider, estimated)

Estimates from each provider's published image-token formulas versus a ~150-token local description.

ProviderRaw Image TokensLatentGate TokensSavings
OpenAI GPT-4o (high detail)~1,105~150~86%
Claude 3.5 Sonnet (1MP image)~1,334~150~89%
Gemini 2.0 Flash~258~150~42%

Text: measured, reproducible

Deterministic optimizer on built-in realistic inputs, counted with tiktoken (o200k_base). Facts kept = share of numbers, identifiers, URLs, file names and code preserved verbatim. Reproduce with latent-gate --optimizer-benchmark (no Ollama or API key needed).

InputTokenslosslessbalanced (default)aggressive
Pretty-printed API JSON1,407−33%, facts 100%−33%, facts 100%−33%, facts 100%
60 repeated log lines + trace2,2350%−92% (ranges kept)−92%
Prompt pasted 3×121−61%, facts 100%−61%, facts 100%−61%, facts 100%
Verbose spec with 6 requirements1270%−20%, facts 100%−51%, facts 90%
Code review request57−5%−16%, facts 100%−28%, facts 100%
5 RAG chunks + question1900%−58% (2 relevant docs kept)−58%
Total4,137−13%, facts 100%−66%−67%

Logs lose individual ids/timestamps when folded (the fold keeps first, last and value ranges), and RAG drops facts from documents irrelevant to the question — both by design.

How it works (safest stage first; see latent_gate/optimizer.py):

  1. Protect code blocks, inline code, URLs and quoted strings — restored byte-for-byte
  2. Lossless: whitespace/Unicode cleanup, JSON minification (values untouched), duplicate folding
  3. Log folding: runs of log lines differing only in numbers/ids → first, last, and ranges
  4. Filler: pure pleasantries ("Hi!", "Thanks in advance!") and hedging phrases removed
  5. Selection (only over a budget, or aggressive): BM25 question-relevance + requirement cues, original order preserved, the user's actual ask is never dropped

Guarantees: output never has more tokens than input; same input → same output (so provider prompt caching keeps working); when a local LLM rewrite is used it must be smaller and keep every fact, otherwise the deterministic result is sent.

from latent_gate import optimize

r = optimize(long_prompt, question="What failed?", level="balanced", max_tokens=2000)
print(r.optimized_tokens, r.savings_pct, r.stages)

Video

Selective decoding skips remote calls for frames similar to the previous one (~2.85x fewer calls on typical footage).

At Scale (10,000 image queries with gpt-4o-mini, estimated)

MetricTraditionalLatentGateSavings
Input tokens12,000,0002,000,00010M tokens
Cost$1.80$0.30$1.50 (83%)

Cost Tracking

from latent_gate import CostTracker

tracker = CostTracker()
tracker.record_usage(
    query_type="image",
    provider="openai",
    model="gpt-4o-mini",
    input_tokens=150,
    output_tokens=200,
    tokens_saved=1000,
    compression_ratio=6.7,
    latency_ms=1500,
)

# Session statistics
stats = tracker.get_session_statistics()
print(f"Total cost: ${stats['total_cost']:.4f}")
print(f"Tokens saved: {stats['total_tokens_saved']}")

# Cost projection
projection = tracker.get_cost_projection(
    daily_queries=1000,
    provider="openai",
    model="gpt-4o-mini"
)
print(f"Monthly savings: ${projection['savings']['monthly']:.2f}")

# Export report
tracker.export_report("usage_report.json", fmt="json")
tracker.export_report("usage_report.csv", fmt="csv")

Multi-Language Support

from latent_gate import detect_language, MultiLanguageProcessor

# Detect language
lang = detect_language("Esto es un texto en español")
print(f"Detected: {lang.name} ({lang.confidence:.0%})")

# Process with auto-translation to English
processor = MultiLanguageProcessor()
text, lang_info = processor.process("Texto en español para analizar")
print(f"Language: {lang_info.name}, Translated: {text[:100]}...")

Project Structure

latent-gate/
├── latent_gate/
│   ├── __init__.py           # Package exports and version
│   ├── config.py             # PipelineConfig dataclass
│   ├── config_loader.py      # YAML/TOML/JSON config loading
│   ├── payload.py            # SemanticPayload (compact representation)
│   ├── text_processor.py     # TextPayload + TextProcessor (local compression)
│   ├── local_processor.py    # X-Encoder + Predictor (Ollama vision pipeline)
│   ├── remote_decoder.py     # Y-Decoder (OpenAI, Anthropic, Google, Ollama)
│   ├── selective_decoder.py  # Cosine/Jaccard similarity for skip decisions
│   ├── fast_client.py        # Connection pooling + model preloading
│   ├── cache.py              # Content-hash disk cache
│   ├── pipeline.py           # LatentGatePipeline (main orchestrator)
│   ├── async_pipeline.py     # AsyncLatentGatePipeline
│   ├── video_processor.py    # Video frame extraction + batch processing
│   ├── cost_tracker.py       # SQLite-based cost analytics
│   ├── mcp_server.py         # MCP server (Model Context Protocol)
│   ├── api_server.py         # FastAPI REST server
│   ├── cli.py                # Command-line interface
│   ├── logging_config.py     # Structured logging with rotation
│   ├── plugin_system.py      # Custom processor plugins
│   └── multilang.py          # Multi-language detection and translation
├── integrations/
│   ├── agent_skills/         # Prompt compression skill + scripts
│   ├── vscode-extension/     # VS Code extension source
│   ├── cursor/               # Cursor rules and MCP config
│   ├── continue_dev/         # Continue.dev config
│   ├── langchain/            # LangChain integration wrapper
│   ├── llamaindex/           # LlamaIndex retriever integration
│   └── openai_functions/     # OpenAI/Anthropic function schemas
├── tests/                    # 240+ tests (unit + integration)
├── website/                  # React-based analytics dashboard & landing page
├── deployments/              # Kubernetes Helm configs
├── .github/workflows/        # CI + publish workflows
├── Dockerfile
├── docker-compose.yml
├── pyproject.toml
└── requirements.txt

Community


Contributing

Contributions welcome! See CONTRIBUTING.md.

Development Setup

git clone https://github.com/KathanModh259/latent-gate.git
cd latent-gate
python -m venv .venv
source .venv/bin/activate       # Linux/macOS
.venv\Scripts\Activate.ps1      # Windows

pip install -e ".[dev]"

Run Tests

pytest tests/ -v

Priority Areas

  • Additional vision model support (Florence-2, InternVL, Qwen-VL)
  • Custom similarity plugins for domain-specific use cases
  • WebSocket support for real-time streaming
  • Advanced cost analytics and optimization suggestions
  • Plugin development for specialized industries
  • Test coverage improvements
  • Documentation and examples

Citation

@software{latentgate2026,
  author  = {Kathan Modh},
  title   = {LatentGate: Local-First Semantic Compression Pipeline},
  year    = {2026},
  url     = {https://github.com/KathanModh259/latent-gate},
  version = {1.3.0}
}

Inspired by VL-JEPA (Meta FAIR, 2025).


License

Custom Proprietary License — see LICENSE.


<div align="center">

Built by Kathan Modh

Process locally. Send smart. Pay less.

</div>

Related MCP servers

Simple MCP server that tells the time

0
JavaScript
MIT
View repository →

Git-aware safe file ops for AI agents: delete anything except git-tracked source code.

0
Go
MIT
View repository →

Re-check pinned sources, package and cite. Live Federation MCP. No account.

0
Astro
View repository →
ABAbscissa logo

Abscissa

Maintained

Safety-aware MCP server for Linear issues, projects, cycles, and dependencies.

1
Python
MIT
View repository →
FOfootnote-mcp logo

Source-grounded web research: search, extraction, verification, and browser automation.

1
Python
MIT
View repository →

Web search, content extraction, PDF parsing, plus datetime and geolocation tools. No API keys.

View repository →