PluginBench
MCP Server
Active

Visual Memory MCP MCP Server

io.github.putervision/vision-memory-mcp

Local visual UI cache for AI agents using perceptual hashing and CLIP to reduce vision token costs.

What is the Visual Memory MCP MCP server?

The Visual Memory MCP server is a local-first Model Context Protocol server that caches UI screenshots using perceptual hashing, local CLIP embeddings, and transition graphs to eliminate repetitive vision LLM calls. It provides AI coding assistants with fast visual state recognition, element grounding, video ingestion, and evidence pack generation while maintaining 100% local privacy.

Visual Memory MCP reduces vision token overhead and latency by caching UI states locally using a 4-tier retrieval pipeline: in-memory cache, perceptual hash lookup, semantic vector search, and vision LLM fallback. It's designed for frontend development, UI testing, and multimodal AI workflows where visual state consistency matters. Features include WebM/MP4 video ingestion, element grounding with CSS selectors, visual regression testing via design specs, and cryptographic evidence packs for audit trails.

How to install Visual Memory MCP

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "vision-memory-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "@putervision/vision-memory-mcp"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • analyze_screenshot — Analyzes screenshots using L1/L2 perceptual dHash and accessibility tree parsing for single or batch processing.
  • recall_memory — Performs text and image semantic vector search across cached visual states.
  • get_session_context — Returns aggregated cache hit metrics, recent states, and sub-1KB compact slice exports.
  • predict_next_action — Predicts deterministic CSS selectors and bounding coordinates for UI element interaction.
  • record_outcome — Records UI action transitions and visual blockers.
  • get_navigation_paths — Computes BFS shortest-path planning for UI navigation.
  • wait_for_visual_state — Polls for target UI state with configurable timeout.
  • manage_video — Ingests WebM/MP4 video recordings, extracts keyframes, and enables timeline search.
  • compare_states — Compares visual layouts and generates diffs for video trajectory analysis.
  • create_evidence_pack — Generates cryptographically hashable evidence packs linking video keyframes to state DAGs.
  • export_trajectories — Exports multimodal fine-tuning datasets from visual state transitions.
  • manage_visual_spec — Registers design mockups as baseline contracts and performs visual regression checks.
  • manage_snapshot — Creates, exports, and restores visual memory checkpoints.
  • undo_visual_mutation — Reverts state ingestion operations.
  • forget_state — Purges visual states for privacy and PII removal.

Use cases

  • Cache UI screenshots during E2E testing to reduce vision token usage and speed up visual assertions.
  • Ingest Playwright or Cypress test recordings to build searchable visual state timelines and detect regressions.
  • Ground UI elements to CSS selectors and coordinates for deterministic automation and action prediction.
  • Register design mockups as visual specs to verify UI against baseline contracts and catch visual regressions.
  • Generate cryptographic evidence packs linking visual states to task DAGs for compliance and audit trails.

Visual Memory MCP MCP server FAQ

What is the Visual Memory MCP server?

It's a local-first MCP server that caches UI screenshots using perceptual hashing and CLIP embeddings to reduce vision token costs for AI coding assistants. It provides fast visual state recognition, element grounding, video ingestion, and evidence pack generation entirely on your machine.

Is it free?

Yes, Visual Memory MCP is open-source under the MIT License and free to use. It runs entirely locally with no cloud dependencies or telemetry.

How do I install it in Cursor or Claude?

Install globally via npm: `npm install -g @putervision/vision-memory-mcp`. Then run `vision-memory-mcp init --yes` in your project. Add the server to your MCP client config (e.g., `.cursor/mcp.json`) with command `vision-memory-mcp` and args `["run"]`.

What are the system requirements?

Node.js >= 18.18.0 is required. The full CLIP model requires ~300 MB RAM; use `--skip-model-load` for lightweight dHash-only mode on memory-constrained systems.

Does it require authentication or API keys?

No. Visual Memory MCP is 100% local-first with zero cloud dependencies. All screenshots, embeddings, and state data are stored locally in `.vision-memory-mcp/` with no external API calls or telemetry.

Can I use it for backend or headless development?

It's designed for visual frontend state caching and UI testing. For pure backend development without UI rendering, the companion `state-memory-mcp` server is more appropriate.

README (reference)

Source of truth, from the repository.

@putervision/vision-memory-mcp

npm version version npm downloads CI Node TypeScript Website License: MIT

@putervision/vision-memory-mcp is a zero-infrastructure, local-first Model Context Protocol (MCP) server and CLI tool that provides AI coding assistants (such as Cursor, Claude Code, Gemini, or Copilot) with visual state caching using perceptual hashing, local CLIP embeddings, and transition graphs to eliminate repetitive vision LLM calls.

🌐 Official Documentation & Website: visionmemorymcp.com


⚡ Quick Start & Installation

Prerequisites: Node.js >= 18.18.0

1. Installation

# Global installation via npm
npm install -g @putervision/vision-memory-mcp

2. Workspace Initialization

Run init in your project root to scaffold database directories, .gitignore, .env, and IDE rules:

vision-memory-mcp init --yes

3. Basic MCP Client Setup

Add to your MCP client config (e.g. .cursor/mcp.json or .vscode/mcp.json):

{
  "mcpServers": {
    "vision-memory-mcp": {
      "command": "vision-memory-mcp",
      "args": ["run"]
    }
  }
}

Alternative Options & CLI Usage Examples

# Run stdio MCP server directly via binary (after global install)
vision-memory-mcp run

# Start server skipping heavy CLIP model downloads (air-gapped / offline mode)
vision-memory-mcp run --skip-model-load

# Re-initialize across all registered workspace projects
vision-memory-mcp init-global

# Health check dependencies, sharp bindings, and git safety
vision-memory-mcp doctor

# Run health diagnostics & aggregate metrics across all registered projects
vision-memory-mcp doctor-global

# Inspect stored visual states and metadata in terminal ASCII table
vision-memory-mcp inspect

# Register baseline design mockup contract (Visual SDD)
vision-memory-mcp spec set --name "Dashboard" --file ./dashboard-spec.png

# Save visual memory checkpoint snapshot
vision-memory-mcp snapshot save --name "v1.0-milestone"

# Ingest WebM / MP4 video recording into visual state memory timeline
vision-memory-mcp video ingest ./playwright-test.webm --category playwright_test

# Open interactive force-directed visual graph viewer in browser
vision-memory-mcp view

🌟 Key Highlights

  • 👁️ Perceptual Visual Caching & Compact Slices: Sub-5ms L1/L2 dHash zero-token fast-path layout recognition and sub-1KB compact_slice export for Pentad System 1 StatePack assembly.
  • 🎬 WebM & MP4 Video Ingestion: Digest E2E test recordings & screen captures into searchable keyframe visual states & state transition graphs.
  • ⚡ 15 Core MCP Tools: High-coherence consolidated toolset covering perception, video memory, evidence packs, trajectory comparison, semantic retrieval, element grounding, visual SDD, snapshots, and unified context & metrics.
  • 🔗 Dual-MCP Synergy & Immutable Evidence Packs: Deeply bridges @putervision/state-memory-mcp task DAGs with visual state memory, generating cryptographically hashable evidence packs for compliance and audit trails.
  • 📉 Reduced Token Overhead: Caches UI states locally using dHash, local CLIP vector search, and accessibility trees to maximize vision token savings.
  • 🚀 Sub-5ms Fast-Path Latency: Eliminates repetitive vision LLM API calls and avoids visual hallucination loops.
  • 🎯 Element Grounding & Action Target Prediction: Maps screen elements to CSS selectors and coordinates for deterministic UI interaction.
  • 🎨 Visual Spec-Driven Development (Visual SDD): Register design mockups or screenshots as perceptual baseline contracts to verify visual regression.
  • 🛡️ 100% Local-First Privacy: Local LanceDB vector store, local CLIP model, zero cloud telemetry, and PII redaction guarantees.

🛠️ MCP Tool Suite

@putervision/vision-memory-mcp provides 15 production-grade consolidated MCP tools structured across 4 core visual perception & automation domains:

  • Perception & Semantic Search: analyze_screenshot (L1/L2 perceptual dHash & AX tree parsing, single/batch), recall_memory (text & image semantic vector search), get_session_context (aggregated cache hit metrics, recent states, and sub-1KB compact_slice export).
  • Element Grounding & Navigation: predict_next_action (deterministic CSS selectors & bounding coordinates), record_outcome (UI action transitions & visual blockers), get_navigation_paths (BFS shortest-path planner), wait_for_visual_state (polling for target UI state).
  • Video Trajectories & Evidence Packs: manage_video (WebM/MP4 keyframe ingestion, timeline search), compare_states (visual layout diffs & video trajectory comparison), create_evidence_pack (cryptographic audit proof linking video keyframes to state-memory DAGs), export_trajectories (multimodal fine-tuning datasets).
  • Snapshots & Visual SDD: manage_visual_spec (mockup baseline contracts & regression checks), manage_snapshot (checkpoints, export, restore), undo_visual_mutation (revert state ingestion), forget_state (privacy & PII purging).

👉 For complete parameter specifications, return schemas, and example payloads, see the Formal API Reference and Features & Architecture Guide.


🚀 Architecture At a Glance

                     Incoming Screen
                            │
                            ▼
              ┌──────────────────────────────┐
              │ L1: In-Memory Cache Lookup   │ ──(Hit)──▶ Return Cached Description & Grounded Elements
              └──────────────┬───────────────┘
                             │ (Miss)
                             ▼
              ┌──────────────────────────────┐
              │ L2: Perceptual Hash Scan     │ ──(Hit)──▶ Return Cached Description & Grounded Elements
              └──────────────┬───────────────┘
                             │ (Miss)
                             ▼
              ┌──────────────────────────────┐
              │ L3: Local CLIP Vector Search │ ──(Hit)──▶ Return Semantically Close
              └──────────────┬───────────────┘
                             │ (Miss)
                             ▼
              ┌──────────────────────────────┐
              │ L4: Vision LLM Fallback      │ ──(Ingest)──▶ Save Redacted State to DB
              └──────────────┬───────────────┘

📚 Documentation Directory

Explore dedicated guides and deep dives in the docs/ directory:

GuideDescription
🏗️ Architecture & Codebase DistillationHigh-signal architectural overview, 4-tier pipeline, module inventory, and design decisions.
🚀 Features & ArchitectureKey features, 4-tier retrieval pipeline, element grounding, and Dual MCP Synergy.
📘 Formal API ReferenceComplete specifications, parameters, and schemas for all 15 consolidated MCP tools.
🔌 Multi-IDE Integration GuideStep-by-step configs for Cursor, Claude Desktop, Antigravity, Windsurf, Zed, Roo Code & Agent Rules.
💻 CLI Commands ReferenceFull guide for all 16 CLI management, visual spec, and snapshot commands.
⚙️ Configuration GuideComplete .env environment variables, thresholds, and L4 vision fallback setup.
🔒 Storage Encryption & SecurityEncryption details, local storage privacy, and PII masking guarantees.
🤝 Contributing GuideDevelopment setup, codebase structure, and submission guidelines.
🛡️ Security PolicySecurity vulnerability reporting and privacy disclosures.
📜 ChangelogChronological record of release features, fixes, and patch updates.

⚠️ When Not to Use This Server

While vision-memory-mcp is designed for visual frontend state caching, UI testing, and multimodal workflows, it may not be appropriate for:

  • Headless / Pure Backend Development: Non-visual CLI tools, database scripts, or pure backend microservices with no UI rendering. (Use state-memory-mcp standalone instead).
  • High-Framerate Live Video Streaming: Continuous 60 fps live video ingest without discrete keyframe or test action boundaries.
  • Ultra Low-Memory Embedded Environments (<512 MB RAM): Running full local CLIP neural embeddings requires ~300 MB RAM (use --skip-model-load for lightweight dHash-only perception if memory is constrained).

🧪 Testing

# Run full unit and integration test suite across all 72 test files (312 tests)
npm run test

⚖️ License & Disclaimers

Developed and maintained by PuterVision. Released under the MIT License.

  • Local Storage Guarantee: Provided "as is" without warranty. Screenshots, perceptual hashes, vector embeddings, and transition graphs are stored locally unencrypted at the application level in .vision-memory-mcp/. Zero telemetry or analytics data is ever transmitted.
  • Trademarks & Non-Affiliation: Product names (Cursor, Claude Code, Gemini, Windsurf, VS Code, Sharp, LanceDB, ONNX, HuggingFace) are property of their respective owners and used solely for compatibility identification.

Related MCP servers

Zero-dependency Web Crypto MCP server with AES-256-GCM, RSA-4096, and post-quantum cryptography for AI agents.

28
JavaScript
MIT
View repository →

Persistent 3D/2D spatial world model for AI agents with entity tracking, collision simulation, and durable memory.

46
TypeScript
MIT
View repository →

Strategic BDI reasoning engine for autonomous AI agents with goal decomposition, utility theory, and adaptive replanning.

18
TypeScript
MIT
View repository →

High-frequency (~60Hz) in-browser behavior tree execution engine with reactive triggers and safety guardrails for AI agents.

21
TypeScript
MIT
View repository →

File uploads for AI agents. Upload, list, and manage files. No signup required.

1
JavaScript
View repository →

MCP server for AI image generation via OpenAI, Stable Diffusion (SD WebUI), or placeholders.

1
Python
MIT
View repository →