PluginBench
MCP Server
Active

Ui.Vision MCP MCP Server

io.github.A9T9/uivision-mcp

Browser and desktop automation with OCR, image recognition, and AI-assisted macro creation via MCP.

What is the Ui.Vision MCP MCP server?

Ui.Vision MCP is a Model Context Protocol server that connects AI assistants like Claude and Cursor to the Ui.Vision browser automation extension. It enables agents to create, edit, and run reusable macros for web scraping, form filling, visual task automation, and desktop workflows using OCR and image recognition.

Ui.Vision automates websites and desktop applications through browser macros, JavaScript, or AI assistance. The MCP bridge lets Claude, Cursor, and other AI agents inspect pages, write automation scripts, execute them in your browser, and review logs and screenshots—combining visual recognition, OCR, and real mouse/keyboard input for tasks HTML selectors cannot reach.

How to install Ui.Vision MCP

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • UIVISION_MCP_PORT

    Local WebSocket port the Ui.Vision extension connects to (default 50888; must match Settings > AI in the extension)

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "uivision-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "uivision-mcp-bridge"
      ],
      "env": {
        "UIVISION_MCP_PORT": "<YOUR_UIVISION_MCP_PORT>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • Page inspection and element interaction — Inspect page elements, fill forms, extract web data, and interact with page controls
  • OCR and image recognition — Find on-screen text with OCR and locate controls by image when HTML selectors are unavailable
  • Macro creation and execution — Create, edit, and run reusable macros through the MCP interface; inspect resulting scripts, logs, and screenshots
  • Browser automation — Record and replay browser interactions, fill forms, download reports, and repeat workflows
  • Desktop automation — Automate applications and remote desktop interfaces using visual recognition and mouse/keyboard input (requires XModules)
  • JavaScript API (uiv.*) — Write automation scripts using the uiv.* JavaScript API for page elements, visual matching, OCR, input, screenshots, tabs, CSV files, and downloads

Use cases

  • Ask an AI assistant to automate form filling, data extraction, or report downloads on websites where selectors are unstable or unavailable
  • Create reusable browser macros for repetitive workflows (login, data entry, report generation) and trigger them from Claude or Cursor
  • Test websites by recording interactions and replaying them with different input data, using OCR and image recognition for visual verification
  • Automate desktop applications and remote desktop interfaces by having an AI agent write and execute macros based on visual recognition
  • Extract text and locate UI elements on any website or application using OCR when traditional HTML-based selectors fail

Ui.Vision MCP MCP server FAQ

What is Ui.Vision MCP?

Ui.Vision MCP is a bridge that connects AI assistants (Claude, Cursor) to the Ui.Vision browser automation extension. It lets agents create, edit, and run macros for web and desktop automation, using OCR, image recognition, and real mouse/keyboard input.

Is Ui.Vision free?

The browser extension is free for personal and commercial use. Desktop components (XModules) and additional editions are available on the Ui.Vision website.

How do I install it in Cursor or Claude?

Install the Ui.Vision browser extension from Chrome Web Store, Edge Add-ons, or Firefox Add-ons. Then run `npx uivision-mcp-bridge --setup` to register the MCP bridge with your MCP clients, enable it in Ui.Vision settings, and restart your client.

What authentication is required?

No authentication is required. The MCP bridge connects locally to your Ui.Vision browser extension; automation runs in your existing browser session.

Can I automate desktop applications?

Yes, with the additional Ui.Vision desktop components (XModules). The browser extension alone automates websites; XModules add visual recognition and input for desktop and remote desktop interfaces.

What programming skills do I need?

None required. You can record macros visually, use the command table, write JavaScript with the uiv.* API, or ask an AI assistant to write macros for you.

README (reference)

Source of truth, from the repository.

Ui.Vision - Browser and Desktop Automation

Automate websites and desktop applications with macros, JavaScript, or an AI assistant. Ui.Vision combines browser automation, OCR and image recognition, with an MCP server that lets AI agents create, edit and run reusable macros in your browser.

Install the extension · Connect an AI assistant with MCP · API reference · User forum

What can you automate?

  • Browser tasks: fill forms, extract web data, download reports and repeat workflows in your existing browser session.
  • Website testing: record and replay interactions, use Selenium IDE commands, and run checks with different input data.
  • Visual tasks: find on-screen text with OCR and locate controls by image when HTML selectors are unavailable.
  • Desktop workflows: automate applications and remote desktop interfaces using visual recognition and mouse and keyboard input. Desktop automation requires the additional Ui.Vision desktop components (XModules).
  • AI-assisted automation: ask the built-in AI assistant or an external MCP client to write and run macros, then inspect the resulting script, logs and screenshots.

Use the macro recorder and command table, write JavaScript with the uiv.* API, or let an AI assistant author the macro. The resulting automation can be saved and run again.

Get started

Install the browser extension; you do not need to build this repository:

Open Ui.Vision and record a browser task, or start with a macro from the extension's demos. For desktop input and additional native capabilities, see the XModules installation guide.

The browser extension is open source and free for personal and commercial use. See the Ui.Vision website for desktop components and available editions.

Connect an AI assistant with MCP

Ui.Vision provides a Model Context Protocol (MCP) server through the uivision-mcp-bridge package. It connects MCP clients such as Claude Code, Claude Desktop and Cursor to the Ui.Vision browser extension.

An agent can inspect a page, create or edit a macro, run it, and use logs and screenshots to check the result. The automation executes through Ui.Vision in your browser.

Start with the MCP setup guide. The installer command is:

npx uivision-mcp-bridge --setup

The installer registers Ui.Vision with supported MCP clients found on your machine. Then enable the MCP bridge in Ui.Vision → Settings → AI, complete the connection steps in the guide, and restart your MCP client. Node.js and npm are required to run the installer.

For a first task, ask your connected assistant:

Use Ui.Vision to open https://example.com, read the page heading, and save a reusable macro for this task.

See mcp/README.md for bridge source, configuration and troubleshooting.

JavaScript automation and documentation for AI assistants

The uiv.* JavaScript API covers page elements, visual matching, OCR, browser and desktop input, screenshots, tabs, CSV files and downloads. Ui.Vision macros execute sequentially: write API calls without async or await.

When asking an AI assistant to write a macro, give it the API reference below so it uses Ui.Vision's supported methods and runtime conventions.

Help and support

Ask questions and share automation examples in the Ui.Vision user forum, where users, support staff and developers participate.

Build from source

Building the extension is not required if you "only" want to use it.

You can install UI.Vision directly from the Chrome, Edge or Firefox stores, which is the easiest and the recommended way of using the Ui.Vision browser extension. Older versions can be found in the RPA software archive.

The information below is only required and intended for developers:

The project uses Node V20.11.1 and NPM V10.2.4

If you have any questions, please contact us at TEAM AT UI.VISION - Thanks!

Build the extension bundle

npm i -f
npm run build
npm run build-ff

npm run build creates the Chrome/Edge build in dist, npm run build-ff creates the Firefox build in dist_ff.

Develop

npm i -f
npm start

Use npm run start-ff for the Firefox variant. Both run webpack in watch mode, so the bundles in dist (Chrome/Edge) and dist_ff (Firefox) are rebuilt on every change.

Once done, the ready-to-use extension code appears in the /dist directory (Chrome, Edge) or /dist_ff directory (Firefox). Load it via chrome://extensions → "Load unpacked" (Chrome/Edge) or about:debugging → "Load Temporary Add-on" (Firefox).

Repository layout

  • src/ - the extension source (React UI, side panel, macro player, commands)
  • extension/ - static extension assets and manifest.json (Manifest V3)
  • mcp/ - the Ui.Vision MCP bridge, which lets Claude Code and other MCP clients create, edit and run macros. See mcp/README.md
  • dist/, dist_ff/ - build output for Chrome/Edge and Firefox

Related MCP servers

Formally verified AI safety APIs. 75+ endpoints, pay-per-call via USDC x402, no signup.

View repository →

Pay-to-rank board where agents discover other agents by bid. Bearer auth required.

View repository →

US insider trades, 13F holdings, 13D/13G, Form 144 and politicians' stock trades. Free, no API key.

Gate/Prove: deny unattended destructive agent tools. Instant Audit $499 on a2zsoc.com.

0
Python
MIT
View repository →

Discover SAP BTP services, browse service catalog, check running instances, query dests, & more

1
Python
View repository →
CTCTlogs.io logo

CTlogs.io

Active

Certificate Transparency search: subdomains, certificate history and hostname keyword search.

0
View repository →