Ui.Vision MCP MCP Server
io.github.A9T9/uivision-mcp
Browser and desktop automation with OCR, image recognition, and AI-assisted macro creation via MCP.
What is the Ui.Vision MCP MCP server?
Ui.Vision MCP is a Model Context Protocol server that connects AI assistants like Claude and Cursor to the Ui.Vision browser automation extension. It enables agents to create, edit, and run reusable macros for web scraping, form filling, visual task automation, and desktop workflows using OCR and image recognition.
Ui.Vision automates websites and desktop applications through browser macros, JavaScript, or AI assistance. The MCP bridge lets Claude, Cursor, and other AI agents inspect pages, write automation scripts, execute them in your browser, and review logs and screenshots—combining visual recognition, OCR, and real mouse/keyboard input for tasks HTML selectors cannot reach.
How to install Ui.Vision MCP
Copy-paste configuration for popular MCP clients.
UIVISION_MCP_PORTLocal WebSocket port the Ui.Vision extension connects to (default 50888; must match Settings > AI in the extension)
Tools & capabilities
Tools this server exposes to the agent.
Page inspection and element interaction— Inspect page elements, fill forms, extract web data, and interact with page controlsOCR and image recognition— Find on-screen text with OCR and locate controls by image when HTML selectors are unavailableMacro creation and execution— Create, edit, and run reusable macros through the MCP interface; inspect resulting scripts, logs, and screenshotsBrowser automation— Record and replay browser interactions, fill forms, download reports, and repeat workflowsDesktop automation— Automate applications and remote desktop interfaces using visual recognition and mouse/keyboard input (requires XModules)JavaScript API (uiv.*)— Write automation scripts using the uiv.* JavaScript API for page elements, visual matching, OCR, input, screenshots, tabs, CSV files, and downloads
Use cases
- Ask an AI assistant to automate form filling, data extraction, or report downloads on websites where selectors are unstable or unavailable
- Create reusable browser macros for repetitive workflows (login, data entry, report generation) and trigger them from Claude or Cursor
- Test websites by recording interactions and replaying them with different input data, using OCR and image recognition for visual verification
- Automate desktop applications and remote desktop interfaces by having an AI agent write and execute macros based on visual recognition
- Extract text and locate UI elements on any website or application using OCR when traditional HTML-based selectors fail
Ui.Vision MCP MCP server FAQ
Ui.Vision MCP is a bridge that connects AI assistants (Claude, Cursor) to the Ui.Vision browser automation extension. It lets agents create, edit, and run macros for web and desktop automation, using OCR, image recognition, and real mouse/keyboard input.
The browser extension is free for personal and commercial use. Desktop components (XModules) and additional editions are available on the Ui.Vision website.
Install the Ui.Vision browser extension from Chrome Web Store, Edge Add-ons, or Firefox Add-ons. Then run `npx uivision-mcp-bridge --setup` to register the MCP bridge with your MCP clients, enable it in Ui.Vision settings, and restart your client.
No authentication is required. The MCP bridge connects locally to your Ui.Vision browser extension; automation runs in your existing browser session.
Yes, with the additional Ui.Vision desktop components (XModules). The browser extension alone automates websites; XModules add visual recognition and input for desktop and remote desktop interfaces.
None required. You can record macros visually, use the command table, write JavaScript with the uiv.* API, or ask an AI assistant to write macros for you.
README (reference)
Source of truth, from the repository.
Ui.Vision - Browser and Desktop Automation
Automate websites and desktop applications with macros, JavaScript, or an AI assistant. Ui.Vision combines browser automation, OCR and image recognition, with an MCP server that lets AI agents create, edit and run reusable macros in your browser.
Install the extension · Connect an AI assistant with MCP · API reference · User forum
What can you automate?
- Browser tasks: fill forms, extract web data, download reports and repeat workflows in your existing browser session.
- Website testing: record and replay interactions, use Selenium IDE commands, and run checks with different input data.
- Visual tasks: find on-screen text with OCR and locate controls by image when HTML selectors are unavailable.
- Desktop workflows: automate applications and remote desktop interfaces using visual recognition and mouse and keyboard input. Desktop automation requires the additional Ui.Vision desktop components (XModules).
- AI-assisted automation: ask the built-in AI assistant or an external MCP client to write and run macros, then inspect the resulting script, logs and screenshots.
Use the macro recorder and command table, write JavaScript with the uiv.* API, or let an AI assistant author the macro. The resulting automation can be saved and run again.
Get started
Install the browser extension; you do not need to build this repository:
Open Ui.Vision and record a browser task, or start with a macro from the extension's demos. For desktop input and additional native capabilities, see the XModules installation guide.
The browser extension is open source and free for personal and commercial use. See the Ui.Vision website for desktop components and available editions.
Connect an AI assistant with MCP
Ui.Vision provides a Model Context Protocol (MCP) server through the uivision-mcp-bridge package. It connects MCP clients such as Claude Code, Claude Desktop and Cursor to the Ui.Vision browser extension.
An agent can inspect a page, create or edit a macro, run it, and use logs and screenshots to check the result. The automation executes through Ui.Vision in your browser.
Start with the MCP setup guide. The installer command is:
npx uivision-mcp-bridge --setup
The installer registers Ui.Vision with supported MCP clients found on your machine. Then enable the MCP bridge in Ui.Vision → Settings → AI, complete the connection steps in the guide, and restart your MCP client. Node.js and npm are required to run the installer.
For a first task, ask your connected assistant:
Use Ui.Vision to open https://example.com, read the page heading, and save a reusable macro for this task.
See mcp/README.md for bridge source, configuration and troubleshooting.
JavaScript automation and documentation for AI assistants
The uiv.* JavaScript API covers page elements, visual matching, OCR, browser and desktop input, screenshots, tabs, CSV files and downloads. Ui.Vision macros execute sequentially: write API calls without async or await.
When asking an AI assistant to write a macro, give it the API reference below so it uses Ui.Vision's supported methods and runtime conventions.
- Ui.Vision JavaScript API and AI authoring guide: API syntax and automation recipes used by the built-in assistant.
- Plain-text reference for AI assistants: documentation that can be supplied directly as context.
- Classic command-to-JavaScript mapping:
uiv.*equivalents and theuiv.run('command', 'target', 'value')bridge for classic commands. - Selenium IDE command reference: documentation for command-table macros.
Help and support
Ask questions and share automation examples in the Ui.Vision user forum, where users, support staff and developers participate.
Build from source
Building the extension is not required if you "only" want to use it.
You can install UI.Vision directly from the Chrome, Edge or Firefox stores, which is the easiest and the recommended way of using the Ui.Vision browser extension. Older versions can be found in the RPA software archive.
The information below is only required and intended for developers:
The project uses Node V20.11.1 and NPM V10.2.4
If you have any questions, please contact us at TEAM AT UI.VISION - Thanks!
Build the extension bundle
npm i -f
npm run build
npm run build-ff
npm run build creates the Chrome/Edge build in dist, npm run build-ff creates the Firefox build in dist_ff.
Develop
npm i -f
npm start
Use npm run start-ff for the Firefox variant. Both run webpack in watch mode, so the bundles in dist (Chrome/Edge) and dist_ff (Firefox) are rebuilt on every change.
Once done, the ready-to-use extension code appears in the /dist directory (Chrome, Edge) or /dist_ff directory (Firefox). Load it via chrome://extensions → "Load unpacked" (Chrome/Edge) or about:debugging → "Load Temporary Add-on" (Firefox).
Repository layout
src/- the extension source (React UI, side panel, macro player, commands)extension/- static extension assets andmanifest.json(Manifest V3)mcp/- the Ui.Vision MCP bridge, which lets Claude Code and other MCP clients create, edit and run macros. See mcp/README.mddist/,dist_ff/- build output for Chrome/Edge and Firefox
Related MCP servers
Formally verified AI safety APIs. 75+ endpoints, pay-per-call via USDC x402, no signup.
View repository →
Pay-to-rank board where agents discover other agents by bid. Bearer auth required.
View repository →US insider trades, 13F holdings, 13D/13G, Form 144 and politicians' stock trades. Free, no API key.

Agent Action Gate
Gate/Prove: deny unattended destructive agent tools. Instant Audit $499 on a2zsoc.com.

io.github.ABRANJAN07/btp-mcp-server
Discover SAP BTP services, browse service catalog, check running instances, query dests, & more

CTlogs.io
Certificate Transparency search: subdomains, certificate history and hostname keyword search.
