PluginBench
MCP Server
Active
Apache-2.0

Munim Computer Use MCP Server

io.github.munimtechnologies/munim-computer-use

Accessibility-first desktop control for AI agents with background input, agent pointer, and Chrome integration across macOS, Windows, and Linux.

What is the Munim Computer Use MCP server?

The Munim Computer Use MCP server is an open-source computer-use agent backend that enables AI agents like Claude and Cursor to control your desktop by reading accessibility trees and performing background input without moving your mouse. It supports macOS, Windows, and Linux, and includes features like agent pointer overlay, zoom, and signed-in Chrome tab control.

Munim Computer Use lets coding agents interact with your computer the way a person does—reading the screen through accessibility trees, clicking and typing in the background so your mouse stays yours, and showing where the agent is working. It works with Claude Code, Cursor, Codex, and any MCP client, with no vision model required.

How to install Munim Computer Use

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • COMPUTER_USE_BROWSER

    Set to 0 to hide the browser_* tools

  • COMPUTER_USE_AGENT_CURSOR

    Set to 0 to disable the agent pointer overlay

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "munim-computer-use": {
      "command": "npx",
      "args": [
        "-y",
        "munim-computer-use"
      ],
      "env": {
        "COMPUTER_USE_BROWSER": "<YOUR_COMPUTER_USE_BROWSER>",
        "COMPUTER_USE_AGENT_CURSOR": "<YOUR_COMPUTER_USE_AGENT_CURSOR>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • get_app_state — Read the screen through accessibility trees to understand the current UI state and available interactive elements.
  • screenshot — Capture the current screen state for verification and visual feedback.
  • zoom — Zoom into a region at full resolution to read small text or interact with fine details.
  • browser_read — Read a page's text in a signed-in Chrome tab in chunks with optional query filtering, without requiring vision.
  • browser_request_credentials — Securely request sign-in credentials for websites; credentials go directly into the page and never into the conversation.
  • click — Click on UI elements by accessibility element ID or coordinate.
  • type — Type text into focused fields in the background.
  • hover — Hover over elements to trigger tooltips or state changes.
  • wait — Wait for UI state changes or elements to appear.

Use cases

  • Automate web searches and data collection across multiple sites while staying signed in.
  • Control desktop applications and perform multi-step workflows without interrupting your work.
  • Fill out forms, navigate complex UIs, and verify results using accessibility-tree-based interaction instead of pixel coordinates.
  • Manage Chrome tabs and browser automation with your real, signed-in browser instead of a sandboxed environment.
  • Interact with password-protected sites securely by letting the agent request credentials without exposing them to the model.

Munim Computer Use MCP server FAQ

What is Munim Computer Use?

It's an MCP server that gives AI agents like Claude and Cursor the ability to control your desktop—reading screens via accessibility trees, clicking and typing in the background, and driving your real Chrome browser.

Is it free?

Yes, Munim Computer Use is open-source under the Apache 2.0 license.

How do I install it in Cursor?

Create or edit ~/.cursor/mcp.json with the entry {"mcpServers": {"munim-computer-use": {"command": "npx", "args": ["-y", "munim-computer-use"]}}} or download a binary and point to it.

How do I install it in Claude Code?

Run `claude mcp add munim-computer-use -- npx -y munim-computer-use` or point to a downloaded binary.

Does it require authentication or API keys?

No API keys are required. On macOS, you must grant Accessibility and Screen Recording permissions once via `munim-computer-use request-permissions`.

What platforms does it support?

macOS, Windows, and Linux. Full background input works on macOS; Windows supports UI Automation patterns and Win32 controls; Linux uses AT-SPI and X11 (Wayland apps get element actions but not coordinate clicks).

README (reference)

Source of truth, from the repository.

<!-- Banner Image --> <p align="center"> <a href="https://github.com/munimtechnologies/munim-computer-use"> <img alt="Munim Technologies Computer Use" height="128" src="./.github/resources/banner.png?v=1"> <h1 align="center">munim-computer-use</h1> </a> </p> <p align="center"> <a aria-label="Latest release" href="https://github.com/munimtechnologies/munim-computer-use/releases/latest" target="_blank"> <img alt="Latest release" src="https://img.shields.io/github/v/release/munimtechnologies/munim-computer-use?style=flat-square&label=Version&labelColor=000000&color=0066CC" /> </a> <a aria-label="License" href="https://github.com/munimtechnologies/munim-computer-use/blob/main/LICENSE" target="_blank"> <img alt="License: Apache-2.0" src="https://img.shields.io/badge/License-Apache%202.0-success.svg?style=flat-square&color=33CC12" /> </a> <a aria-label="package downloads" href="https://www.npmtrends.com/munim-computer-use" target="_blank"> <img alt="Downloads" src="https://img.shields.io/npm/dm/munim-computer-use.svg?style=flat-square&labelColor=gray&color=33CC12&label=Downloads" /> </a> <a aria-label="total package downloads" href="https://www.npmjs.com/package/munim-computer-use" target="_blank"> <img alt="Total Downloads" src="https://img.shields.io/npm/dt/munim-computer-use.svg?style=flat-square&labelColor=gray&color=0066CC&label=Total%20Downloads" /> </a> <a aria-label="MCP" href="https://modelcontextprotocol.io" target="_blank"> <img alt="MCP stdio server" src="https://img.shields.io/badge/MCP-stdio%20server-8A2BE2?style=flat-square" /> </a> <img alt="Platforms" src="https://img.shields.io/badge/Platforms-macOS%20%7C%20Windows%20%7C%20Linux-lightgrey?style=flat-square" /> </p> <p align="center"> <a aria-label="download" href="https://github.com/munimtechnologies/munim-computer-use/releases/latest"><b>Download</b></a> &ensp;•&ensp; <a aria-label="documentation" href="https://github.com/munimtechnologies/munim-computer-use#readme">Read the Documentation</a> &ensp;•&ensp; <a aria-label="report issues" href="https://github.com/munimtechnologies/munim-computer-use/issues">Report Issues</a> &ensp;•&ensp; <a aria-label="website" href="https://munimtech.com/computer-use">munimtech.com/computer-use</a> </p> <h6 align="center">Follow Munim Technologies</h6> <p align="center"> <a aria-label="Follow Munim Technologies on GitHub" href="https://github.com/munimtechnologies" target="_blank"> <img alt="Munim Technologies on GitHub" src="https://img.shields.io/badge/GitHub-222222?style=for-the-badge&logo=github&logoColor=white" /> </a>&nbsp; <a aria-label="Follow Munim Technologies on LinkedIn" href="https://linkedin.com/in/sheehanmunim" target="_blank"> <img alt="Munim Technologies on LinkedIn" src="https://img.shields.io/badge/LinkedIn-0077B5?style=for-the-badge&logo=linkedin&logoColor=white" /> </a>&nbsp; <a aria-label="Visit Munim Technologies Website" href="https://munimtech.com" target="_blank"> <img alt="Munim Technologies Website" src="https://img.shields.io/badge/Website-0066CC?style=for-the-badge&logo=globe&logoColor=white" /> </a> </p>

Introduction

Computer Use is an open-source MCP server — a computer-use agent (CUA) backend — that lets any coding agent use your computer the way a person does. It reads the screen through accessibility trees, clicks and types in the background so your mouse stays yours, shows an agent pointer where it is working, zooms in on small text, and drives tabs in your signed-in Chrome — on macOS, Windows and Linux.

Works with Claude Code, Codex, Cursor and MT Code, or any other MCP client, with any model — no vision model is required for interaction.

Built by Munim Technologies as the Computer Use engine of MT Code, and published here on its own.

Table of contents

Quick start

  1. Download the latest binary for your platform from Releases (munim-computer-use-macos-universal.zip, munim-computer-use-windows-x64.zip) or build from source.
  2. Put it somewhere on your PATH (/usr/local/bin/munim-computer-use, or %LOCALAPPDATA%\Programs\munim-computer-use\munim-computer-use.exe).
  3. macOS only: run munim-computer-use request-permissions once to be prompted for Accessibility and Screen Recording.
  4. Register it with your agent:
# Claude Code — fastest: the npm launcher fetches the signed binary on first run.
# Add `--scope user` to register it for every project instead of just this one.
claude mcp add munim-computer-use -- npx -y munim-computer-use
# or point at a downloaded binary
claude mcp add munim-computer-use -- /usr/local/bin/munim-computer-use
# Codex — writes the entry below into ~/.codex/config.toml for you (needs a Codex CLI
# with `codex mcp`; check with `codex mcp --help`)
codex mcp add munim-computer-use -- npx -y munim-computer-use
# Codex — ~/.codex/config.toml, if you would rather edit it yourself
[mcp_servers.munim-computer-use]
command = "npx"
args = ["-y", "munim-computer-use"]
# Cursor — no MCP subcommand in its CLI, so write the config. ~/.cursor/mcp.json applies
# to every project; .cursor/mcp.json in a repo applies to that one.
mkdir -p ~/.cursor && [ -s ~/.cursor/mcp.json ] || echo '{}' > ~/.cursor/mcp.json
jq '.mcpServers["munim-computer-use"] = {"command":"npx","args":["-y","munim-computer-use"]}' \
  ~/.cursor/mcp.json > ~/.cursor/mcp.json.tmp && mv ~/.cursor/mcp.json.tmp ~/.cursor/mcp.json
// Cursor — the entry that produces, in .cursor/mcp.json
{ "mcpServers": { "munim-computer-use": { "command": "npx", "args": ["-y", "munim-computer-use"] } } }

Then ask: "Open Safari, find the cheapest flight to Denver on Tuesday and put it in a note." The agent reads the UI with get_app_state, acts by element id, and verifies with screenshot.

Capability matrix

Munim Computer Use is the highlighted first column; the others are the computer-use servers people reach for. each cell comes from that project's own README or docs in September 2026 (sources under Credits and license). ✅ present · ❌ absent or not documented · ⚠️ partial.

CapabilityMunim Computer UseCodex Computer UseOpenAI Agents APIAnthropic reference demoWindows-MCPMacOS-MCPopen-computer-usecomputer-use-mcp (zavora)Notes
macOS✅✅n/a❌❌✅✅✅The Anthropic demo drives a Linux desktop inside Docker, not your machine. The OpenAI Agents API (September 2026) runs a browser OpenAI hosts, so it never drives your machine either.
Windows✅✅n/a❌✅❌✅✅Windows-MCP is Windows only; MacOS-MCP is macOS only.
Linux✅❌n/a✅ (sandbox)❌❌✅✅Munim Computer Use uses AT-SPI + X11; native Wayland apps get element actions but not coordinate clicks.
Accessibility tree with element ids✅❌❌❌✅✅✅✅Codex, the Agents API and the Anthropic demo are screenshot-driven. Ids let the agent press the button instead of a pixel.
Background input (your mouse never moves)✅✅n/an/a❌❌❌❌Munim Computer Use addresses events to the target window: always on macOS (SkyLight and per-process events), for UI Automation patterns and classic Win32 controls on Windows; other Windows UI and all Linux pointer input (XTEST) still move the real pointer. Codex does this too, with a second cursor of its own. Every other server in this table drives the real cursor.
Agent pointer overlay✅✅❌❌⚠️❌❌❌Windows-MCP flashes a border around captures; Codex draws its own cursor on your screen.
Zoom into a region at full resolution✅❌❌✅❌❌❌❌Anthropic's toolset has zoom; here it is a tool on every platform.
Screenshots carry screen-coordinate mapping✅n/an/an/a❌❌❌❌Origin and pixels-per-point in every capture, so clicks from Retina or downscaled images land.
Hover, wait, label query✅⚠️❌⚠️✅⚠️❌⚠️Windows-MCP has Wait/WaitFor; MacOS-MCP has Wait; Anthropic has wait/mouse_move.
Your signed-in Chrome, own tab group or yours✅⚠️❌❌⚠️❌❌❌Codex uses its in-app browser; Windows-MCP reads the DOM of open browsers. Munim Computer Use opens its own labelled tab group in your real Chrome, and can also take over a tab you already have open when you ask it to.
Act and observe in one call✅n/an/a✅❌❌❌❌return_state on any action appends the fresh accessibility tree or page snapshot, so one call acts and checks. The Anthropic demo returns a screenshot after each action; Codex and the Agents API run the loop themselves.
Read a page's text in a signed-in tab✅❌❌❌⚠️❌❌❌browser_read returns the page as text with headings, in chunks, with a query filter, from a background tab. Windows-MCP scrapes pages.
Sign-in without the model seeing the password✅❌✅❌❌❌❌❌browser_request_credentials opens a Chrome window naming the site's real origin; what the user types goes into the page and never into the conversation. The Agents API hands sign-in to the developer's app the same way. Desktop password fields refuse typing by default.
Per-app and per-site allow / ask / block rules✅⚠️⚠️❌❌❌❌❌A policy file the user controls; ask shows an approval prompt once per task. Codex asks before each new app and keeps an "Always allow" list; the Agents API asks the developer's app before each new website.
Works with any MCP client✅❌❌❌✅✅✅✅Codex Computer Use is Codex only; the Anthropic demo is Claude only.
Identical tool surface on every platform✅n/an/an/an/an/a⚠️✅32 tools with byte-identical schemas across the Swift and Rust servers.
Prebuilt signed binaries + npm launcher✅✅n/a❌❌❌✅ (npm)✅ (npm)macOS universal (Developer ID signed) and Windows x64 on Releases.
Open source✅ Apache-2.0❌❌✅✅ MIT✅ MIT✅ MIT✅ MIT

Also looked at: mediar-ai/mcp-server-macos-use (macOS, accessibility, real input), deploymenttheory/windows-mcp-server (Windows, UIA Invoke patterns, WaitFor), nuphus-mcp (OCR + bring-your-own vision model, CDP Chrome), computer-control-mcp (PyAutoGUI + OCR), and microsoft/playwright-mcp (browser only). Corrections welcome — open an issue with a link.

Why it works well

  • Accessibility first, pixels second. get_app_state returns the app's accessibility tree with stable element ids, so the agent presses the button instead of guessing at a coordinate. It costs a fraction of the tokens of a screenshot and it is what scores highest on OSWorld-style tasks. Screenshots are for verifying and for content the tree cannot describe.
  • Background control. Events are addressed to the target window (SkyLight on macOS, UI Automation patterns and posted window messages on Windows). On macOS the agent never takes your mouse or keyboard; see Works alongside you for the exact guarantee on each platform.
  • Pointer overlay, not your pointer. A soft lavender agent pointer shows where the agent is acting. Your cursor is untouched.
  • Coordinates that land. Every screenshot and zoom carries its screen origin and pixels-per-point. zoom captures any region at full physical resolution.
  • Your browser, your logins. The Chrome extension gives the agent its own labelled tab group in your signed-in Chrome, and leaves your tabs alone unless you point it at one.
  • Or the tab you already have open. browser_list_tabs all=true shows every tab in the browser and browser_use_tab takes one over in place — useful when the page is already signed in or mid-flow and re-opening the URL would throw that away. An adopted tab is not moved into the agent's group, not activated and not reloaded; cleanup releases it rather than closing it, and browser_release_tab hands it back early.
  • Parallel tasks, one extension. Every MCP process gets its own tab group, and one process can run several tasks by passing a stable session_id on its browser calls. A task cannot drive or adopt another task's tabs, and cleanup (including a process exiting) closes only its own. Any number of MCP processes share the one extension: the first owns it and the rest go through it, and if the owner exits another takes over without closing anyone's tabs. Tasks share Chrome's cookies and logins, and desktop apps and the clipboard are not isolated.
  • Model-agnostic. No vision model is required for interaction; local models work too.
  • Look → act → verify. hover for mouse-over menus, wait for loads, query to find a control by label without reading a whole tree.
  • Act and look in one call. Pass return_state: true to any action (click, type_text, set_value, browser_click, browser_navigate, …) and the result carries the app's fresh accessibility tree, or the page's fresh snapshot, taken once the UI has settled. That halves the round trips of a look-act-verify loop. state_query narrows the desktop tree the same way query does.
  • Pages as text. browser_read returns a page's readable text with its headings, including what is scrolled out of view, from a tab in the background. It comes in chunks you can continue with offset, query keeps only the matching lines under their heading, and include_links lists the links.
  • Passwords stay out of the conversation. browser_request_credentials asks the user to sign in in a small Chrome window that shows the site's real origin. What they type goes straight into the page's fields and is never returned to the model. browser_snapshot never shows a password field's value, and on the desktop, typing into password fields is refused by default.
  • Pick up what you are looking at. get_app_state and screenshot take app: "frontmost" for the app in front of the user, so a client can hand the agent "this window" in one step.

Works alongside you

The agent has its own pointer; yours stays yours.

macOS — guaranteed. Every action goes through accessibility (press, set value, select text, show menu, scroll bars) or through events addressed to the target app's process and window. The server never moves your pointer, never posts into the system-wide input stream, and never holds or blocks your input, so you can keep clicking and typing in other apps while the agent works — even in the same app, on another window. Events come from a private source, so a modifier you are holding does not leak into the agent's clicks. The target app may be brought forward when that is the point of the step (activate_app, or handing you a password field), but not on every action.

The exceptions refuse instead of borrowing your pointer: a click, hover or scroll with no target app (coordinates over the desktop before any get_app_state), and drags that leave the source window (between apps, or onto the desktop). The error says what to pass instead.

Windows — best effort, reported. Element presses (Invoke), set_value, select_text and scrolling through UI Automation's ScrollPattern never touch your pointer, and classic Win32 controls also take clicks, hovers, wheel and drags as posted window messages. Other UI (Chromium, Electron, WPF, UWP) only reacts to real mouse input, so those clicks, hovers and drags move your pointer, and type_text/press_key go to the focused window. Whenever the real pointer was used, the result says via cursor.

Linux — pointer actions use it. XTEST input moves the real pointer and goes to the focused window, and results say via cursor. Element actions through AT-SPI (press, set value, insert text, select) do not move it.

Apps and sites the agent may use

You decide which apps and websites the agent may touch. Put a policy.json in the server's support directory (~/Library/Application Support/computer-use on macOS, %LOCALAPPDATA%\munim-computer-use on Windows, ~/.local/share/munim-computer-use on Linux; munim-computer-use identity prints it as supportDir), or point COMPUTER_USE_POLICY at a file anywhere. An app that embeds the server under its own profile reads the file from its own support directory:

{
  "apps": { "Keychain Access": "block", "com.apple.MobileSMS": "ask" },
  "sites": { "bank.example": "block", "mail.google.com": "ask", "*": "allow" }
}
  • allow, ask or block per app or site. ask shows the user a prompt the first time the agent reaches for it (a system dialog for apps, a small Chrome window for sites), and an approval lasts until that agent task ends. An unanswered prompt counts as no after two minutes.
  • Apps match their name or bundle id, ignoring case. The rule applies when the agent reads, captures or activates the app. Actions then target elements from a snapshot that was allowed.
  • Sites match a host and all its subdomains, and the most specific pattern wins. They are checked on every browser action against the page the tab is showing at that moment, so following a link into a blocked site does not get around it.
  • * sets the default, which is otherwise allow, so {"apps": {"*": "block", "Notes": "allow"}} is an allow-list.
  • Changes apply at once, with no restart. A file that is not valid blocks everything rather than being ignored, so a typo cannot quietly turn a block into an allow.

This is a guard rail for an agent that follows its instructions, not a sandbox. A coordinate click lands wherever it points, and a whole-display screenshot shows every window.

Tools (32)

AreaTools
Seelist_apps, get_app_state (with query), screenshot, zoom, list_displays
Actclick, right_click, hover, drag, scroll, type_text, set_value, select_text, press_key, activate_app, wait
Clipboardclipboard_read, clipboard_write (plain text)
Browserbrowser_open_tab, browser_list_tabs, browser_use_tab, browser_release_tab, browser_select_tab, browser_navigate, browser_snapshot, browser_read, browser_click, browser_type, browser_request_credentials, browser_press_key, browser_close_tab, browser_close_all_tabs

Every action tool also takes return_state, which returns the state after the action in the same call.

Names, argument shapes and descriptions are identical on every platform, and CI enforces it (node scripts/check-tool-parity.mjs); a model that learned them on a Mac needs nothing new on Windows. To change a tool, edit both literals at once with scripts/tool-defs.mjs rather than by hand.

Repository layout

DirectoryWhatBuild
macos/Swift server on the Accessibility API and ScreenCaptureKit (macOS 14+)swift build -c release → .build/release/munim-computer-use
windows-linux/Rust server: UI Automation on Windows, AT-SPI + X11 on Linuxcargo build --release → target/release/munim-computer-use
chrome-extension/Chrome extension + native messaging host for the browser_* toolsLoad unpacked; sh install.sh / install.ps1 registers the host; node background.test.mjs
DockerfileHeadless Linux build of the Rust server for registry introspection (Glama and similar); no desktop control inside a containerdocker build -t munim-computer-use .

Build from source

# macOS
cd macos && swift build -c release
# Windows / Linux
cd windows-linux && cargo build --release

Linux notes: element actions work everywhere; coordinate clicks need an X11 or XWayland client, since native Wayland apps do not expose absolute geometry.

Tests

cd windows-linux && cargo test                  # Rust server
node chrome-extension/background.test.mjs       # extension, against a fake Chrome
node scripts/check-tool-parity.mjs              # Swift and Rust tool lists match
# The browser tools end to end, in a throwaway Chrome for Testing profile that
# never touches your own Chrome (npx playwright install chromium to get one):
node scripts/e2e-browser.mjs --server <munim-computer-use binary> --chrome <Chrome for Testing binary>

Chrome extension (optional)

  1. chrome://extensions → Developer mode → Load unpacked → select chrome-extension/.
  2. Register the native messaging host: munim-computer-use install-native-host (for example npx -y munim-computer-use install-native-host), or from a checkout sh chrome-extension/install.sh (macOS/Linux) / powershell -File chrome-extension/install.ps1 (Windows), which find the build and run the same command. Point COMPUTER_USE_PATH at the binary if it is not in the default build location.

The standalone server's host is com.munimtech.computer_use.desktop; MT Code, which bundles this server, registers com.munim.mtcode.desktop. The extension connects to every host it knows at once and answers each on its own connection, so MT Code and a standalone server (npx, Claude Code, Cursor, a checkout) can both drive Chrome at the same time, each in its own tab groups.

Environment flags

VariableEffect
COMPUTER_USE_BROWSER=0Hide the browser_* tools
COMPUTER_USE_AGENT_CURSOR=0Do not draw the agent pointer
COMPUTER_USE_AGENT_CURSOR_TASK_FADE_SECSHow long the pointer stays after the last tool call (default 8)
COMPUTER_USE_ALLOW_SECURE_FIELD_INPUT=1Allow typing into password fields (refused by default)
COMPUTER_USE_REMOTE_CONTROL=1Remote-desktop mode: input takes over the real pointer
COMPUTER_USE_POLICY=<file>Where to read the app and site policy (default policy.json in the support directory)

Remote control

Normally this server never touches the pointer: coordinate clicks are routed to a specific window, keystrokes are posted to a specific process, and an action that cannot be targeted is refused rather than taking over the machine. That is what lets an agent work while the user keeps using their computer.

COMPUTER_USE_REMOTE_CONTROL=1 inverts that contract for one process, for the case where a person is watching this machine's screen from another one and is steering it themselves. Then click, right_click, drag, hover and scroll move the real cursor and type_text and press_key go to whatever is focused, the way Chrome Remote Desktop or Screen Sharing behave. scroll also accepts x/y so the wheel acts over the point the viewer scrolled at, and the agent-cursor overlay stays hidden — there is only one pointer now.

A host that runs both an agent and a viewer runs them as two processes, so turning this on for the viewer never takes the pointer away from the user on the agent's behalf.

Embedding in an app

An app can ship this binary inside its own bundle and run it under its own identity, so it never shares a browser bridge, agent-cursor app or native-messaging host with a standalone install on the same machine. MT Code does exactly this. Pass a profile, either as a JSON object or a path to a JSON file, with --profile <json|file> (any position) or COMPUTER_USE_PROFILE:

{
  "name": "example-desktop",
  "envPrefix": "EXAMPLE_DESKTOP_",
  "agentCursorName": "ExampleAgentCursor",
  "agentCursorBundleId": "com.example.agent-cursor",
  "nativeHostNames": ["com.example.desktop"],
  "extensionIds": ["abcdefghijklmnopabcdefghijklmnop"],
  "nativeHostDescription": "Example desktop control bridge"
}
KeyDefaultMeaning
namenone (standalone paths)Moves the support dir, bridge socket and Windows pipe under this name
supportDir~/Library/Application Support/computer-use, $XDG_DATA_HOME/munim-computer-use, %LOCALAPPDATA%\munim-computer-useNative-host wrapper, profile copy; on macOS also the bridge socket and a bare build's overlay app
bridgeSocket<supportDir>/bridge.sock (macOS), $XDG_RUNTIME_DIR/<name>/bridge.sock (Linux), <name>-bridge-<user> pipe (Windows)Where the MCP server and the Chrome relay meet
envPrefixnoneTunables are read as <prefix>BROWSER, <prefix>AGENT_CURSOR, … before COMPUTER_USE_*
agentCursorName / agentCursorBundleIdMunimAgentCursor / com.munimtech.computer-use.agent-cursorThe pointer overlay's app, executable and window-class name, and its macOS bundle id
historyDirnoneDefault --root for computer-history
nativeHostNames / extensionIdscom.munimtech.computer_use.desktop, com.munim.mtcode.desktop / kgdolgnijopbghhomnblabjkmjhnoageWhat install-native-host registers, and for which extension. The first name is this identity's own; later ones are aliases, written only if no other installed app owns them

Each of supportDir, bridgeSocket, envPrefix, agentCursorName, agentCursorBundleId and historyDir can also be overridden by COMPUTER_USE_<SNAKE_CASE> (for example COMPUTER_USE_SUPPORT_DIR), which wins over the profile. munim-computer-use identity prints the resolved values.

  • Browser bridge. Run munim-computer-use install-native-host with the same profile. It writes a wrapper that relays Chrome into this identity's bridge (replaying the profile), and a host manifest for each name in every Chrome/Chromium profile directory (the registry on Windows). It rewrites nothing that is already current, so an app can call it on every launch.
  • Extension. Use the stock extension (it connects to com.munim.mtcode.desktop and com.munimtech.computer_use.desktop), or build a variant with its own host names, tab-group title and key: node scripts/build-extension.mjs --out <dir> --host com.example.desktop --group-title "Example" --key <base64>. It prints the variant's extension id for extensionIds.
  • macOS permissions. The MCP server is a bare executable, so Accessibility and Screen Recording are granted to the app that spawns it. Only the agent-cursor overlay has a bundle of its own; ship <agentCursorName>.app (a copy of the binary plus an LSUIElement Info.plist) beside the binary, or it is materialised under supportDir on first use.

Prompting your agent

Look → act → verify. get_app_state for ids, act by id with return_state: true to see the result in the same call, and screenshot when the tree cannot show it. Read pages with browser_read, and let the user type passwords through browser_request_credentials. Use zoom for small text, hover for menus that appear on mouse-over, wait after loads, keyboard shortcuts for stubborn widgets. The system-prompt text MT Code gives its agents lives in CodexDeveloperInstructions.ts and is a good starting point.

Contributing

This repository mirrors the native/ tree of munimtechnologies/mtcode, where the server is developed and shipped inside MT Code. Issues and discussions are welcome here; code changes land in mtcode first and are synced.

Credits and license

Designed and built by Munim Technologies (Munim, Inc.) for MT Code. Copyright 2026 Munim, Inc. Licensed under the Apache License 2.0; see LICENSE.

Comparison sources: Codex Computer Use and its docs · OpenAI Agents API computer use · Anthropic computer-use demo · CursorTouch/Windows-MCP · CursorTouch/MacOS-MCP · QwenLM/open-computer-use · zavora-ai/computer-use-mcp · mediar-ai/mcp-server-macos-use · deploymenttheory/windows-mcp-server · nuphus-mcp · computer-control-mcp · microsoft/playwright-mcp

Related MCP servers

LALayerbridge logo

Read, edit and export the Figma file open in a local plugin. No REST API, no rate limits.

0
TypeScript
MIT
View repository →

Bridges any GraphQL API to Claude Code — auto-generates MCP tools via schema introspection.

1
TypeScript
MIT
View repository →

PlayStation Network (PSN) gaming history: playtime, trophies and purchased library.

1
Python
MIT
View repository →
VPVPS Guardian logo

Secure SSH bridge for AI agents to observe and safely administer Linux VPSs.

0
Python
MIT
View repository →

Signed Buildability Oracle for AI-for-science papers. ed25519 receipts, Wave proofs, divergence.

View repository →
AWAWS CLI MCP Server logo

Python MCP server for AWS CLI inspection and operations

0
Python
MIT
View repository →