glass MCP Server
io.github.fixed-width/glass
Give AI coding agents hands to build, test, and debug native GUI apps independently.
What is the glass MCP server?
The glass MCP server is a Rust-based tool that enables AI coding agents to launch, interact with, and debug native GUI applications across multiple platforms (Linux, Windows, macOS, Android, iOS). It provides a closed build→see→interact→debug loop by capturing screenshots, injecting input, reading logs, and detecting visual changes without requiring app integration or manual verification.
glass lets AI agents drive external GUI applications as a black box, working with any native app regardless of toolkit or language. Instead of asking users "does this look right?", agents can now launch apps, reproduce bugs from accessibility trees, fix code, and verify results independently. It supports multiple platforms with platform-specific backends (X11/Wayland on Linux, native Windows, Android AVD emulator, iOS Simulator, macOS) and provides both accessibility-based semantic addressing and pixel-based interaction for custom-rendered UIs.
How to install glass
Copy-paste configuration for popular MCP clients.
No machine-readable install method is published for this server in the registry. Check the repository or website for setup instructions.
Tools & capabilities
Tools this server exposes to the agent.
glass_start— Launch a GUI application under glass controlglass_do— Execute a sequence of actions on the running app (click, set values, wait for elements)glass_screenshot— Capture the current screen state of the running applicationglass_click— Click at specific pixel coordinates on the appglass_diff— Detect visual changes between screenshots, returning changed percentage and bounding boxglass_logs— Read the application's logs for debuggingglass_find_elements— Inspect and query accessibility tree elements to find UI controlsglass_a11y_snapshot— Explore the app's accessibility tree and available controlsglass_stop— Stop the running application session
Use cases
- Automatically test GUI applications by driving them and verifying results from accessibility trees without screenshots
- Reproduce and fix UI bugs by having an agent interact with the app, detect the issue, modify code, and re-verify
- Build and validate form-based applications by filling fields, clicking buttons, and confirming success states
- Debug canvas or custom-rendered apps using pixel-based interaction and visual diffing
- Develop cross-platform native applications with automated testing across Linux, Windows, macOS, Android, and iOS
glass MCP server FAQ
glass is an MCP server that gives AI coding agents the ability to launch, drive, and debug native GUI applications independently. It works with any GUI app regardless of toolkit or language by providing screenshot capture, input injection, accessibility tree access, and visual change detection.
Yes, glass is open-source under the Apache-2.0 license (open core model). Binaries are available for download from the Releases page.
Download the prebuilt binary for your platform from the Releases page, then follow the platform-specific setup guide (Linux, Windows, macOS, Android, or iOS). Connect it to your MCP host and run `glass-mcp doctor` to verify your environment.
glass supports Linux (X11 and Wayland), Windows, macOS, Android (AVD emulator), and iOS (Simulator). Each platform has full support for capture, input, windows, clipboard, and logs; accessibility features vary by platform.
No. glass drives apps as an external black box, so it works with any native GUI application without requiring any integration or modifications to the app code.
glass-drive is an open Agent Skill that teaches agents how to use glass effectively, helping them verify changes cheaply (via accessibility trees and diffs) before taking expensive screenshot actions. Installing it is recommended for best results.
README (reference)
Source of truth, from the repository.
glass
Give your coding agent hands for the app it's building — launch it, drive it, and verify the result, without burning a screenshot on every step.
A Rust MCP server that gives an AI coding agent a closed build → see → interact → debug loop over external native GUI applications.
glass lets an agent launch a GUI app, capture what is on screen, inject mouse and keyboard input, read the app's logs, and detect visual changes — so a coding agent can build and debug UI applications independently instead of asking the user "does this look right?".
glass drives apps as an external black box, so it works with any native GUI app regardless of toolkit
or language. It has two Linux backends (X11 and Wayland), a Windows backend, an
Android backend (an AVD emulator, driven over adb from any host), an iOS backend (native
apps in the Simulator over xcrun simctl, with input and the accessibility tree via idb_companion,
including a two-finger pinch), and a macOS backend, behind a platform-agnostic core.
See it

An agent building a GUI app runs it under glass, reproduces a bug from the accessibility tree (no screenshots), fixes the code, and re-verifies — the loop it otherwise can't close on its own. Try it yourself.
Try it in 60 seconds
-
Download glass for your platform from the Releases page and connect it to your agent — glass works with any MCP host (see which are verified).
-
Get the example app — clone this repo, or download
examples/tasks-demo/tasks_demo.py(on Linux it needssudo apt install python3-gi gir1.2-gtk-4.0). -
Paste this to your agent:
Use glass to run
examples/tasks-demo/tasks_demo.pywith accessibility on. There's a bug: clicking Add doesn't add the typed task. Reproduce it by driving the UI and checking the accessibility tree (don't just screenshot), then find and fix the bug in the code and verify a task actually appears.
Your agent launches the app, reproduces the bug from the accessibility tree, fixes the one-line
wiring bug, and confirms the task appears — the whole build → see → interact → debug loop, start to
finish. (glass-mcp doctor checks your environment if anything's off.)
The loop in practice
An agent can fill in a form, click Save, and check that the app reports success. When the app exposes an accessibility tree, the agent works with named controls and verifies changes from text:
glass_start { "run": ["python3", "app.py"] }
glass_do { "actions": [
{ "action": "set_value",
"target": { "query": "account", "role": "TextField", "states": ["enabled"] },
"text": "hello" },
{ "action": "click_element",
"target": { "query": "Save", "role": "Button", "states": ["enabled"] },
"mode": "auto" },
{ "action": "wait_for_element", "name": "Saved" }
] }
glass_logs
Use glass_find_elements to inspect candidates when the intended target is not unique or known well
enough to act on directly. Use glass_a11y_snapshot to explore the app's controls. See the
tool reference for click modes and platform behavior.
For a canvas or custom-rendered app with no accessibility tree, drive it by pixels instead —
glass_screenshot, glass_click {x,y}, and glass_diff, which returns changed_pct + a bbox as
text, so routine checks between screenshots cost no vision tokens. Why the loop is shaped this way:
the build → see → interact → debug loop.
To review a run after the app stops, opt into session evidence recording
with --trace-dir. Glass retains tool inputs and requested results in bounded storage, then
glass-mcp trace inspect and trace export validate the trace and package its evidence in a ZIP.
Recording is off by default; retained app content can contain sensitive data.
Install at a glance
Download the latest build for your platform from the Releases page, then set up your host:
- Linux — docs/how-to/setup-linux.md (X11 or Wayland;
Xvfb/sway+ bubblewrap) - Windows — docs/how-to/setup-windows.md (a prebuilt
.exe+ Sandboxie) - macOS — docs/how-to/setup-macos.md (install the notarized
.dmg; no build needed) - Android — docs/how-to/setup-android.md (an AVD emulator, from any host)
- iOS — docs/how-to/setup-ios.md (the Simulator, macOS host only)
Every asset is listed in docs/reference/platforms.md.
Prefer to compile, or on an architecture with no published asset? See
docs/how-to/build-from-source.md — it is a single cargo build.
Then connect glass to your agent and run glass-mcp doctor to check
the environment. New here? Follow the tutorial for a guaranteed first
success.
The full toolbox remains the default. For fewer agent tool definitions, start with
glass-mcp --tool-profile lean and use glass_do for actions. Inspect either inventory without
starting a session using glass-mcp tools --json; see tool profiles.
Drive it well — the glass-drive skill
glass needs no app integration and no skill to run, but an agent drives it far more reliably with the open glass-drive Agent Skill — it stops the agent spending its first turns rediscovering the verify-cheaply-then-look loop. Installing it is the single highest-leverage thing you can add when pointing an agent at glass.
Platform support
✓ supported · ◑ partial · – not supported · 🚧 planned.
<!-- KEEP IN SYNC with docs/reference/platforms.md (the canonical matrix) and the code. -->| Capability | Linux (X11 + Wayland) | Windows | Android (AVD) | iOS (Simulator) | macOS |
|---|---|---|---|---|---|
| Capture · input · windows · clipboard · logs | ✓ | ✓ | ✓ | ✓ | ✓ |
| Accessibility (semantic addressing) | ✓ AT-SPI | ✓ UI Automation | ✓ UIAutomator | ✓ idb | ✓ AX |
| Containment / sandboxing | ✓ bubblewrap | ✓ Sandboxie | ✓ the emulator VM | ✓ the Simulator | ✓ Seatbelt |
| Display isolation (app off your desktop) | ✓ headless Xvfb / sway | ◑ virtual display · VM tier | ✓ headless emulator | ✓ headless simctl boot | 🚧 |
Full matrix, per-capability detail, and system requirements: docs/reference/platforms.md. Transport is MCP over stdio (default) or network HTTP.
Documentation
The full docs — tutorial, how-to guides, reference, and explanations — are under
docs/. See CHANGELOG.md for release notes, and
Stability and versioning for what a 1.0 release guarantees.
Contributing? CONTRIBUTING.md has the gates a PR has to pass, and
Verify a change covers platform code your own host cannot run.
License
glass is open core, licensed Apache-2.0 — see LICENSE-APACHE.
Related MCP servers

io.github.fjnunezp75/gpu-bridge
30 GPU-powered AI services as MCP tools. LLM, image, video, audio, embeddings & more.

llmtrim
Local proxy that compresses LLM API traffic to cut token costs by 31–74% without changing answers.

Versioning de code et génération de documentation intelligents. Supporte CLI, API REST, et MCP.

io.github.fkom13/mcp-sftp-orchestrator
Serveur MCP pour l'orchestration de tâches distantes (SSH/SFTP) avec une file d'attente persistante.

RankReactor
Hosted RankReactor MCP: SEO articles, UGC video, keywords, analytics, and site audits.
Ad marketplace: search Telegram/VK/Max channels by topic, geo, price; reach, ER, CPV stats.