PluginBench
MCP Server
Active
MIT

agent-device MCP Server

io.github.callstack/agent-device

Mobile app automation and verification for AI coding agents on iOS, Android, TV, and desktop.

What is the agent-device MCP server?

agent-device is an MCP server for mobile app automation that lets AI coding agents inspect, control, debug, and verify apps on iOS, Android, HarmonyOS, tvOS, Android TV, web, macOS, and Linux. Agents read accessibility snapshots instead of screenshots alone, act through refs and selectors, and save evidence for review. It works with Claude Code, Cursor, Windsurf, Cline, and other agents via CLI, MCP, or Node.js API.

agent-device gives coding agents a live app feedback loop through accessibility snapshots, interactive controls, and diagnostic tools. Agents can inspect app state, tap or fill UI elements, scroll, make gestures, capture screenshots and video, read logs and traces, and save repeatable workflows as scripts. It coordinates device access across parallel agents and supports local simulators/emulators, physical devices, and remote device clouds.

How to install agent-device

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "agent-device": {
      "command": "npx",
      "args": [
        "-y",
        "agent-device"
      ]
    }
  }
}
Cursor
~/.cursor/mcp.json
{
  "mcpServers": {
    "agent-device": {
      "command": "npx",
      "args": [
        "-y",
        "agent-device"
      ]
    }
  }
}
Windsurf
~/.codeium/windsurf/mcp_config.json
{
  "mcpServers": {
    "agent-device": {
      "command": "npx",
      "args": [
        "-y",
        "agent-device"
      ]
    }
  }
}
VS Code
.vscode/mcp.json
{
  "servers": {
    "agent-device": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y",
        "agent-device"
      ]
    }
  }
}
Claude Code
claude mcp add agent-device -- npx -y agent-device

Tools & capabilities

Tools this server exposes to the agent.

  • snapshotCapture accessibility tree snapshots with refs and selectors instead of screenshots alone
  • pressTap or press UI elements by ref or selector
  • fillFill text fields with input
  • scrollScroll within the app
  • gestureMake custom gestures on the app
  • screenshotCapture visual evidence as PNG
  • videoRecord video of app interactions
  • logsRetrieve app and system logs
  • tracesCapture performance traces and network data
  • crashRetrieve crash details and stack traces
  • assertAssert app state conditions
  • alertHandle system alerts
  • openLaunch an app on a target platform
  • closeClose an app session
  • waitWait for UI to settle or conditions to be met

Use cases

  • Implement UI screens and verify them on iOS simulator and Android emulator with screenshots
  • Reproduce crashes and capture logs leading up to the failure
  • Check for unnecessary React Native re-renders and performance issues
  • Save repeatable workflows as replay scripts and run them in CI/CD pipelines
  • Verify pull requests on physical devices and attach reviewable evidence

agent-device MCP server FAQ

What is agent-device?

agent-device is a command-line tool and MCP server that lets AI coding agents inspect, control, and verify mobile apps and save evidence for review. It supports iOS, Android, HarmonyOS, tvOS, Android TV, web, macOS, and Linux.

Is there an MCP server for mobile app automation?

Yes. agent-device mcp starts the official stdio MCP server. Configure it in your client's mcpServers section with command 'agent-device' and args ['mcp'].

How do I install agent-device in Cursor or Claude?

Install via npm: npm install -g agent-device@latest. Then configure the MCP server in your client settings with command 'agent-device' and args ['mcp']. See the AI Agent Setup docs for per-client details.

Does it work with React Native, Expo, and Flutter apps?

Yes. agent-device supports native iOS and Android apps, plus React Native, Expo, and Flutter apps on supported targets. Commands and evidence vary by target.

Can I run agent-device in CI/CD?

Yes. Record a run as an .ad script, replay it in CI, and keep screenshots and logs as artifacts. The EAS workflow template provides a working example.

Is agent-device free?

Yes. agent-device is open source under the MIT license. It requires Node.js 22.12 or newer (24+ for web automation) and target-specific setup like iOS Xcode or Android SDK.

README (reference)

Source of truth, from the repository.

<a href="https://www.callstack.com/open-source?utm_campaign=generic&utm_source=github&utm_medium=referral&utm_content=agent-device" align="center"> <picture> <img alt="agent-device: mobile app automation and verification for AI coding agents" src="website/docs/public/agent-device-banner.jpg"> </picture> </a>

agent-device

npm version CI License: MIT Glama MCP server

Mobile app automation and verification for AI coding agents. Give coding agents a live app feedback loop through a CLI, built-in MCP server, or typed Node.js API.

<!-- Intro rule: category line, job pitch with the platform list, works-with/proof line, nothing else. Per-platform transports, tool names, vendor lists, and support caveats belong in "How it works" and website/docs, never here. Every name in the proof line must link to public evidence. -->

Let your coding agent verify its changes in the running app. agent-device lets agents inspect, control, debug, and verify apps on iOS, Android, and HarmonyOS (simulators, emulators, and physical devices), plus tvOS, Android TV, Amazon Vega OS TV (Vega Virtual Device), web, macOS, and Linux. Agents read token-efficient accessibility snapshots instead of reasoning over screenshots alone, act through refs and selectors, and save evidence for review. It also coordinates device access across parallel agent worktrees and connects to remote device clouds.

Works with Claude Code, Codex, Cursor, Windsurf, Cline, Goose, and any agent that can run a CLI or connect over MCP, or as the runtime under agents you build with the AI SDK or Eve. Developers at Expensify, Shopify, and others use it to verify their apps.

Quick start

Install the CLI and check setup. It requires Node.js 22.12 or newer; web automation requires Node.js 24 or newer. See Installation for target requirements.

npm install -g agent-device@latest
agent-device doctor
agent-device help workflow

Run doctor yourself before handing the CLI to an agent; help workflow links to the guides for debugging, replay, and profiling, and the installed help always matches the installed version.

Drive an app from the CLI

Add a contact in the built-in iOS Contacts app:

# Start a session.
agent-device open Contacts --platform ios

# Inspect the screen. The example below shows the output; refs vary.
agent-device snapshot -i
# @e2 [button] "Add"

# Use the ref and wait for the UI to settle.
agent-device press @e2 --settle
# The diff includes:
# + @e7 [text-field] "First name"

agent-device fill @e7 "Ada" --settle
# The next diff shows changed values and current refs:
# - @e7 [text-field] "First name"
# + @e14 [text-field] "Ada"
# = @e15 [text-field] "Last name"

# Capture evidence and close the session.
agent-device screenshot ./contact-form.png
agent-device close

Refs are only valid from the latest output: after a --settle command, use the refs in its diff, and take a new snapshot only if the diff omits what you need. Snapshots come from the app's accessibility tree, so clear labels, roles, and test IDs make agent runs more reliable; use screenshots and video as evidence or when accessibility data is poor.

agent-device demo showing Codex using agent-device to create a new contact in the iOS Contacts app from a simple prompt

Add MCP tools to your agent

agent-device mcp starts the official stdio MCP server, exposing the installed commands as structured tools over the same execution path as the CLI:

{
  "mcpServers": {
    "agent-device": {
      "command": "agent-device",
      "args": ["mcp"]
    }
  }
}

See AI Agent Setup for per-client setup and when to prefer plain CLI over MCP.

Script it from Node.js

createAgentDeviceClient() gives Node.js code typed access to the same commands, as model tools in your own agent or from orchestration code:

import { createAgentDeviceClient } from 'agent-device';

const client = createAgentDeviceClient({ session: 'qa-run' });
try {
  await client.apps.open({ app: 'com.apple.Preferences', platform: 'ios' });
  const snapshot = await client.capture.snapshot({ interactiveOnly: true });
  const button = snapshot.nodes.find((node) => node.role === 'button');
  if (button) await client.interactions.press({ ref: button.ref });
} finally {
  await client.sessions.close();
}

See the Node.js API, the runnable examples, and the AI SDK and Eve integration guides.

What agents can do

  • Inspect app state through accessibility snapshots, refs, selectors, and React Native component trees.
  • Act on visible UI by tapping or pressing elements, filling fields, scrolling, making gestures, waiting, asserting state, and handling alerts.
  • Diagnose failures with screenshots, video, logs, traces, network data, performance samples, crash details, and React profiles.
  • Repeat workflows by saving working steps as .ad scripts for local use or CI. Export strict Maestro YAML when needed.

See Commands for the commands and evidence each target supports.

Diagram of the agentic development loop: humans assign tasks, agents write and review code, agent-device verifies mobile apps, pull requests receive evidence, and bugs or performance issues lead to fixes

What to ask your agent

With the CLI installed, prompts like these work end to end:

  • "Implement the onboarding screen, run it on the iOS simulator and Android emulator, and attach screenshots."
  • "Reproduce this crash and capture the logs that lead up to it."
  • "Check whether this change causes unnecessary React Native re-renders."
  • "Explore the checkout flow once, save it as a replay script, and run it in CI."
  • "Verify this pull request on a physical device and attach reviewable evidence."

Next steps

  • AI Agent Setup: skills, project rules, and per-client setup for Cursor, Codex, Claude Code, Windsurf, and others.
  • Quick Start: a guided run on the bundled Expo test app with screenshots, replay, and performance data.
  • Replay & E2E and Debugging & Profiling: repeatable tests and bug hunting.

Where to run agent-device

The same session and evidence model works at every step: the agent explores the app, captures evidence, saves a replay, runs it in CI, and moves onto remote devices.

PathBest forStart with
LocalTrying commands and debugging apps on simulators, emulators, physical devices, macOS, and Linux.Follow the Quick Start.
CI/CDAutomated pull request and merge validation with replay scripts and captured artifacts.Try the EAS workflow template.
Cloud / remoteLinux runners, managed devices, and remote jobs.Set up a remote proxy, connect a device cloud (BrowserStack, AWS Device Farm, Limrun), or contact Callstack for team QA.

How it works

agent-device keeps device state in sessions. It sends commands to XCTest on iOS and tvOS, ADB and the snapshot helper on Android, HDC and ArkUI uitest on HarmonyOS, Vega CLI/VDA on the Vega Virtual Device, a local helper on macOS, and AT-SPI on Linux.

Support depth varies by target. Newer backends such as HarmonyOS and Vega OS cover a subset of commands; run agent-device capabilities --platform <platform> to see what a target supports.

Sessions are scoped to the caller's git worktree, and host-local device claims stop parallel agents from taking over each other's simulators and emulators. The same commands drive hosted devices on BrowserStack, AWS Device Farm, and Limrun.

agent-device uses the inspect-act-verify process from Vercel's agent-browser for mobile, TV, and desktop apps. Basic --platform web support runs agent-browser in the same session and replay system.

FAQ

What is agent-device?

agent-device is a command-line tool and MCP server that lets AI coding agents inspect, control, and verify mobile apps and save evidence for review. It supports iOS, Android, HarmonyOS, TV, web, macOS, and Linux.

Is there an MCP server for mobile app automation?

Yes. agent-device mcp starts the official stdio MCP server. The Quick start above has the client config, and AI Agent Setup covers per-client details.

Does it work with React Native, Expo, Flutter, and native apps?

Yes. agent-device supports native iOS and Android apps, plus React Native, Expo, and Flutter apps on supported targets. The commands and evidence vary by target.

How is it different from mobile MCP servers?

The MCP server is one entry point to the same runtime used by the CLI and typed Node.js API. Sessions, device ownership, selectors, evidence, replay, CI workflows, and cloud routing stay consistent across all three.

Can I build my own agent or QA product on agent-device?

Yes. The typed Node.js client is a public surface over that same runtime, so an agent you build inherits everything above. Start from the Node.js API, AI SDK, or Eve guides.

How is it different from Appium, Detox, or Maestro?

With agent-device, an agent reads app state and chooses each command at run time. Teams use Appium, Detox, and Maestro to write and maintain test suites. agent-device can complement them by saving its runs as .ad scripts or exporting them as strict Maestro YAML.

Can agent-device run in CI?

Yes. Record a run as an .ad script, replay it in CI, and keep the screenshots and logs as artifacts; the EAS workflow template is a working example.

Articles and videos

Articles

Videos

Who uses agent-device?

Teams and developers at Callstack, JPMorgan Chase, Expensify, Shopify, Kindred, Total Wine & More, LegendList, HerLyfe, App & Flow, and others use agent-device.

Documentation

Contributing

See CONTRIBUTING.md.

Made at Callstack

agent-device is open source under the MIT license. Visit agent-device.dev or contact Callstack.

Related MCP servers

Control a real Chrome browser to complete any task: fill forms, extract data, book flights.

110k
Python
MIT
View repository →

Real-time global intelligence: markets, conflicts, country risk, energy, and infrastructure monitoring via 39 MCP tools.

83k
TypeScript
AGPL-3.0
View repository →

Netdata

Active

Real-time infrastructure monitoring with per-second metrics, ML-powered anomaly detection, and zero-configuration setup.

80k
Go
GPL-3.0
View repository →

Trending hip-hop artist momentum scores across four cultural dimensions.

79k
TypeScript
MIT
View repository →

AI orchestration platform with 100+ agents, swarm coordination, and self-learning memory for enterprise development.

68k
TypeScript
MIT
View repository →

Web scraping with stealth HTTP, real browsers, and Cloudflare bypass capabilities.

67k
Python
BSD-3-Clause
View repository →