PluginBench
Skill
Review
Audit score 70

core-agent-browser

actionbook/rust-skills

Internal browser automation CLI for agent workflows requiring interactive page testing and form filling.

What is core-agent-browser?

core-agent-browser is an internal support skill providing CLI-based browser automation for rust-learner, docs-researcher, and crate-researcher agents. Use it only when actionbook selectors are unavailable or interactive testing, screenshots, and form filling are explicitly required.

  • Navigate pages and manage browser state (open, back, forward, reload, close)
  • Capture page snapshots with interactive element references for reliable interaction
  • Click, fill, type, hover, check, select, and scroll elements using snapshot-derived refs
  • Press keys and key combinations, wait for elements or network conditions
  • Take full-page and targeted screenshots
  • Retrieve element text, values, page titles, and URLs

How to install core-agent-browser

npx skills add https://github.com/actionbook/rust-skills --skill core-agent-browser
Prerequisites
  • agent-browser CLI installed and available in PATH
Claude Code
Cursor
Windsurf
Cline

How to use core-agent-browser

  1. 1.Run `agent-browser open <url>` to navigate to the target page
  2. 2.Execute `agent-browser snapshot -i` to retrieve interactive elements with reference IDs (e.g., @e1, @e2)
  3. 3.Interact with elements using their refs: `agent-browser click @e1`, `agent-browser fill @e2 "text"`
  4. 4.Use `agent-browser wait` to pause for elements or network conditions if needed
  5. 5.Re-snapshot after navigation or DOM changes to get updated element references
  6. 6.Close the browser with `agent-browser close` when done

Use cases

Good for
  • Testing form submission workflows when actionbook has no pre-computed selectors
  • Automating interactive browser tasks like filling multi-step forms or clicking dynamic elements
  • Capturing screenshots of rendered pages for verification or debugging
  • Waiting for network idle or specific elements to appear before proceeding
  • Interacting with sites where static selector extraction is insufficient
Who it's for
  • Rust learning agents (rust-learner)
  • Documentation research agents (docs-researcher)
  • Crate information research agents (crate-researcher)
  • Agent developers needing fallback browser automation when MCP selectors unavailable

core-agent-browser FAQ

When should I use core-agent-browser vs. actionbook or rust-learner?

Use actionbook first for pre-computed selectors on known sites, then rust-learner for orchestrated Rust/crate lookups. Use core-agent-browser only as a last resort when actionbook has no selectors and you need direct browser automation, screenshots, or form filling.

How do I interact with elements after taking a snapshot?

The snapshot returns element references like @e1, @e2, etc. Use these refs in commands: `agent-browser click @e1`, `agent-browser fill @e2 "value"`, `agent-browser select @e1 "option"`.

What should I do if the page changes after my interaction?

Re-run `agent-browser snapshot -i` to get updated element references, since the DOM may have changed and old refs may no longer be valid.

Can I use semantic locators instead of refs?

Yes. You can use semantic locators like `agent-browser find role button click --name "Submit"` or `agent-browser find text "Sign In" click` as an alternative to snapshot refs.

Is this skill user-invocable?

No. core-agent-browser is an internal support skill for agent workflows only and cannot be directly invoked by users.

Full instructions (SKILL.md)

Source of truth, from actionbook/rust-skills.


name: core-agent-browser description: "Internal support skill for agent-browser CLI workflows used by rust-learner, docs-researcher, and crate-researcher. Use only when browser automation is explicitly required." user-invocable: false disable-model-invocation: true

Browser Automation with agent-browser

Priority Note

For fetching Rust/crate information, use this priority order:

  1. rust-learner skill - Orchestrates actionbook + browser-fetcher
  2. actionbook MCP - Pre-computed selectors for known sites
  3. agent-browser CLI - Direct browser automation (last resort)

Use agent-browser directly only when:

  • actionbook has no pre-computed selectors for the target site
  • You need interactive browser testing/automation
  • You need screenshots or form filling

Quick start

agent-browser open <url>        # Navigate to page
agent-browser snapshot -i       # Get interactive elements with refs
agent-browser click @e1         # Click element by ref
agent-browser fill @e2 "text"   # Fill input by ref
agent-browser close             # Close browser

Core workflow

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

Commands

Navigation

agent-browser open <url>      # Navigate to URL
agent-browser back            # Go back
agent-browser forward         # Go forward
agent-browser reload          # Reload page
agent-browser close           # Close browser

Snapshot (page analysis)

agent-browser snapshot        # Full accessibility tree
agent-browser snapshot -i     # Interactive elements only (recommended)
agent-browser snapshot -c     # Compact output
agent-browser snapshot -d 3   # Limit depth to 3

Interactions (use @refs from snapshot)

agent-browser click @e1           # Click
agent-browser dblclick @e1        # Double-click
agent-browser fill @e2 "text"     # Clear and type
agent-browser type @e2 "text"     # Type without clearing
agent-browser press Enter         # Press key
agent-browser press Control+a     # Key combination
agent-browser hover @e1           # Hover
agent-browser check @e1           # Check checkbox
agent-browser uncheck @e1         # Uncheck checkbox
agent-browser select @e1 "value"  # Select dropdown
agent-browser scroll down 500     # Scroll page
agent-browser scrollintoview @e1  # Scroll element into view

Get information

agent-browser get text @e1        # Get element text
agent-browser get value @e1       # Get input value
agent-browser get title           # Get page title
agent-browser get url             # Get current URL

Screenshots

agent-browser screenshot          # Screenshot to stdout
agent-browser screenshot path.png # Save to file
agent-browser screenshot --full   # Full page

Wait

agent-browser wait @e1                     # Wait for element
agent-browser wait 2000                    # Wait milliseconds
agent-browser wait --text "Success"        # Wait for text
agent-browser wait --load networkidle      # Wait for network idle

Semantic locators (alternative to refs)

agent-browser find role button click --name "Submit"
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "user@test.com"

Example: Form submission

agent-browser open https://example.com/form
agent-browser snapshot -i
# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]

agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i  # Check result