PluginBench
MCP Server
Active
MIT

io.github.HarimxChoi/google-surf-mcp MCP Server

io.github.HarimxChoi/google-surf-mcp

Google search via Playwright with a warm Chrome profile—no API key, no proxies, no solver.

What is the io.github.HarimxChoi/google-surf-mcp MCP server?

The google-surf-mcp MCP server provides Google search capabilities through Playwright and a persistent Chrome profile, requiring no API keys or proxies. It combines search, URL extraction, and academic PDF parsing in a single tool, with built-in CAPTCHA recovery and intelligent ad/panel filtering.

google-surf-mcp replaces three separate MCPs (search + URL fetcher + academic extractor) by bundling Google search, web content extraction, and academic PDF parsing into one server. It uses a warm Chrome profile for fast, reliable searches (~1.5s per query), handles CAPTCHAs via OS notifications or headless recovery, and filters out sponsored ads and knowledge panels using geometric verification. Ideal for local research workflows or AI agents that need real-time web data without external API dependencies.

How to install io.github.HarimxChoi/google-surf-mcp

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • CHROME_PATH

    Absolute path to the Chrome binary. Auto-detected on Windows/macOS/Linux when omitted.

  • SURF_PROFILE_ROOT

    Directory for the warm Chrome profile. Defaults to ~/.google-surf-mcp.

  • SURF_LOCALE

    Browser locale, e.g. en-US.

  • SURF_TZ

    IANA timezone, e.g. America/New_York. Defaults to system timezone.

  • SURF_HEADLESS

    Set to 'false' to run Chrome visibly (demos/debugging). Defaults to true. CAPTCHA recovery always runs visible regardless.

  • SURF_IDLE_CLOSE_MS

    Idle ms before closing the sequential ctx and pool. 0 disables idle auto-close. Defaults to 30000.

  • SURF_ALLOW_PRIVATE

    Set to 'true' to allow extract on private/loopback addresses (localhost, 10.x, 192.168.x, 169.254.x, etc). Default blocks them as an SSRF guard.

Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "google-surf-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "google-surf-mcp"
      ],
      "env": {
        "CHROME_PATH": "<YOUR_CHROME_PATH>",
        "SURF_PROFILE_ROOT": "<YOUR_SURF_PROFILE_ROOT>",
        "SURF_LOCALE": "<YOUR_SURF_LOCALE>",
        "SURF_TZ": "<YOUR_SURF_TZ>",
        "SURF_HEADLESS": "<YOUR_SURF_HEADLESS>",
        "SURF_IDLE_CLOSE_MS": "<YOUR_SURF_IDLE_CLOSE_MS>",
        "SURF_ALLOW_PRIVATE": "<YOUR_SURF_ALLOW_PRIVATE>"
      }
    }
  }
}
Cursor
~/.cursor/mcp.json
{
  "mcpServers": {
    "google-surf-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "google-surf-mcp"
      ],
      "env": {
        "CHROME_PATH": "<YOUR_CHROME_PATH>",
        "SURF_PROFILE_ROOT": "<YOUR_SURF_PROFILE_ROOT>",
        "SURF_LOCALE": "<YOUR_SURF_LOCALE>",
        "SURF_TZ": "<YOUR_SURF_TZ>",
        "SURF_HEADLESS": "<YOUR_SURF_HEADLESS>",
        "SURF_IDLE_CLOSE_MS": "<YOUR_SURF_IDLE_CLOSE_MS>",
        "SURF_ALLOW_PRIVATE": "<YOUR_SURF_ALLOW_PRIVATE>"
      }
    }
  }
}
Windsurf
~/.codeium/windsurf/mcp_config.json
{
  "mcpServers": {
    "google-surf-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "google-surf-mcp"
      ],
      "env": {
        "CHROME_PATH": "<YOUR_CHROME_PATH>",
        "SURF_PROFILE_ROOT": "<YOUR_SURF_PROFILE_ROOT>",
        "SURF_LOCALE": "<YOUR_SURF_LOCALE>",
        "SURF_TZ": "<YOUR_SURF_TZ>",
        "SURF_HEADLESS": "<YOUR_SURF_HEADLESS>",
        "SURF_IDLE_CLOSE_MS": "<YOUR_SURF_IDLE_CLOSE_MS>",
        "SURF_ALLOW_PRIVATE": "<YOUR_SURF_ALLOW_PRIVATE>"
      }
    }
  }
}
VS Code
.vscode/mcp.json
{
  "servers": {
    "google-surf-mcp": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y",
        "google-surf-mcp"
      ],
      "env": {
        "CHROME_PATH": "<YOUR_CHROME_PATH>",
        "SURF_PROFILE_ROOT": "<YOUR_SURF_PROFILE_ROOT>",
        "SURF_LOCALE": "<YOUR_SURF_LOCALE>",
        "SURF_TZ": "<YOUR_SURF_TZ>",
        "SURF_HEADLESS": "<YOUR_SURF_HEADLESS>",
        "SURF_IDLE_CLOSE_MS": "<YOUR_SURF_IDLE_CLOSE_MS>",
        "SURF_ALLOW_PRIVATE": "<YOUR_SURF_ALLOW_PRIVATE>"
      }
    }
  }
}
Claude Code
claude mcp add google-surf-mcp --env CHROME_PATH=<YOUR_CHROME_PATH> --env SURF_PROFILE_ROOT=<YOUR_SURF_PROFILE_ROOT> --env SURF_LOCALE=<YOUR_SURF_LOCALE> --env SURF_TZ=<YOUR_SURF_TZ> --env SURF_HEADLESS=<YOUR_SURF_HEADLESS> --env SURF_IDLE_CLOSE_MS=<YOUR_SURF_IDLE_CLOSE_MS> --env SURF_ALLOW_PRIVATE=<YOUR_SURF_ALLOW_PRIVATE> -- npx -y google-surf-mcp

Tools & capabilities

Tools this server exposes to the agent.

  • searchSingle Google search query returning title, URL, and snippet. Drops sponsored ads and knowledge panels. Results cached 24h by default.
  • search_parallelParallel Google search for up to 10 queries using a pool of 4 workers. Returns same format as search.
  • extractFetch and extract content from a URL. Supports three modes: 'full' (whole body), 'abstract' (~1500 chars), and 'metadata' (PDF page count only). Handles HTML via Readability and PDFs via spatial parsing.
  • search_extractCombined search + parallel extract in one call. Defaults to 'abstract' mode for cheap triage (~1500 chars per result); use 'full' mode for complete article text.
  • healthServer status endpoint. Returns cascade, pool, rate limiter, cache, telemetry, and self-healing strategy stats.

Use cases

  • Research and gather real-time information from Google search without API keys or rate limits
  • Extract and summarize academic papers from arxiv, bioRxiv, Nature, OpenReview, NeurIPS, JMLR, and other repositories inline
  • Triage search relevance quickly with abstract mode (~1500 chars) before fetching full article text
  • Perform parallel searches across multiple queries in ~1.5 seconds wall time
  • Build AI agents that need web search + content extraction without managing external APIs or CAPTCHA solvers

io.github.HarimxChoi/google-surf-mcp MCP server FAQ

What is google-surf-mcp and how does it differ from other search MCPs?

google-surf-mcp is a Google search MCP that uses Playwright and a persistent Chrome profile to perform real searches without API keys, proxies, or CAPTCHA solvers. It combines search, URL extraction, and academic PDF parsing in one tool, replacing three separate MCPs. The author tested 6 free Google search MCPs and found all failed; this one is designed to actually work.

Is google-surf-mcp free to use?

Yes, it is completely free. It requires no API key, no proxy service, and no CAPTCHA solver. You only need Node 18+ and Google Chrome (or Chromium) installed locally.

How do I install google-surf-mcp in Claude or Cursor?

Add this to your ~/.claude.json (or equivalent config for your MCP client): {"mcpServers": {"google-surf": {"command": "npx", "args": ["-y", "google-surf-mcp"]}}}. Restart your client and the tools will be available. Alternatively, clone the repo and point to the local build/index.js file.

What happens when Google shows a CAPTCHA?

google-surf-mcp has 4 CAPTCHA recovery modes: (1) OS notification + headed Chrome window (default for local use), (2) headed Chrome only (SURF_HEADLESS=false), (3) remote DevTools (SURF_REMOTE_DEBUG=true for headless servers), or (4) fail-fast error (SURF_CLOUD_MODE=true for serverless). Each solve preserves the profile's reputation with Google.

How fast is google-surf-mcp?

Sequential searches take ~1.5s per query (first call ~4s including setup). Parallel searches with 4 workers complete ~10 queries in ~1.5s wall time. search_extract with abstract mode (default) is ~3s for 5 results; full mode is ~5s.

Does google-surf-mcp extract academic papers?

Yes. It automatically extracts academic PDFs inline from arxiv, bioRxiv, Nature, OpenReview, NeurIPS, JMLR, PMLR, Springer, and PubMed (via PMC). The extract tool supports 'abstract' mode (~1500 chars, token-cheap) and 'full' mode for complete text. Optional OCR is available for scanned PDFs.

README (reference)

Source of truth, from the repository.

<img src="./assets/icon256.png" width="128" align="right" alt="google-surf-mcp"/>

google-surf-mcp

English | 한국어

npm version npm downloads ci google-surf-mcp MCP server

demo

Demo only. Actual searches run headless by default (no visible browser). Set SURF_HEADLESS=false to make Chrome visible like in the clip above.

Google search MCP. No API key. Just works.

One MCP replaces three: search + URL fetcher + academic-paper extractor.

  • ✅ Actually works (tested 6 free Google search MCPs, all failed)
  • ✅ Search + URL + academic PDF extract in one MCP (replaces the search MCP + fetch MCP + academic-search MCP combo)
  • ✅ Academic PDFs extracted inline: arxiv, biorxiv, Nature, OpenReview, NeurIPS, JMLR, PMLR, Springer, PubMed (via PMC)
  • search_extract defaults to abstract mode (~1500 chars/result, token-cheap), mode="full" for whole bodies
  • ✅ Sponsored ads + knowledge panels dropped (geometric verification, not just text matching)
  • ✅ CAPTCHA recovery in 4 modes: OS notification (default) / SURF_HEADLESS=false / SURF_REMOTE_DEBUG / SURF_CLOUD_MODE (fail-fast)
  • ✅ No API key, no proxies, no solver

5 tools: search / search_parallel / extract / search_extract / health

What

Plug it into any MCP client and you get Google search as a tool.

No CAPTCHA solver. When CAPTCHA fires on any tool, a Chrome window opens for a human to solve. Each solve preserves the profile's reputation with Google.

First call auto-bootstraps the warm profile. Designed for local use. For headless / serverless environments set SURF_CLOUD_MODE=true (fail-fast on CAPTCHA, worker pool disabled).

Numbers

result
sequential~1.5s/query (first call ~4s, includes setup)
parallel x4~1.5s wall (first call ~9s, includes pool warm)
parallel x10~4.5s wall
search_extract x5 (abstract, default)~3s wall
search_extract x5 (full)~5s wall (search + 5 parallel extracts)

Measured on a workstation with a 1Gb/s connection.

Stack

  • Playwright + persistent Chrome profile
  • playwright-extra stealth as a cascade fallback tier
  • Multi-strategy SERP parser + geometric verification (drops sponsored / knowledge_panel / related)
  • @llamaindex/liteparse for PDF text extraction (PDFium spatial parsing, optional OCR); Mozilla Readability + Turndown for HTML
  • Resource-blocked images / media / fonts for speed
  • Auto-bootstrap on first call; pool falls back to single-context after repeated warm failures
  • Self-healing: runtime parser-strategy reorder (deterministic) + daily cron repair PR (synthesis → optional LLM → triple-gate validation, human review)

Install

Requires Node 18+ and Google Chrome (or Chromium) on the system.

npx google-surf-mcp   # actual MCP - register in client config

First tool call auto-bootstraps the warm profile (you may see Chrome open briefly).

Or local clone:

git clone https://github.com/HarimxChoi/google-surf-mcp
cd google-surf-mcp
npm install

If auto-bootstrap fails (rare), run it manually:

npm run bootstrap

Override paths if needed:

CHROME_PATH=/path/to/chrome SURF_TZ=America/New_York npm run bootstrap

Use with Claude Code

Paste this into your ~/.claude.json:

{
  "mcpServers": {
    "google-surf": {
      "command": "npx",
      "args": ["-y", "google-surf-mcp"]
    }
  }
}

Restart Claude Code. Done. search, search_parallel, extract, search_extract, health are now available.

For other MCP clients, use the same JSON shape in their config file.

Local clone variant:

{
  "mcpServers": {
    "google-surf": {
      "command": "node",
      "args": ["/abs/path/to/google-surf-mcp/build/index.js"]
    }
  }
}

Tools

  • search(query, limit?) - single query, ~1.5s. Returns title / url / snippet. Sponsored ads + knowledge-panel dropped (response includes dropped count + dropped_reasons). Results cached 24h (SURF_CACHE_TTL_SEARCH_MS=0 to bypass).
  • search_parallel(queries[], limit?) - pool of 4, max 10 queries per call.
  • extract(url, max_chars?, mode?) - fetch a URL, return article content.
    • mode="full" (default): whole body. HTML via Readability, PDFs via liteparse (spatial parsing, multi-column reading order).
    • mode="abstract": ~1500-char survey (PDF page 1 or HTML meta description). Triage relevance before paying for full text.
    • mode="metadata": PDF page count only.
    • Response: content, title, excerpt, length, is_pdf, page_count, extraction_quality. Failures return { error }, never throw.
  • search_extract(query, limit?, max_chars?, mode?) - search + parallel extract in one call. Default mode="abstract" returns SERP enriched with ~1500-char summaries (cheap triage). Use mode="full" when you actually need the article texts (slower, more tokens).
  • health() - server status. Response: cascade / pool (warmFailures + fallback) / rateLimiter / cache / telemetry / selfHealing (current strategy order + stats) / config. Call it if searches start failing — pool.fallback=true or rising cascade.totalCaptchas are the usual culprits.

Env vars

vardefaultnotes
CHROME_PATHauto-detectedabsolute path to Chrome binary
SURF_PROFILE_ROOT~/.google-surf-mcpwhere the warm profile lives
SURF_LOCALEen-USbrowser locale
SURF_TZsystem tze.g. America/New_York
SURF_HEADLESStrueset false to run Chrome visibly (demos / debugging). When false, CAPTCHA recovery skips the OS notification (user is already watching).
SURF_REMOTE_DEBUGfalseset true on a headless server with remote DevTools. CAPTCHA path emits the DevTools port and throws instead of spawning a window; attach chrome://inspect from a local machine over SSH port-forward to solve.
SURF_IDLE_CLOSE_MS30000idle ms before closing the sequential ctx and pool. 0 disables idle auto-close. Lower = faster cleanup, higher = warmer cache for spaced-out calls.
SURF_ALLOW_PRIVATEfalseset true to allow extract to fetch private/loopback addresses (localhost, 127.0.0.1, 10.x, 192.168.x, 169.254.x, etc). Default blocks them as an SSRF guard.
SURF_EXTRACT_MAX_CHARS8000default extract truncation (200–50000); per-call max_chars still overrides
SURF_EXTRACT_OCRfalseOCR scanned/image PDFs via Tesseract (slower; off by default)
SURF_CLOUD_MODEfalseheadless/serverless mode: TLS bypass + --no-sandbox + --disable-dev-shm-usage + worker pool disabled + fail-fast on CAPTCHA
SURF_CASCADE_DISABLEDfalsepin a single stealth mode (chosen by SURF_USE_STEALTH) instead of the 3-tier auto-cascade
SURF_USE_STEALTHtrueinitial stealth tier — only consulted when SURF_CASCADE_DISABLED=true
SURF_HUMANLIKE_MODEoffoff / background (fire-and-forget after returning results) / inline (await before returning, slower) — opt-in humanlike browsing
SURF_RATE_LIMIT_PER_MIN10internal cap on Google-facing requests per minute
SURF_CACHE_TTL_SEARCH_MS86400000search cache TTL (24h); 0 disables caching
SURF_CACHE_MAX_ENTRIES1000LRU cap per cache namespace
SURF_CACHE_ROOT<profile>/cachecache directory
SURF_INSECURE_TLS=SURF_CLOUD_MODE--ignore-certificate-errors (auto-on in cloud mode)
SURF_NO_SANDBOX=SURF_CLOUD_MODE--no-sandbox (auto-on in cloud mode)
SURF_TELEMETRYfalseset true to enable jsonl event logging (search outcomes, cache hits/misses, tool errors, parser staleness) under {SURF_TELEMETRY_ROOT}. Designed as the input feed for the self-healing pipeline. Off by default.
SURF_TELEMETRY_ROOT<profile>/telemetrydirectory for jsonl telemetry files. UTC-dated one file per day (YYYY-MM-DD.jsonl).
SURF_SELF_HEALINGtrueper-strategy outcome tracking + persisted reordering. Healing must win by 3 outcomes before reorder kicks in, so single-call flapping is impossible. Set false to pin the default strategy order.
SURF_SELF_HEALING_FILE<profile>/.heal/strategy-order.jsonpersistence path for healing state. Atomic tmp+rename writes; debounced 5s.
SURF_LLM_HEALfalseopt-in for LLM-assisted selector repair in the workflow-only repairWithLLM helper. Off by default → no third-party LLM request ever fires. When true, requires ANTHROPIC_API_KEY (your own); the package never ships a maintainer key.
ANTHROPIC_API_KEYyour Anthropic key. Read only when SURF_LLM_HEAL=true. The runtime self-healing in SURF_SELF_HEALING is deterministic and never reads this variable.

Troubleshooting

  • CAPTCHA in 4 modes (picked automatically from env):
    • default (local desktop): OS notification fires, headed Chrome opens, human solves, call retries
    • SURF_HEADLESS=false: headed Chrome opens, no notification (user is already watching)
    • SURF_REMOTE_DEBUG=true: DevTools port + instructions printed, attach chrome://inspect locally to solve
    • SURF_CLOUD_MODE=true: fail-fast with CAPTCHA_REQUIRED error
  • Headed Chrome opens to a plain search box instead of CAPTCHA: just type any query in the box and press Enter. Subsequent calls work.
  • "Chrome not found": install Chrome or set CHROME_PATH.
  • Stale selectors: two-layer mitigation — runtime per-strategy reorder (SURF_SELF_HEALING, deterministic) + daily cron that opens draft PRs with candidate fixes (SURF_LLM_HEAL optional, human review required, never auto-merged).
  • Searches feel slower than the Numbers table: check health().pool.fallback. true means the worker pool gave up after 3 warm failures and is using a single context. Usually fixed by npm run bootstrap to refresh the seed profile.
  • SSRF: extract blocks localhost, private IPs, AWS metadata by default. Set SURF_ALLOW_PRIVATE=true to allow them.

Changelog

See CHANGELOG.md.

License

MIT

Related MCP servers

Give your AI agent stealth web scraping with Cloudflare bypass and CSS selection, powered by Scrapling.

67k
Python
BSD-3-Clause
View repository →

Give your AI coding agent full control of a live Chrome browser for automation, debugging, and performance analysis.

45k
TypeScript
Apache-2.0
View repository →

Let AI agents manage your Puter files, websites, and serverless workers over MCP.

43k
TypeScript
AGPL-3.0
View repository →

Browser automation for AI agents via MCP, powering ByteDance's Agent TARS hybrid GUI/DOM browser control.

37k
TypeScript
Apache-2.0
View repository →

Run arbitrary shell commands from an MCP-connected AI agent.

37k
TypeScript
Apache-2.0
View repository →

Filesystem access MCP server from ByteDance's UI-TARS/Agent TARS ecosystem.

37k
TypeScript
Apache-2.0
View repository →