PluginBench
MCP Server
Active
MIT

io.github.Lykhoyda/ask-llm MCP Server

io.github.Lykhoyda/ask-llm

Get a second opinion from a different LLM—Codex, Claude, Grok, Ollama, Gemini, or Antigravity—without leaving your editor.

What is the io.github.Lykhoyda/ask-llm MCP server?

Ask LLM is an MCP server that lets you route code reviews and questions to multiple LLM providers (Codex, Claude, Grok, Ollama, Gemini, Antigravity) from Claude Code, Cursor, Codex CLI, Claude Desktop, or 40+ other MCP clients. It enables independent second opinions on diffs, architecture, and code without copying between tools.

Ask LLM gives your primary AI assistant a second opinion from a different model. Send a diff, architecture proposal, or code snippet to Codex, Claude, Grok, Ollama, Gemini, or Antigravity—each with different strengths (Codex for code reasoning, Gemini for 1M+ token context, Ollama for offline privacy). One prompt, multiple reviewers in parallel, consensus findings highlighted. Works in Claude Code, Cursor Agent, Codex CLI, Claude Desktop, and Pi.

How to install io.github.Lykhoyda/ask-llm

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • XAI_API_KEY
    secret

    xAI API key for metered Grok requests; Ask LLM never enables billing or credits

  • ASK_GROK_HARNESS

    Explicit Grok harness: xai-api (default) or grok-cli; no failover

  • ASK_GROK_MODEL

    Exact selected-harness model override (API default grok-4.7; CLI default grok-4.7; no fallback)

  • ASK_GROK_REASONING_EFFORT

    Grok reasoning effort: low, medium, high, or xhigh (default: high)

  • ASK_GROK_MAX_OUTPUT_TOKENS

    xAI API output-token ceiling to bound accidental spend (default: 16384)

  • ASK_GROK_TIMEOUT_MS

    Timeout for xAI API requests in milliseconds (default: 600000)

  • CURSOR_API_KEY
    secret

    Optional Cursor CLI credential for ask-cursor-agent; alternative to agent login

  • ASK_CURSOR_TIMEOUT_MS

    Timeout for model-neutral Cursor Agent harness calls (default: 600000)

  • ASK_CLAUDE_MODEL

    Default Claude model alias or full model name (default: opus)

  • ASK_CLAUDE_FALLBACK_MODEL

    Claude CLI fallback when Opus is overloaded or unavailable (default: sonnet)

  • ASK_CLAUDE_TIMEOUT_MS

    Timeout for Claude CLI execution in milliseconds (default: 600000 = 10 minutes)

  • OLLAMA_HOST

    Ollama server address (default: http://localhost:11434)

  • ASK_ANTIGRAVITY_TIMEOUT_MS

    Timeout for Antigravity (agy) execution in milliseconds (default: 300000 = 5 minutes)

  • ASK_ANTIGRAVITY_SANDBOX

    Set to '0' to drop agy's --sandbox flag if it blocks --add-dir context reads (default: sandbox on)

  • GMCPT_TIMEOUT_MS

    Timeout for CLI execution in milliseconds (default: 300000 = 5 minutes)

  • GMCPT_LOG_LEVEL

    Log verbosity: debug, info, warn, error (default: warn)

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "ask-llm": {
      "command": "npx",
      "args": [
        "-y",
        "@ask-llm/mcp"
      ],
      "env": {
        "XAI_API_KEY": "<YOUR_XAI_API_KEY>",
        "ASK_GROK_HARNESS": "<YOUR_ASK_GROK_HARNESS>",
        "ASK_GROK_MODEL": "<YOUR_ASK_GROK_MODEL>",
        "ASK_GROK_REASONING_EFFORT": "<YOUR_ASK_GROK_REASONING_EFFORT>",
        "ASK_GROK_MAX_OUTPUT_TOKENS": "<YOUR_ASK_GROK_MAX_OUTPUT_TOKENS>",
        "ASK_GROK_TIMEOUT_MS": "<YOUR_ASK_GROK_TIMEOUT_MS>",
        "CURSOR_API_KEY": "<YOUR_CURSOR_API_KEY>",
        "ASK_CURSOR_TIMEOUT_MS": "<YOUR_ASK_CURSOR_TIMEOUT_MS>",
        "ASK_CLAUDE_MODEL": "<YOUR_ASK_CLAUDE_MODEL>",
        "ASK_CLAUDE_FALLBACK_MODEL": "<YOUR_ASK_CLAUDE_FALLBACK_MODEL>",
        "ASK_CLAUDE_TIMEOUT_MS": "<YOUR_ASK_CLAUDE_TIMEOUT_MS>",
        "OLLAMA_HOST": "<YOUR_OLLAMA_HOST>",
        "ASK_ANTIGRAVITY_TIMEOUT_MS": "<YOUR_ASK_ANTIGRAVITY_TIMEOUT_MS>",
        "ASK_ANTIGRAVITY_SANDBOX": "<YOUR_ASK_ANTIGRAVITY_SANDBOX>",
        "GMCPT_TIMEOUT_MS": "<YOUR_GMCPT_TIMEOUT_MS>",
        "GMCPT_LOG_LEVEL": "<YOUR_GMCPT_LOG_LEVEL>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • ask-llm — Unified orchestrator: pick a provider per call or fan out to every installed provider
  • multi-llm — Send one prompt to multiple providers in parallel; returns per-provider responses and usage in one call
  • ask-codex — Codex CLI with GPT-6 Astra model; supports ephemeral or persistent sessions
  • ask-claude — Claude Code CLI with Opus 5.5 model; native sessions and read-only workspace access
  • ask-grok — One-shot Grok prompt through xAI API or Grok CLI with exact model attribution
  • ask-antigravity — Google Antigravity for subscription-backed second opinion; experimental one-shot mode
  • ask-ollama — Local Ollama for fully private, zero-cost reviews with server-side conversation replay
  • ask-gemini — Gemini CLI with 1M+ token context and @ file syntax; live progressive output
  • ask-gemini-edit — Structured OLD/NEW code edit blocks from Gemini
  • fetch-chunk — Retrieve chunks from cached large Gemini responses
  • get-usage-stats — Per-session token totals, fallback counts, and per-provider/model breakdowns
  • diagnose — Self-diagnosis: Node version, PATH resolution, provider CLI presence and versions
  • ping — Connection test

Use cases

  • Review a diff or commit for security issues, bugs, or architectural problems from a second model's perspective
  • Debate an architecture proposal or design plan with independent critique and trade-off analysis
  • Get a second opinion on code before committing, catching issues your primary AI missed
  • Ingest and summarize large codebases (1M+ tokens) using Gemini or Antigravity
  • Keep code reviews completely private and offline by routing through local Ollama

io.github.Lykhoyda/ask-llm MCP server FAQ

What is Ask LLM?

Ask LLM is an MCP server that routes code reviews and questions to multiple LLM providers (Codex, Claude, Grok, Ollama, Gemini, Antigravity) from your editor. It gives your primary AI a second opinion from a different model without copying between tools.

Is Ask LLM free?

Ask LLM itself is free and open-source (MIT license). Using it requires authentication with at least one provider: Codex and Claude require their respective CLI accounts, Grok requires an xAI API key, Ollama is free and local, Gemini and Antigravity require Google subscriptions.

How do I install Ask LLM in Cursor or Claude?

Run `npm install -g @ask-llm/mcp`, then `ask-llm setup --host cursor` (Cursor) or `ask-llm setup --host claude-desktop` (Claude Desktop). It previews changes and asks before applying. Restart your editor to load the server.

Which provider should I use?

Codex excels at code reasoning, Claude gives an independent opinion from non-Claude hosts, Grok offers xAI's model, Ollama runs locally and privately, Gemini handles 1M+ token contexts, and Antigravity is Google's subscription-backed alternative.

Can I use Ask LLM offline?

Yes, with Ollama. Route reviews through a local Ollama instance running on your machine—nothing leaves your system and there are no API costs.

What hosts does Ask LLM support?

Claude Code, Cursor Agent, Codex CLI, Claude Desktop, OpenCode, Pi, and 40+ other MCP clients via standard STDIO configuration.

README (reference)

Source of truth, from the repository.

<div align="center">

Ask LLM

Give your AI coding assistant a second opinion — from a different model.

Claude Code, Codex CLI, Cursor, Claude Desktop, or any of 40+ MCP clients can call Codex, Claude, Grok, Antigravity, Ollama, or Gemini to review a diff, debate a plan, or catch the bug the first model missed. Standard MCP; no prompt hacks.

CI Release GitHub Release npm License: MIT

Quick Start · Choose a reviewer · Claude Code plugin · Docs · Packages

</div>
You:    ask codex to review src/auth.ts for security issues
Codex:  ⚠ verifyToken() compares tokens with === — not timing-safe (line 42)
        ⚠ the session cookie is missing a SameSite attribute
Claude: Good catches — applying both fixes to src/auth.ts.

One prompt. A second model reviews independently; your assistant applies the fix. No copy-pasting between tools.

Why a second opinion?

Your primary AI is confident, but confidence isn't correctness. A second model with no stake in the first answer catches what it glossed over.

You want to…Ask LLM does
Review a diffA different model analyzes your changes and surfaces issues your primary AI missed
Debate a planSend an architecture proposal for critique, alternatives, and trade-off analysis
Get a second opinion on codeHave another model review an approach independently before you commit to it
Read more than fitsGemini and Antigravity ingest whole codebases in one call (1M+ tokens)
Keep it localRoute reviews through Ollama when nothing can leave your machine
Compare models side by sidemulti-llm fans one prompt out to several providers in parallel

Quick Start

Prerequisites: Node.js 24+ (the current LTS), on Linux or macOS, and at least one provider CLI installed and authenticated (see Provider setup).

npm install -g @ask-llm/mcp

Claude Code

ask-llm setup --host claude

Setup previews the user-scope registration and asks before applying it. Then try: ask codex to review my last commit. Run ask-llm doctor if anything looks off. See command compatibility for -y, removal, and other hosts.

<details> <summary>Advanced: install split provider packages instead</summary>
claude mcp add --scope user codex -- npx -y @ask-llm/codex-mcp
claude mcp add --scope user grok -e XAI_API_KEY="$XAI_API_KEY" -- npx -y @ask-llm/grok-mcp
claude mcp add --scope user antigravity -- npx -y @ask-llm/antigravity-mcp
claude mcp add --scope user ollama -- npx -y @ask-llm/ollama-mcp
claude mcp add --scope user gemini -- npx -y @ask-llm/gemini-mcp
</details>

Cursor

After the global install above, run ask-llm setup --host cursor. It previews the user-scope change to ~/.cursor/mcp.json and asks before writing. Restart Cursor Agent to load the server. For a project-scoped or manual install, see the Cursor host guide.

Codex CLI

# ~/.codex/config.toml
[mcp_servers.ask-llm]
command = "npx"
args = ["-y", "@ask-llm/mcp"]

Want Codex to consult Claude specifically? codex mcp add claude -- npx -y @ask-llm/claude-mcp

Claude Desktop

After the global install above, run ask-llm setup --host claude-desktop. It previews the user-scope change to claude_desktop_config.json and asks before writing. Restart Claude Desktop to load the server. See command compatibility for removal, backups, and formats that require manual setup.

OpenCode

After the global install above, run ask-llm setup --host opencode. Setup changes only a plain JSON opencode.json; for JSONC it prints the entry to add manually. OpenCode registration has fixture coverage and awaits verification on an installed host.

Pi

Pi has no built-in MCP client, so it installs the host package instead, which registers native ask-* tools plus the shared skills:

pi install npm:@ask-llm/plugin

Then use /skill:codex-review, /skill:multi-review, /skill:compare, /skill:brainstorm, or just describe what you want. See the Pi host guide for trust, data-transfer, and compatibility details.

<details> <summary>Any other MCP client (STDIO)</summary>
{ "command": "npx", "args": ["-y", "@ask-llm/mcp"] }

Swap @ask-llm/mcp for @ask-llm/codex-mcp, @ask-llm/claude-mcp, @ask-llm/grok-mcp, @ask-llm/antigravity-mcp, @ask-llm/ollama-mcp, or @ask-llm/gemini-mcp to install a single provider.

</details>

Choose your reviewer

The unified @ask-llm/mcp server is the recommended install: one registration, every provider you have, and parallel fan-out via multi-llm. Each provider is also available standalone.

ProviderBest forModel (default → fallback)Requires
CodexCode reasoning, targeted reviews, architecture critiquegpt-6-astra → gpt-5.6-terraOpenAI/Codex account
ClaudeAn independent Claude opinion from Codex or another non-Claude hostopus (Opus 5.5) → sonnetClaude Code CLI; read-only workspace tools
GrokGrok 4.7 critique via xAI API or the official Grok CLIgrok-4.7, reasoning high (no fallback)XAI_API_KEY or Grok CLI; one explicit harness per call
AntigravitySubscription-backed second opinion; large-context readsgemini-3.1-pro → gemini-3.8-flash (--effort high)Google AI Pro/Ultra; agy CLI. Experimental, one-shot; MCP execution requires unisolated opt-in
OllamaPrivate, offline, zero-cost reviewqwen3.8:27b (no auto-fallback)Ollama running locally
GeminiWhole-codebase reads (1M+ tokens)gemini-3.1-pro-preview → gemini-3.8-flashEnterprise Gemini seat (see note)

Gemini CLI is enterprise-only since 2026-06-18. Google restricted Gemini CLI to Gemini Code Assist Standard/Enterprise seats; free, Google AI Pro, and Ultra accounts lost access. @ask-llm/gemini-mcp still installs, but non-enterprise accounts get actionable guidance instead of output. On a subscription plan, use Antigravity (Google's sanctioned successor, covered by AI Pro/Ultra), Codex, Claude, or Ollama. Announcement

Fallbacks fire only under each provider's documented conditions (quota for Gemini/Codex, overload for Claude, rate limit for Antigravity). Grok and Ollama never substitute a model: Grok sends the exact harness catalog ID unchanged, and Ollama returns a clear ollama pull error if the requested model isn't local. Full details in Model Selection.

The Ask LLM plugin (Claude Code, Cursor Agent, Pi)

MCP gives your assistant the tools. The canonical @ask-llm/mcp package also owns the workflows: slash-command reviews with a validation pipeline, multi-model brainstorming, and opt-in continuous pair review. @ask-llm/plugin remains a dependent bridge for existing installs.

/plugin marketplace add Lykhoyda/ask-llm
/plugin install ask-llm@ask-llm-plugins
CommandWhat it does
<nobr>/multi-review</nobr>Parallel Antigravity + Codex review with a 4-phase validation pipeline and consensus highlighting
<nobr>/codex-review</nobr> · <nobr>/gemini-review</nobr> · <nobr>/ollama-review</nobr> · <nobr>/antigravity-review</nobr>Single-provider reviews with confidence filtering
<nobr>/sol-review</nobr>Model-pinned GPT-6 Sol review through Codex
<nobr>/grok-review</nobr>Metered Grok review through xAI with exact model attribution and no fallback
<nobr>/fable-review</nobr>Isolated, read-only review that requests the native Fable model and discloses runtime verification limits
<nobr>/brainstorm</nobr>Claude Opus researches your real files in parallel with external providers, then synthesizes, weighting verified findings higher. Also supports an exact no-Gemini Grok + GPT-6 Sol panel routed through Cursor Agent
<nobr>/compare</nobr>Raw side-by-side answers from multiple providers, no synthesis
<nobr>codex-pair</nobr>Opt-in continuous review on gpt-6-sol at medium effort: Codex checks every Edit/Write/MultiEdit when a .codex-pair/context.md marker is present

Review agents follow a 4-phase pipeline inspired by Anthropic's code-review plugin: context gathering, prompt construction with explicit false-positive exclusions, synthesis, and source-level validation of each finding.

<details> <summary>Host support matrix</summary>

@ask-llm/mcp owns the canonical host assets and skill corpus; @ask-llm/plugin is a dependent compatibility bridge. Claude Code loads its marketplace agents and hooks; Cursor Agent loads the adapted /codex-pair and /grok-pair skills through Agent Skills plus mcp.json (agent --plugin-dir ./packages/llm-mcp; see the Cursor Agent host guide); Pi loads explicit native tools, portable skill adapters, and a thin lifecycle extension.

CapabilityClaude CodeCursor AgentCodex CLI hostPi
Provider transportMCPMCP (mcp.json, unified ask-llm only)MCPnative Ask LLM tools (no built-in MCP)
Review/compare/brainstorm skillsyesAgent Skillstools only/skill:<name> + natural language
Isolated reviewer contexts / Fableyesno; fable-review excludednono; fable-review excluded
codex-pairhookson-demand persisted sessionnolifecycle extension
/grok-pairyes (explicit Cursor/xAI/CLI route)direct xAI/CLI routes via pinned unified ask-llm (or user-installed ask-grok)noexcluded
Blocking HIGH Stop gateopt-innonono; surfaced non-blockingly
Async pairing in one-shot printn/aon-demand skillnounsupported

Pi specifics: codex-pair requires the repository marker, Pi project trust, and interactive user-owned consent via /codex-pair; a committed marker alone never authorizes source transfer or cost. Pi surfaces findings non-blockingly and does not claim Claude's blocking Stop gate or one-shot print parity. fable-review is Claude Code-only. Provider CLI authentication is separate from Pi's host-model login. Update or remove with pi update npm:@ask-llm/plugin / pi remove npm:@ask-llm/plugin.

</details>

See the plugin docs for hooks, agents, and configuration.

MCP tools

ToolPackagePurpose
ask-llm@ask-llm/mcpUnified orchestrator: pick a provider per call, or fan out to every installed provider
multi-llm@ask-llm/mcpSend one prompt to multiple providers in parallel; returns per-provider responses and usage in one call
ask-codex@ask-llm/codex-mcpCodex CLI. GPT-6 Astra with Terra fallback. Omit sessionId for ephemeral use, or pass sessionId: "" first to persist and resume
ask-claude@ask-llm/claude-mcpClaude Code CLI. opus (Opus 5.5) with Sonnet fallback; native sessions; Read/Glob/Grep-only workspace access
ask-grok@ask-llm/grok-mcpOne-shot Grok prompt through explicit xai-api (default) or grok-cli; exact harness model ID; no harness/model fallback
ask-cursor-agent@ask-llm/mcpModel-neutral Cursor Agent harness: separate provider (claude, codex, gemini, grok) + exact Cursor catalog model verified against that family; read-only ask mode; no force/trust/spend changes or fallback
ask-antigravity@ask-llm/antigravity-mcpGoogle Antigravity (agy) for a subscription-backed second opinion. Experimental; one-shot
ask-ollama@ask-llm/ollama-mcpLocal Ollama. Fully private, zero cost. Server-side conversation replay via sessionId
ask-gemini@ask-llm/gemini-mcpGemini CLI with @ file syntax. 1M+ token context. Live progressive output via stream-json
ask-gemini-edit@ask-llm/gemini-mcpStructured OLD/NEW code edit blocks from Gemini
fetch-chunk@ask-llm/gemini-mcpRetrieve chunks from cached large responses
get-usage-statsallPer-session token totals, fallback counts, breakdowns by provider/model. In-memory only
diagnose@ask-llm/mcpSelf-diagnosis: Node version, PATH resolution, provider CLI presence and versions. Read-only
pingallConnection test

Session-capable ask-* tools accept an optional sessionId and return a structured AskResponse (provider, response, model, sessionId, usage) via MCP outputSchema alongside the human-readable text. Codex requires sessionId: "" on the first call for a resumable thread. The orchestrator also exposes usage://current-session as an MCP Resource for live JSON snapshots.

Things to say

ask codex to review the changes in src/auth.ts for security issues
ask claude for an independent opinion on this architecture        (from Codex or another non-Claude client)
ask antigravity to debate the plan in docs/design.md
ask ollama to explain src/config.ts                                 (runs locally, nothing leaves your machine)
ask gemini to summarize @. the current directory                    (1M+ context; @ syntax is Gemini-only)
use multi-llm to compare what codex and grok think about this approach

More patterns in How to Ask and Multi-Turn Sessions.

CLI

The @ask-llm/mcp binary (ask-llm-mcp) starts the MCP server when run with no arguments. With arguments it's a CLI. For host registration (ask-llm setup, ask-llm remove) and host diagnostics, use the separate ask-llm command; see command compatibility.

# Diagnose your setup: Node version, PATH, provider CLI versions, env vars
npx @ask-llm/mcp doctor                       # human-readable
npx @ask-llm/mcp doctor --json                # full JSON, exit 1 on error
npx @ask-llm/mcp doctor --format toon         # bounded, versioned agent-facing TOON pilot
npx @ask-llm/mcp doctor --format toon --full  # full TOON escape hatch

# Interactive multi-provider REPL: switch providers, persist sessions, watch usage live
npx @ask-llm/mcp repl

The REPL keeps a session per provider (/provider codex, /new, /sessions, /usage) and inherits all executor behavior: quota fallback, stream-json output for Gemini, native session resume.

Provider setup

Install and authenticate whichever providers you want to consult. The unified server detects what's present.

ProviderSetup
Codex CLIInstall and sign in
Claude Code CLIInstall and sign in (for Codex or other clients consulting Claude)
xAI API / Grok CLISet XAI_API_KEY for the default metered harness, or install and authenticate official Grok Build and pin harness: "grok-cli" per request (or set ASK_GROK_HARNESS=grok-cli). No failover between harnesses
Cursor CLIOptional model-neutral harness. Authenticate and pick an exact ID from agent --list-models
Antigravity CLI (agy)Version >= 1.1.5, logged in once (Google AI Pro/Ultra). Verify with agy --version
OllamaRunning locally with a model pulled: ollama pull qwen3.8:27b
Gemini CLInpm install -g @google/gemini-cli && gemini login. Enterprise-gated since 2026-06-18

Packages

PackageWhat it isVersionDownloads
@ask-llm/mcpCanonical package: unified MCP server, provider executors, and Claude Code, Cursor Agent, and Pi host assetsnpmdownloads
@ask-llm/pluginDependent bridge for existing plugin and Pi installationsnpmdownloads
@ask-llm/codex-mcpCodex-only MCP servernpmdownloads
@ask-llm/claude-mcpClaude-only MCP servernpmdownloads
@ask-llm/grok-mcpGrok-only MCP servernpmdownloads
@ask-llm/antigravity-mcpAntigravity-only MCP servernpmdownloads
@ask-llm/ollama-mcpOllama-only MCP servernpmdownloads
@ask-llm/gemini-mcpGemini-only MCP servernpmdownloads
<details> <summary>Migrating from the old package names</summary>

All public MCP packages now live in the @ask-llm npm organization. The old names are deprecated, but executable names are unchanged: update the package argument in your MCP config and commands such as ask-codex-mcp and ask-llm-mcp doctor keep working after a global install.

Old packageUse instead
ask-llm-mcp@ask-llm/mcp
ask-codex-mcp@ask-llm/codex-mcp
@anton-lykhoyda/ask-claude-mcp@ask-llm/claude-mcp
ask-antigravity-mcp@ask-llm/antigravity-mcp
ask-ollama-mcp@ask-llm/ollama-mcp
ask-gemini-mcp@ask-llm/gemini-mcp

The installation guide has the complete package-to-executable mapping.

</details>

Documentation

Contributing

Contributions are welcome. Start with the open issues and CONTRIBUTING.md.

License

MIT. See LICENSE.

Disclaimer: Ask LLM is an unofficial, third-party tool and is not affiliated with, endorsed, or sponsored by Anthropic, Google, OpenAI, or xAI.

Related MCP servers

Bridge Claude with local Ollama LLMs for private AI-to-AI collaboration — no API keys, fully local

18
TypeScript
MIT
View repository →

Bridge Claude with Google's Antigravity CLI for independent code review and second opinions from a different model.

18
TypeScript
MIT
View repository →

Get a second opinion from a different AI model on code reviews, architecture, and diffs.

18
TypeScript
MIT
View repository →

Get a second opinion on code from a different AI model—Claude, Codex, Grok, Gemini, Ollama, or Antigravity.

18
TypeScript
MIT
View repository →

McDonald's China MCP Server with event calendar, coupon inquiries and redemption features.

Build native Linux packages (.deb/.rpm/.apk/.pkg.tar.zst) from a single PKGBUILD via MCP.

22
Go
GPL-3.0
View repository →