runpod-mcp
runpod/runpod-plugins-official
Manage Runpod infrastructure—pods, endpoints, jobs, volumes, templates—via structured MCP tool calls.
What is runpod-mcp?
The Runpod MCP server exposes Runpod's control plane as typed tool calls, letting agents manage infrastructure without shell commands. Use it when MCP tools are connected in your session; it wraps Runpod's REST v2 API and is preferred over runpodctl for plain CRUD when available.
- List, create, update, start, stop, restart, and delete pods with structured parameters and error handling
- Manage serverless endpoints (QUEUE or LOAD_BALANCER type), workers, and releases; stream worker logs
- Run, monitor, and cancel serverless jobs; stream job output and health status
- Deploy Hub repositories (vLLM, ComfyUI, etc.) as endpoints with one call
- Manage templates, network volumes, container registry auth, and ECR delegations
- Query GPU/CPU catalog, data centers, and billing/usage breakdowns
How to install runpod-mcp
npx skills add https://github.com/runpod/runpod-plugins-official --skill runpod-mcp- Runpod API key (same key works for MCP and runpodctl/flash CLIs)
- MCP connection established via `claude mcp add` (hosted, OAuth, or local stdio)
- Verify connection in Claude Code with `/mcp` — runpod should show Connected
- Confirm a test call works (e.g., `list-endpoints`) before relying on MCP
How to use runpod-mcp
- 1.Connect the Runpod MCP server using your API key as a Bearer header or OAuth
- 2.Verify the connection is live by running `/mcp` in Claude Code and confirming runpod shows Connected
- 3.Consult the golden paths (runpod/golden-paths/README.md) for multi-step sequences (image → template → endpoint, pod → volume → serverless)
- 4.Call tools in the correct order: e.g., create template, then create endpoint from template, then invoke with jobs
- 5.For file transfer, SSH, or multi-GPU priority lists, fall back to runpodctl instead of MCP
Use cases
- Deploy a serverless endpoint from a Hub repository and invoke it with structured job calls
- Create a pod with specific GPU type, start it, stream logs, and clean up when done
- Set up a network volume, attach it to a pod, and monitor usage costs
- Manage multiple endpoints across regions and autoscale by updating endpoint configuration
- Authenticate private container registries and reference them in pod/endpoint creation
- ML engineers deploying models and managing inference infrastructure
- DevOps teams automating Runpod resource provisioning and monitoring
- Agents and scripts that need typed, structured API calls instead of shell commands
- Users already using runpodctl who want MCP-native tool integration in Claude Code or Cursor
runpod-mcp FAQ
Use MCP for infrastructure CRUD and serverless job calls when tools are connected. Use runpodctl for file transfer (send/receive), SSH key management, multi-GPU priority lists, or when you need a reproducible shell command.
No — Runpod's REST v2 (which MCP drives) has no CPU-endpoint concept. Use `runpodctl serverless create --compute-type CPU` for CPU endpoints instead.
Runpod's REST API returns 204 No Content on successful deletion, which MCP reports as an error. Don't treat it as failure — confirm deletion with a follow-up `get-` or `list-` call; the deleted resource will 404.
Use `set-endpoint-gpus` to pin a SKU on an existing endpoint. For new endpoints, `create-endpoint` only exposes `gpuPoolIds` and cannot express a specific SKU; use `deploy-hub-repo` with `gpuIds` exclusions if you need SKU control at creation time.
No — the tier is immutable. Choose the correct `volumeType` when calling `create-network-volume`; `update-network-volume` cannot change it.
Full instructions (SKILL.md)
Source of truth, from runpod/runpod-plugins-official.
name: runpod-mcp description: >- Manage Runpod infrastructure — pods, serverless endpoints, jobs, templates, network volumes, container-registry auth, GPU/CPU catalog, and billing — via the Runpod MCP server's structured tool calls. Use when the Runpod MCP tools (create-pod, list-endpoints, …) are connected in this session, or to connect them (hosted OAuth or local npx). Prefer this over runpodctl for plain infra CRUD when MCP is available; use runpodctl for the terminal, file transfer, or SSH setup. allowed-tools: Bash(npx:), Bash(claude:) compatibility: Linux, macOS, Windows metadata: author: runpod version: "1.4.0" # x-release-please-version license: Apache-2.0
Runpod MCP
The Runpod MCP server exposes Runpod's control plane as structured tool calls,
so an MCP-capable agent can manage infrastructure without shelling out. It is the
same Runpod REST API that runpodctl uses — pick MCP when its tools are
connected (typed params, structured errors, no shell quoting).
For a multi-step job, read the worked example before calling tools. Tool calls are easy to issue and easy to issue in the wrong order — the verified end-to-end sequences live in runpod/golden-paths/README.md (image → template → endpoint, pod → volume → serverless, multi-region, autoscaling, monitoring). This skill covers what each tool does; the paths cover what order to do them in and what it costs.
Connect
Connect the hosted server with your API key as a Bearer header if you also use runpodctl/flash — that one key auths the MCP and the CLIs (the 80% path):
claude mcp add --transport http runpod -s user https://mcp.getrunpod.io/ \
--header "Authorization: Bearer $RUNPOD_API_KEY"
Plain OAuth ("Sign in with Runpod", via npx @runpod/mcp-server@latest add) is MCP-only — the CLIs stay unauthed, so use it only for MCP-only work. Local stdio runs the server as a subprocess with your key. Those variants + the key-vs-OAuth tradeoff: reference/connect.md. After connecting, reconnect the client (in Claude Code, /mcp) so the tools load.
Verify it's live (do this before relying on MCP): in Claude Code run /mcp —
runpod should show Connected, not Needs authentication (if it's the latter,
sign in there first; the bundled plugin server registers the URL but stays inert
until you authenticate). Confirm a real call works by asking for list-endpoints.
If the runpod tools aren't present at all, the server isn't connected — (re)run the
install above, or fall back to runpodctl for this task.
Check the server version (which REST API it drives): the MCP initialize handshake
returns it in serverInfo.version. /mcp in Claude Code shows it, or probe the hosted
server directly:
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"0"}}}' \
| curl -s -X POST https://mcp.getrunpod.io/ -H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" -H "Authorization: Bearer $RUNPOD_API_KEY" -d @-
# → serverInfo.version e.g. "3.0.0 [RUNPOD_REST_VERSION=v2]" (verified 2026-07-29)
The MCP server drives Runpod's REST v2 internally (RUNPOD_REST_VERSION=v2), so most
tools avoid the buggy public rest.runpod.io/v1 control API. Two exceptions worth
knowing: the Hub, public-endpoint and set-endpoint-gpus tools go through GraphQL (so they
work under either REST version), and CPU serverless endpoints are not creatable through
MCP — v2 has no CPU-endpoint concept at all (create-endpoint requires gpuPoolIds), so
use runpodctl serverless create --compute-type CPU for those.
Prefer MCP or runpodctl over hand-rolled rest.runpod.io/v1 calls for creating endpoints.
Tool surface
Structured tools, grouped by resource:
- Pods — list, get, create, update, start, stop, restart, delete, stream logs.
- Serverless endpoints — list, get, create, update, delete; list workers; list releases; stream worker logs.
- Logs are no longer an MCP-only capability — runpodctl grew
pod logsandserverless logsin v2.10.0. MCP still returns already-parsed, bounded frames, which is the easier shape inside an agent; reach for the CLI when you are shell-only or want--follow. Job output streaming (stream-job) remains MCP-only. create-endpointtakesendpointType: QUEUE(default) orLOAD_BALANCER— see golden path 14. The routing type is fixed at creation;update-endpointcannot change it.- Read an endpoint's invoke URLs from
requestUrlson the get/list reply instead of assembling them. - To pin a specific GPU SKU on an existing endpoint use
set-endpoint-gpus;create-endpoint/update-endpointexpose onlygpuPoolIdsand can't express a SKU (deploy-hub-repocan pin one at deploy time viagpuIdsexclusions).
- Logs are no longer an MCP-only capability — runpodctl grew
- Jobs (serverless runtime) — run, runsync, status, stream, cancel, retry, health, purge queue.
- Hub —
list-hub-repos(public catalog of prebuilt Serverless workers and Pod templates: vLLM, ComfyUI, …) anddeploy-hub-repo, which deploys a repo's listed release as an endpoint — the same as clicking Deploy on the Hub. - Public endpoints —
list-public-endpoints: managed pay-per-use model APIs (text/image/video/audio) that need no deployment. Call the returned endpointId withrun-endpoint/runsync-endpoint. - Templates — list, get, create, update, delete.
- Network volumes — list, get, create, update, delete.
create-network-volumetakesvolumeType(STANDARD|HIGH_PERFORMANCE) and a size of 10–4096 GB; omitvolumeTypeto get the data center's default tier. The tier is immutable after creation —update-network-volumecan't change it. - Container registry auth — list, get, create, delete. A username + password for any registry; pass the resulting id as
containerRegistryAuthIdon create-pod/create-endpoint. - ECR delegations (
list-/create-/delete-registry-delegation) — AWS ECR only, v2 only, and stores no credentials: you register a repository ARN and Runpod gets scoped pull access instead. Prefer it over a stored username/password for ECR. The reply carries adockerRegistryUri— that's the image URI to deploy with. - Catalog — list/get GPU types, list/get CPU types, list/get data centers.
- Billing — scoped usage/cost breakdowns (
get-billing).
The tool list above is a map, not a contract. The server is the source of truth —
/mcp(or your client's tool list) shows exactly what the connected version exposes, and each tool carries its own parameter descriptions. Check there before assuming a capability exists or doesn't.
Delete tools (
delete-template,delete-pod, …) can returnisError: truewith "Unexpected end of JSON input" even on success — the Runpod REST API returns 204 No Content. Don't treat it as failure; confirm with a follow-upget-/list-(a deleted resource then 404s).
Use MCP vs runpodctl
- Use runpod-mcp when the tools are connected AND the task is infra CRUD or a serverless job call the server exposes. Cap large job/log output to a file.
- Use runpodctl instead for:
send/receivefile transfer, SSH key management,doctorsetup, model cache — or any shell-only agent, or when the user wants a reproducible command. - Hand pod creation to runpodctl for a multi-GPU priority list (MCP's v2
create-pod takes one GPU type; extra
gpuTypeIdsare dropped with a_warningon success), or for a template + CPU pod together —create-podrejects that combination, since a template deploy is GPU-and-v2-only. Each alone is fine in MCP:templateId(v2-only,imageNamethen optional, and each field you pass replaces the template's whole value rather than merging) orcomputeType: "CPU". - Not this lane: writing/deploying your own Python (→ flash); downloading models or building/pushing images (→ companion-clis).
For concepts (pods vs serverless, GPU selection, storage), read
../runpod-usage/.
Source & docs
- Server source: https://github.com/runpod/runpod-mcp
- Package (npm): https://www.npmjs.com/package/@runpod/mcp-server
- Hosted endpoint: https://mcp.getrunpod.io/
- Docs: https://docs.runpod.io
Related skills
More from runpod/runpod-plugins-official and the wider catalog.

runpod-migrate
Migrate a codebase from Runpod GraphQL or REST v1 to REST v2 with inventory, rewriting, and verification.

runpod-templates
Reference for Runpod's official prebuilt pod templates (ComfyUI, PyTorch, Ubuntu) — what ships, ports, paths, and first-boot gotchas.

runpod-usage
Runpod concepts and workflows: pods vs serverless, GPU selection, container building, and the agentic development loop.

runpodctl
Terminal CLI for managing Runpod GPU pods, serverless endpoints, volumes, and models.

companion-clis
Companion CLIs for Runpod workflows — HuggingFace, GitHub, Docker, and AWS.

flash
Deploy AI workloads on Runpod serverless GPUs/CPUs with hot-reload dev iteration and one-command shipping.