PluginBench
MCP Server
Active
Apache-2.0

GPU MCP Server MCP Server

io.github.pmady/gpu-mcp-server

Query NVIDIA GPU metrics—utilization, memory, temperature, power—directly from Claude or Cursor.

What is the GPU MCP Server MCP server?

The GPU MCP Server is an MCP server that exposes real-time NVIDIA GPU metrics as tools accessible to AI agents like Claude, Goose, and Cursor. It provides utilization, memory, temperature, power, PCIe, and NVLink throughput data without requiring Prometheus or dcgm-exporter, and supports Multi-Instance GPU (MIG) configurations.

This server lets AI agents query live GPU performance data on your machine. It connects directly to NVIDIA's NVML library via Go, exposing four tools for listing GPUs, fetching detailed metrics, inspecting GPU processes, and viewing aggregate statistics. Useful for agents that need to monitor hardware health, optimize workloads, or make decisions based on current GPU state.

How to install GPU MCP Server

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "gpu-mcp-server": {
      "command": "docker",
      "args": [
        "run",
        "-i",
        "--rm",
        "ghcr.io/pmady/gpu-mcp-server:v0.1.0"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • list_gpus — List all GPUs with utilization and memory info
  • get_gpu_metrics — Detailed metrics for a GPU by index or UUID, including utilization, memory, temperature, power, PCIe and NVLink throughput
  • get_gpu_processes — PID-level GPU process attribution
  • gpu_summary — Aggregate stats across all devices

Use cases

  • Monitor GPU utilization and memory in real-time while running AI workloads
  • Check GPU temperature and power draw to prevent thermal throttling
  • Attribute GPU usage to specific processes for debugging performance issues
  • Get aggregate GPU statistics across a multi-GPU system for capacity planning
  • Query MIG instance metrics when using NVIDIA's Multi-Instance GPU feature

GPU MCP Server MCP server FAQ

What does the GPU MCP Server do?

It exposes real-time NVIDIA GPU metrics (utilization, memory, temperature, power, PCIe/NVLink throughput) as MCP tools that AI agents like Claude and Cursor can call directly, without needing Prometheus or other metric pipelines.

Is it free?

Yes, it is open-source under the Apache 2.0 license.

How do I install it in Claude Desktop?

Add an entry to your claude_desktop_config.json with the command pointing to the gpu-mcp-server binary, or use the Docker image via the docker command.

How do I install it in Cursor?

Add the server to .cursor/mcp.json (project-level) or ~/.cursor/mcp.json (global) with type 'stdio' and the path to the binary.

Does it require authentication or API keys?

No, it runs locally and communicates directly with NVIDIA's NVML library on your machine. No external services or credentials needed.

What are the system requirements?

Requires an NVIDIA GPU with drivers installed, Go 1.23+, CGO, and NVIDIA NVML headers. Prebuilt Docker images (linux/amd64, linux/arm64) are available on GHCR.

README (reference)

Source of truth, from the repository.

gpu-mcp-server

CI Helm Go Report Card Go Reference License DOI OpenSSF Scorecard OpenSSF Best Practices

Note: the OpenSSF Best Practices questionnaire is in progress. Once the project entry is registered at https://www.bestpractices.dev/en, swap the static badge above for the live one: [![OpenSSF Best Practices](https://www.bestpractices.dev/projects/<ID>/badge)](https://www.bestpractices.dev/projects/<ID>)

An MCP server that exposes NVIDIA GPU metrics as tools. Any MCP-compatible AI agent (Claude, Goose, Cursor, etc.) can query real-time GPU utilization, memory, temperature, power, PCIe and NVLink throughput no Prometheus or dcgm-exporter required.

Built on the official Go MCP SDK and NVIDIA go-nvml.

Tools

ToolDescription
list_gpusList all GPUs with utilization and memory info
get_gpu_metricsDetailed metrics for a GPU by index or UUID
get_gpu_processesPID-level GPU process attribution
gpu_summaryAggregate stats across all devices

All tools support MIG (Multi-Instance GPU) - MIG instances appear as separate devices with their parent GPU's shared metrics (temperature, power, PCIe).

Sample output

Each tool returns structured JSON. The examples below show the shape of the data an agent receives from a node with two NVIDIA A100 GPUs.

list_gpus:

{
  "count": 2,
  "devices": [
    {
      "index": 0,
      "uuid": "GPU-aaaa-1111",
      "name": "NVIDIA A100-SXM4-80GB",
      "gpu_utilization_percent": 85,
      "memory_used_mib": 57344,
      "memory_total_mib": 81920
    },
    {
      "index": 1,
      "uuid": "GPU-bbbb-2222",
      "name": "NVIDIA A100-SXM4-80GB",
      "gpu_utilization_percent": 20,
      "memory_used_mib": 12288,
      "memory_total_mib": 81920
    }
  ]
}

get_gpu_metrics (with {"index": 0} or {"uuid": "GPU-aaaa-1111"}):

{
  "index": 0,
  "uuid": "GPU-aaaa-1111",
  "name": "NVIDIA A100-SXM4-80GB",
  "gpu_utilization_percent": 85,
  "memory_utilization_percent": 70,
  "memory_used_mib": 57344,
  "memory_total_mib": 81920,
  "temperature_celsius": 72,
  "power_draw_watts": 300,
  "power_limit_watts": 400,
  "pcie_tx_kbps": 0,
  "pcie_rx_kbps": 0,
  "nvlink_tx_mbps": 0,
  "nvlink_rx_mbps": 0
}

gpu_summary:

{
  "device_count": 2,
  "avg_gpu_utilization": 52.5,
  "avg_memory_utilization": 42.5,
  "total_memory_used_mib": 69632,
  "total_memory_total_mib": 163840,
  "max_temperature_celsius": 72,
  "total_power_draw_watts": 375
}

MIG instances add is_mig, parent_gpu, and mig_profile fields to the get_gpu_metrics and list_gpus payloads.

Quick start

# build (requires CGO + NVML headers on Linux)
make build

# run the server communicates over stdio
./gpu-mcp-server

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "gpu": {
      "command": "/path/to/gpu-mcp-server"
    }
  }
}

Goose

extensions:
  gpu-metrics:
    type: stdio
    cmd: /path/to/gpu-mcp-server

Cursor

Add to .cursor/mcp.json for a project, or ~/.cursor/mcp.json for all projects:

{
  "mcpServers": {
    "gpu": {
      "type": "stdio",
      "command": "/path/to/gpu-mcp-server"
    }
  }
}

Windsurf

Add to ~/.codeium/windsurf/mcp_config.json:

{
  "mcpServers": {
    "gpu": {
      "command": "/path/to/gpu-mcp-server"
    }
  }
}

Build

Requires Go 1.23+, CGO, and NVIDIA drivers on the target machine.

make build       # compile binary
make test        # run tests (no GPU needed uses mock)
make lint        # golangci-lint
make docker      # container image

Tests use a mock collector, so they run anywhere no GPU hardware required.

Docker

Prebuilt multi-arch images (linux/amd64, linux/arm64) are published to GHCR on every release.

docker pull ghcr.io/pmady/gpu-mcp-server:latest
docker run --rm -i --gpus all ghcr.io/pmady/gpu-mcp-server:latest

The host needs the NVIDIA Container Toolkit installed for --gpus all to work. The server speaks MCP over stdio, so the -i flag is required — don't drop it.

{
  "mcpServers": {
    "gpu": {
      "command": "docker",
      "args": ["run", "--rm", "-i", "--gpus", "all", "ghcr.io/pmady/gpu-mcp-server:latest"]
    }
  }
}

Pin a specific version via tag instead of :latest, e.g. ghcr.io/pmady/gpu-mcp-server:v0.1.0.

Architecture

Agent (Claude/Goose) ─── MCP (stdio) ──→ gpu-mcp-server ──→ NVML ──→ GPU
                                              │
                                         Tools:
                                         • list_gpus
                                         • get_gpu_metrics
                                         • gpu_summary

The server runs as a local process alongside the agent. It calls NVML directly through cgo — no sidecar, no network hops, no metric pipeline to configure.

Project info

Citing

If you use gpu-mcp-server in your research or writing, please cite it via its DOI. Full citation metadata is in CITATION.cff.

@software{madduri_gpu_mcp_server,
  author    = {Madduri, Pavan},
  title     = {gpu-mcp-server: NVIDIA GPU metrics for AI agents over the Model Context Protocol},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22866670},
  url       = {https://doi.org/10.5281/zenodo.22866670}
}

Roadmap

See ROADMAP.md for the 12-month public roadmap.

Contributing

See CONTRIBUTING.md for how to get involved.

Contributors

Thanks to all our contributors! Add yourself via PR.

Governance

This project follows Linux Foundation Minimum Viable Governance.

Documentation

Star History

<a href="https://www.star-history.com/?repos=pmady%2Fgpu-mcp-server&type=date&legend=top-left"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/chart?repos=pmady/gpu-mcp-server&type=date&theme=dark&legend=top-left" /> <source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/chart?repos=pmady/gpu-mcp-server&type=date&legend=top-left" /> <img alt="Star History Chart" src="https://api.star-history.com/chart?repos=pmady/gpu-mcp-server&type=date&legend=top-left" /> </picture> </a>

Related MCP servers

Clinical voice analysis MCP server — AVQI, DSI, jitter/shimmer, pronunciation assessment, and more.

0
TypeScript
MIT
View repository →

Code knowledge graph over MCP: token-budgeted context for AI coding agents.

0
Python
MIT
View repository →

Servidor MCP per surtdecasa.cat: agenda cultural de Catalunya, cartellera de cinema i poblacions.

0
TypeScript
View repository →

Manage Microsoft 365 using natural language with CLI for Microsoft 365 commands.

124
TypeScript
MIT
View repository →

Shared memory and coordination layer for multi-agent AI systems across Claude, Codex, Copilot, and VS Code.

11
Python
MIT
View repository →

Blocker-aware decision layer for AI coding agents, grounded in source-linked, time-sensitive facts.

4
TypeScript
MIT
View repository →