PluginBench
Skill
Pass
Audit score 90

grafana-dashboards

wshobson/agents

Create and manage production Grafana dashboards for real-time system and application metric visualization.

What is grafana-dashboards?

This skill enables you to design and deploy production-ready Grafana dashboards for monitoring applications, infrastructure, and business metrics. Use it when you need to visualize Prometheus metrics, create custom monitoring interfaces, implement SLO dashboards, or track operational KPIs.

  • Design dashboards using RED (Rate, Errors, Duration) and USE (Utilization, Saturation, Errors) methodologies
  • Create panel types including stat panels, time series graphs, tables, and heatmaps with Prometheus queries
  • Implement dashboard variables for dynamic filtering by namespace, service, and other dimensions
  • Configure alerts within dashboards with conditions, thresholds, and notification routing
  • Provision dashboards as code using Terraform, Ansible, or Grafana's file-based provisioning
  • Apply best practices for information hierarchy, naming conventions, and threshold configuration

How to install grafana-dashboards

npx skills add https://github.com/wshobson/agents --skill grafana-dashboards
Prerequisites
  • Grafana instance (v7.0+) with Prometheus data source configured
  • Prometheus metrics being scraped from your applications and infrastructure
  • Access to Grafana API or file system for dashboard provisioning
Claude Code
Cursor
Windsurf
Cline

How to use grafana-dashboards

  1. 1.Define your monitoring goals using RED method (for services) or USE method (for resources)
  2. 2.Design dashboard layout with critical metrics at top, trends in middle, detailed metrics below
  3. 3.Create panels with appropriate types (stat, graph, table, heatmap) and Prometheus queries
  4. 4.Add dashboard variables for dynamic filtering and multi-select capabilities
  5. 5.Configure alert conditions on critical panels with thresholds and notification channels
  6. 6.Test dashboard with different time ranges and variable combinations
  7. 7.Provision dashboard using JSON export, Terraform, Ansible, or Grafana provisioning files

Use cases

Good for
  • Build API monitoring dashboards tracking request rate, error rate, and P95 latency
  • Create infrastructure dashboards displaying CPU, memory, disk I/O, and network metrics per node
  • Design database dashboards monitoring queries per second, connection pools, replication lag, and slow queries
  • Implement application dashboards showing request rates, error rates, response time percentiles, and cache hit rates
  • Set up SLO dashboards with alert conditions for error rate thresholds and latency targets
Who it's for
  • DevOps engineers building operational observability platforms
  • SRE teams implementing monitoring and alerting strategies
  • Backend engineers creating application performance dashboards
  • Infrastructure teams monitoring Kubernetes clusters and resource utilization
  • Platform teams provisioning dashboards across multiple environments

grafana-dashboards FAQ

What's the difference between RED and USE methods?

RED (Rate, Errors, Duration) is for monitoring services and APIs. USE (Utilization, Saturation, Errors) is for monitoring resources like CPU, memory, and disk. Use RED for application dashboards and USE for infrastructure dashboards.

How do I make dashboards reusable across environments?

Use dashboard variables for dynamic values like namespace, service, or instance. Reference variables in queries with $variable syntax. Export the dashboard JSON and provision it across environments using Terraform, Ansible, or Grafana's file provisioning.

Can I set up alerts directly in dashboards?

Yes, you can configure alert conditions on individual panels with evaluators, thresholds, and notification channels. Alerts evaluate at specified frequencies and send notifications when conditions are met.

What panel types work best for different metrics?

Use stat panels for single values, time series graphs for trends over time, tables for multi-row data, and heatmaps for distribution analysis. Choose based on whether you're showing a point-in-time value, trend, comparison, or distribution.

How do I organize dashboards for large teams?

Use folders to group related dashboards, consistent naming conventions, panel descriptions for context, and templating for flexibility. Provision dashboards as code to maintain consistency and enable version control.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: grafana-dashboards description: Create and manage production Grafana dashboards for real-time visualization of system and application metrics. Use when building monitoring dashboards, visualizing metrics, or creating operational observability interfaces.

Grafana Dashboards

Create and manage production-ready Grafana dashboards for comprehensive system observability.

Purpose

Design effective Grafana dashboards for monitoring applications, infrastructure, and business metrics.

When to Use

  • Visualize Prometheus metrics
  • Create custom dashboards
  • Implement SLO dashboards
  • Monitor infrastructure
  • Track business KPIs

Dashboard Design Principles

1. Hierarchy of Information

┌─────────────────────────────────────┐
│  Critical Metrics (Big Numbers)     │
├─────────────────────────────────────┤
│  Key Trends (Time Series)           │
├─────────────────────────────────────┤
│  Detailed Metrics (Tables/Heatmaps) │
└─────────────────────────────────────┘

2. RED Method (Services)

  • Rate - Requests per second
  • Errors - Error rate
  • Duration - Latency/response time

3. USE Method (Resources)

  • Utilization - % time resource is busy
  • Saturation - Queue length/wait time
  • Errors - Error count

Dashboard Structure

API Monitoring Dashboard

{
  "dashboard": {
    "title": "API Monitoring",
    "tags": ["api", "production"],
    "timezone": "browser",
    "refresh": "30s",
    "panels": [
      {
        "title": "Request Rate",
        "type": "graph",
        "targets": [
          {
            "expr": "sum(rate(http_requests_total[5m])) by (service)",
            "legendFormat": "{{service}}"
          }
        ],
        "gridPos": { "x": 0, "y": 0, "w": 12, "h": 8 }
      },
      {
        "title": "Error Rate %",
        "type": "graph",
        "targets": [
          {
            "expr": "(sum(rate(http_requests_total{status=~\"5..\"}[5m])) / sum(rate(http_requests_total[5m]))) * 100",
            "legendFormat": "Error Rate"
          }
        ],
        "alert": {
          "conditions": [
            {
              "evaluator": { "params": [5], "type": "gt" },
              "operator": { "type": "and" },
              "query": { "params": ["A", "5m", "now"] },
              "type": "query"
            }
          ]
        },
        "gridPos": { "x": 12, "y": 0, "w": 12, "h": 8 }
      },
      {
        "title": "P95 Latency",
        "type": "graph",
        "targets": [
          {
            "expr": "histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))",
            "legendFormat": "{{service}}"
          }
        ],
        "gridPos": { "x": 0, "y": 8, "w": 24, "h": 8 }
      }
    ]
  }
}

Panel Types

1. Stat Panel (Single Value)

{
  "type": "stat",
  "title": "Total Requests",
  "targets": [
    {
      "expr": "sum(http_requests_total)"
    }
  ],
  "options": {
    "reduceOptions": {
      "values": false,
      "calcs": ["lastNotNull"]
    },
    "orientation": "auto",
    "textMode": "auto",
    "colorMode": "value"
  },
  "fieldConfig": {
    "defaults": {
      "thresholds": {
        "mode": "absolute",
        "steps": [
          { "value": 0, "color": "green" },
          { "value": 80, "color": "yellow" },
          { "value": 90, "color": "red" }
        ]
      }
    }
  }
}

2. Time Series Graph

{
  "type": "graph",
  "title": "CPU Usage",
  "targets": [
    {
      "expr": "100 - (avg by (instance) (rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) * 100)"
    }
  ],
  "yaxes": [
    { "format": "percent", "max": 100, "min": 0 },
    { "format": "short" }
  ]
}

3. Table Panel

{
  "type": "table",
  "title": "Service Status",
  "targets": [
    {
      "expr": "up",
      "format": "table",
      "instant": true
    }
  ],
  "transformations": [
    {
      "id": "organize",
      "options": {
        "excludeByName": { "Time": true },
        "indexByName": {},
        "renameByName": {
          "instance": "Instance",
          "job": "Service",
          "Value": "Status"
        }
      }
    }
  ]
}

4. Heatmap

{
  "type": "heatmap",
  "title": "Latency Heatmap",
  "targets": [
    {
      "expr": "sum(rate(http_request_duration_seconds_bucket[5m])) by (le)",
      "format": "heatmap"
    }
  ],
  "dataFormat": "tsbuckets",
  "yAxis": {
    "format": "s"
  }
}

Variables

Query Variables

{
  "templating": {
    "list": [
      {
        "name": "namespace",
        "type": "query",
        "datasource": "Prometheus",
        "query": "label_values(kube_pod_info, namespace)",
        "refresh": 1,
        "multi": false
      },
      {
        "name": "service",
        "type": "query",
        "datasource": "Prometheus",
        "query": "label_values(kube_service_info{namespace=\"$namespace\"}, service)",
        "refresh": 1,
        "multi": true
      }
    ]
  }
}

Use Variables in Queries

sum(rate(http_requests_total{namespace="$namespace", service=~"$service"}[5m]))

Alerts in Dashboards

{
  "alert": {
    "name": "High Error Rate",
    "conditions": [
      {
        "evaluator": {
          "params": [5],
          "type": "gt"
        },
        "operator": { "type": "and" },
        "query": {
          "params": ["A", "5m", "now"]
        },
        "reducer": { "type": "avg" },
        "type": "query"
      }
    ],
    "executionErrorState": "alerting",
    "for": "5m",
    "frequency": "1m",
    "message": "Error rate is above 5%",
    "noDataState": "no_data",
    "notifications": [{ "uid": "slack-channel" }]
  }
}

Dashboard Provisioning

dashboards.yml:

apiVersion: 1

providers:
  - name: "default"
    orgId: 1
    folder: "General"
    type: file
    disableDeletion: false
    updateIntervalSeconds: 10
    allowUiUpdates: true
    options:
      path: /etc/grafana/dashboards

Common Dashboard Patterns

Infrastructure Dashboard

Key Panels:

  • CPU utilization per node
  • Memory usage per node
  • Disk I/O
  • Network traffic
  • Pod count by namespace
  • Node status

Database Dashboard

Key Panels:

  • Queries per second
  • Connection pool usage
  • Query latency (P50, P95, P99)
  • Active connections
  • Database size
  • Replication lag
  • Slow queries

Application Dashboard

Key Panels:

  • Request rate
  • Error rate
  • Response time (percentiles)
  • Active users/sessions
  • Cache hit rate
  • Queue length

Best Practices

  1. Start with templates (Grafana community dashboards)
  2. Use consistent naming for panels and variables
  3. Group related metrics in rows
  4. Set appropriate time ranges (default: Last 6 hours)
  5. Use variables for flexibility
  6. Add panel descriptions for context
  7. Configure units correctly
  8. Set meaningful thresholds for colors
  9. Use consistent colors across dashboards
  10. Test with different time ranges

Dashboard as Code

Terraform Provisioning

resource "grafana_dashboard" "api_monitoring" {
  config_json = file("${path.module}/dashboards/api-monitoring.json")
  folder      = grafana_folder.monitoring.id
}

resource "grafana_folder" "monitoring" {
  title = "Production Monitoring"
}

Ansible Provisioning

- name: Deploy Grafana dashboards
  copy:
    src: "{{ item }}"
    dest: /etc/grafana/dashboards/
  with_fileglob:
    - "dashboards/*.json"
  notify: restart grafana

Related Skills

  • prometheus-configuration - For metric collection
  • slo-implementation - For SLO dashboards