grafana-dashboards
wshobson/agents
Create and manage production Grafana dashboards for real-time system and application metric visualization.
What is grafana-dashboards?
This skill enables you to design and deploy production-ready Grafana dashboards for monitoring applications, infrastructure, and business metrics. Use it when you need to visualize Prometheus metrics, create custom monitoring interfaces, implement SLO dashboards, or track operational KPIs.
- Design dashboards using RED (Rate, Errors, Duration) and USE (Utilization, Saturation, Errors) methodologies
- Create panel types including stat panels, time series graphs, tables, and heatmaps with Prometheus queries
- Implement dashboard variables for dynamic filtering by namespace, service, and other dimensions
- Configure alerts within dashboards with conditions, thresholds, and notification routing
- Provision dashboards as code using Terraform, Ansible, or Grafana's file-based provisioning
- Apply best practices for information hierarchy, naming conventions, and threshold configuration
How to install grafana-dashboards
npx skills add https://github.com/wshobson/agents --skill grafana-dashboards- Grafana instance (v7.0+) with Prometheus data source configured
- Prometheus metrics being scraped from your applications and infrastructure
- Access to Grafana API or file system for dashboard provisioning
How to use grafana-dashboards
- 1.Define your monitoring goals using RED method (for services) or USE method (for resources)
- 2.Design dashboard layout with critical metrics at top, trends in middle, detailed metrics below
- 3.Create panels with appropriate types (stat, graph, table, heatmap) and Prometheus queries
- 4.Add dashboard variables for dynamic filtering and multi-select capabilities
- 5.Configure alert conditions on critical panels with thresholds and notification channels
- 6.Test dashboard with different time ranges and variable combinations
- 7.Provision dashboard using JSON export, Terraform, Ansible, or Grafana provisioning files
Use cases
- Build API monitoring dashboards tracking request rate, error rate, and P95 latency
- Create infrastructure dashboards displaying CPU, memory, disk I/O, and network metrics per node
- Design database dashboards monitoring queries per second, connection pools, replication lag, and slow queries
- Implement application dashboards showing request rates, error rates, response time percentiles, and cache hit rates
- Set up SLO dashboards with alert conditions for error rate thresholds and latency targets
- DevOps engineers building operational observability platforms
- SRE teams implementing monitoring and alerting strategies
- Backend engineers creating application performance dashboards
- Infrastructure teams monitoring Kubernetes clusters and resource utilization
- Platform teams provisioning dashboards across multiple environments
grafana-dashboards FAQ
RED (Rate, Errors, Duration) is for monitoring services and APIs. USE (Utilization, Saturation, Errors) is for monitoring resources like CPU, memory, and disk. Use RED for application dashboards and USE for infrastructure dashboards.
Use dashboard variables for dynamic values like namespace, service, or instance. Reference variables in queries with $variable syntax. Export the dashboard JSON and provision it across environments using Terraform, Ansible, or Grafana's file provisioning.
Yes, you can configure alert conditions on individual panels with evaluators, thresholds, and notification channels. Alerts evaluate at specified frequencies and send notifications when conditions are met.
Use stat panels for single values, time series graphs for trends over time, tables for multi-row data, and heatmaps for distribution analysis. Choose based on whether you're showing a point-in-time value, trend, comparison, or distribution.
Use folders to group related dashboards, consistent naming conventions, panel descriptions for context, and templating for flexibility. Provision dashboards as code to maintain consistency and enable version control.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: grafana-dashboards description: Create and manage production Grafana dashboards for real-time visualization of system and application metrics. Use when building monitoring dashboards, visualizing metrics, or creating operational observability interfaces.
Grafana Dashboards
Create and manage production-ready Grafana dashboards for comprehensive system observability.
Purpose
Design effective Grafana dashboards for monitoring applications, infrastructure, and business metrics.
When to Use
- Visualize Prometheus metrics
- Create custom dashboards
- Implement SLO dashboards
- Monitor infrastructure
- Track business KPIs
Dashboard Design Principles
1. Hierarchy of Information
┌─────────────────────────────────────┐
│ Critical Metrics (Big Numbers) │
├─────────────────────────────────────┤
│ Key Trends (Time Series) │
├─────────────────────────────────────┤
│ Detailed Metrics (Tables/Heatmaps) │
└─────────────────────────────────────┘
2. RED Method (Services)
- Rate - Requests per second
- Errors - Error rate
- Duration - Latency/response time
3. USE Method (Resources)
- Utilization - % time resource is busy
- Saturation - Queue length/wait time
- Errors - Error count
Dashboard Structure
API Monitoring Dashboard
{
"dashboard": {
"title": "API Monitoring",
"tags": ["api", "production"],
"timezone": "browser",
"refresh": "30s",
"panels": [
{
"title": "Request Rate",
"type": "graph",
"targets": [
{
"expr": "sum(rate(http_requests_total[5m])) by (service)",
"legendFormat": "{{service}}"
}
],
"gridPos": { "x": 0, "y": 0, "w": 12, "h": 8 }
},
{
"title": "Error Rate %",
"type": "graph",
"targets": [
{
"expr": "(sum(rate(http_requests_total{status=~\"5..\"}[5m])) / sum(rate(http_requests_total[5m]))) * 100",
"legendFormat": "Error Rate"
}
],
"alert": {
"conditions": [
{
"evaluator": { "params": [5], "type": "gt" },
"operator": { "type": "and" },
"query": { "params": ["A", "5m", "now"] },
"type": "query"
}
]
},
"gridPos": { "x": 12, "y": 0, "w": 12, "h": 8 }
},
{
"title": "P95 Latency",
"type": "graph",
"targets": [
{
"expr": "histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))",
"legendFormat": "{{service}}"
}
],
"gridPos": { "x": 0, "y": 8, "w": 24, "h": 8 }
}
]
}
}
Panel Types
1. Stat Panel (Single Value)
{
"type": "stat",
"title": "Total Requests",
"targets": [
{
"expr": "sum(http_requests_total)"
}
],
"options": {
"reduceOptions": {
"values": false,
"calcs": ["lastNotNull"]
},
"orientation": "auto",
"textMode": "auto",
"colorMode": "value"
},
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 80, "color": "yellow" },
{ "value": 90, "color": "red" }
]
}
}
}
}
2. Time Series Graph
{
"type": "graph",
"title": "CPU Usage",
"targets": [
{
"expr": "100 - (avg by (instance) (rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) * 100)"
}
],
"yaxes": [
{ "format": "percent", "max": 100, "min": 0 },
{ "format": "short" }
]
}
3. Table Panel
{
"type": "table",
"title": "Service Status",
"targets": [
{
"expr": "up",
"format": "table",
"instant": true
}
],
"transformations": [
{
"id": "organize",
"options": {
"excludeByName": { "Time": true },
"indexByName": {},
"renameByName": {
"instance": "Instance",
"job": "Service",
"Value": "Status"
}
}
}
]
}
4. Heatmap
{
"type": "heatmap",
"title": "Latency Heatmap",
"targets": [
{
"expr": "sum(rate(http_request_duration_seconds_bucket[5m])) by (le)",
"format": "heatmap"
}
],
"dataFormat": "tsbuckets",
"yAxis": {
"format": "s"
}
}
Variables
Query Variables
{
"templating": {
"list": [
{
"name": "namespace",
"type": "query",
"datasource": "Prometheus",
"query": "label_values(kube_pod_info, namespace)",
"refresh": 1,
"multi": false
},
{
"name": "service",
"type": "query",
"datasource": "Prometheus",
"query": "label_values(kube_service_info{namespace=\"$namespace\"}, service)",
"refresh": 1,
"multi": true
}
]
}
}
Use Variables in Queries
sum(rate(http_requests_total{namespace="$namespace", service=~"$service"}[5m]))
Alerts in Dashboards
{
"alert": {
"name": "High Error Rate",
"conditions": [
{
"evaluator": {
"params": [5],
"type": "gt"
},
"operator": { "type": "and" },
"query": {
"params": ["A", "5m", "now"]
},
"reducer": { "type": "avg" },
"type": "query"
}
],
"executionErrorState": "alerting",
"for": "5m",
"frequency": "1m",
"message": "Error rate is above 5%",
"noDataState": "no_data",
"notifications": [{ "uid": "slack-channel" }]
}
}
Dashboard Provisioning
dashboards.yml:
apiVersion: 1
providers:
- name: "default"
orgId: 1
folder: "General"
type: file
disableDeletion: false
updateIntervalSeconds: 10
allowUiUpdates: true
options:
path: /etc/grafana/dashboards
Common Dashboard Patterns
Infrastructure Dashboard
Key Panels:
- CPU utilization per node
- Memory usage per node
- Disk I/O
- Network traffic
- Pod count by namespace
- Node status
Database Dashboard
Key Panels:
- Queries per second
- Connection pool usage
- Query latency (P50, P95, P99)
- Active connections
- Database size
- Replication lag
- Slow queries
Application Dashboard
Key Panels:
- Request rate
- Error rate
- Response time (percentiles)
- Active users/sessions
- Cache hit rate
- Queue length
Best Practices
- Start with templates (Grafana community dashboards)
- Use consistent naming for panels and variables
- Group related metrics in rows
- Set appropriate time ranges (default: Last 6 hours)
- Use variables for flexibility
- Add panel descriptions for context
- Configure units correctly
- Set meaningful thresholds for colors
- Use consistent colors across dashboards
- Test with different time ranges
Dashboard as Code
Terraform Provisioning
resource "grafana_dashboard" "api_monitoring" {
config_json = file("${path.module}/dashboards/api-monitoring.json")
folder = grafana_folder.monitoring.id
}
resource "grafana_folder" "monitoring" {
title = "Production Monitoring"
}
Ansible Provisioning
- name: Deploy Grafana dashboards
copy:
src: "{{ item }}"
dest: /etc/grafana/dashboards/
with_fileglob:
- "dashboards/*.json"
notify: restart grafana
Related Skills
prometheus-configuration- For metric collectionslo-implementation- For SLO dashboards
Related skills
More from wshobson/agents and the wider catalog.

grpo-rlvr-training
Train reasoning models with GRPO and reinforcement learning from verifiable rewards when task success is algorithmically checkable.

hads
Human-AI Document Standard for technical docs readable by both humans and AI models.

helm-chart-scaffolding
Design, organize, and manage Helm charts for templating and packaging Kubernetes applications.

hermes-tweet
X/Twitter research and monitoring for Hermes Agent with approval-gated actions.

hybrid-cloud-networking
Configure secure, high-performance connectivity between on-premises and cloud platforms using VPN and dedicated connections.

hybrid-search-implementation
Combine vector and keyword search for better retrieval in RAG and search systems.