PluginBench
Skill
Fail
Audit score 45

sre-engineer

jeffallan/claude-skills

Define SLOs, manage error budgets, and automate production reliability at scale.

What is sre-engineer?

SRE Engineer skill provides structured workflows for defining service level objectives (SLOs), calculating error budgets, designing incident response, and building monitoring and automation for production systems. Use it when establishing reliability targets, managing on-call operations, reducing toil, or implementing chaos engineering.

  • Define quantitative SLOs with SLI measurements and error budget calculations
  • Build golden signal monitoring (latency, traffic, errors, saturation) with Prometheus alerting rules
  • Create automation scripts to reduce toil and enable auto-remediation
  • Design incident response procedures and blameless postmortem processes
  • Develop capacity models and chaos engineering test scenarios
  • Balance reliability targets with feature velocity using error budget policies

How to install sre-engineer

npx skills add https://github.com/jeffallan/claude-skills --skill sre-engineer
Claude Code
Cursor
Windsurf
Cline

How to use sre-engineer

  1. 1.Assess current reliability posture by reviewing architecture, existing SLOs, incident history, and toil levels
  2. 2.Define meaningful SLIs and set quantitative SLO targets (e.g., 99.9% availability) with user impact justification
  3. 3.Verify SLO targets align with business expectations before implementation
  4. 4.Implement golden signal dashboards and multiwindow burn rate alerting rules
  5. 5.Identify repetitive operational tasks and build automation scripts to reduce toil
  6. 6.Design and execute chaos engineering experiments to validate recovery meets RTO/RPO targets

Use cases

Good for
  • Setting SLO targets for a microservice and calculating monthly error budgets
  • Implementing multiwindow burn rate alerts to detect fast and slow error budget consumption
  • Automating pod restarts when error rates exceed thresholds
  • Designing chaos experiments to verify RTO/RPO targets before incidents occur
  • Identifying and automating repetitive operational tasks to reduce on-call toil
Who it's for
  • Site reliability engineers managing production systems
  • DevOps engineers building monitoring and automation infrastructure
  • Platform teams defining reliability standards across services
  • On-call engineers designing incident response runbooks
  • Engineering leaders balancing reliability with deployment velocity

sre-engineer FAQ

What is an error budget and how do I calculate it?

An error budget is the allowed amount of downtime or errors within an SLO window. For a 99.9% availability SLO over 30 days, the error budget is (1 - 0.999) × 30 × 24 × 60 = 43.2 minutes. When budget burns faster than expected, trigger policies like freezing non-critical releases.

What are golden signals and why do I need them?

Golden signals are latency, traffic, errors, and saturation—the four metrics that best indicate system health. Monitor them via PromQL queries and dashboards to detect problems before users are impacted.

When should I use multiwindow burn rate alerts?

Multiwindow alerts detect both fast burns (2% budget in 1 hour, triggering immediately) and slow burns (5% budget in 6 hours, sustained degradation). This prevents alert fatigue while catching real problems early.

How do I automate toil reduction?

Identify repetitive manual tasks (e.g., restarting failed pods), measure their frequency, then build automation scripts (Python, Go, Terraform) to handle them. Track toil metrics and aim to keep operational work below 50% of on-call time.

What should a postmortem include?

Blameless postmortems document the incident timeline, root causes, impact, and action items. Focus on systems and processes, not individuals. Use postmortems to improve monitoring, runbooks, and automation.

Full instructions (SKILL.md)

Source of truth, from jeffallan/claude-skills.


name: sre-engineer description: Defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems. Use when defining SLIs/SLOs, managing error budgets, building reliable systems at scale, incident management, chaos engineering, toil reduction, or capacity planning. license: MIT metadata: author: https://github.com/Jeffallan version: "1.1.0" domain: devops triggers: SRE, site reliability, SLO, SLI, error budget, incident management, chaos engineering, toil reduction, on-call, MTTR role: specialist scope: implementation output-format: code related-skills: devops-engineer, cloud-architect, kubernetes-specialist

SRE Engineer

Core Workflow

  1. Assess reliability - Review architecture, SLOs, incidents, toil levels
  2. Define SLOs - Identify meaningful SLIs and set appropriate targets
  3. Verify alignment - Confirm SLO targets reflect user expectations before proceeding
  4. Implement monitoring - Build golden signal dashboards and alerting
  5. Automate toil - Identify repetitive tasks and build automation
  6. Test resilience - Design and execute chaos experiments; verify recovery meets RTO/RPO targets before marking the experiment complete; validate recovery behavior end-to-end

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
SLO/SLIreferences/slo-sli-management.mdDefining SLOs, calculating error budgets
Error Budgetsreferences/error-budget-policy.mdManaging budgets, burn rates, policies
Monitoringreferences/monitoring-alerting.mdGolden signals, alert design, dashboards
Automationreferences/automation-toil.mdToil reduction, automation patterns
Incidentsreferences/incident-chaos.mdIncident response, chaos engineering

Constraints

MUST DO

  • Define quantitative SLOs (e.g., 99.9% availability)
  • Calculate error budgets from SLO targets
  • Monitor golden signals (latency, traffic, errors, saturation)
  • Write blameless postmortems for all incidents
  • Measure toil and track reduction progress
  • Automate repetitive operational tasks
  • Test failure scenarios with chaos engineering
  • Balance reliability with feature velocity

MUST NOT DO

  • Set SLOs without user impact justification
  • Alert on symptoms without actionable runbooks
  • Tolerate >50% toil without automation plan
  • Skip postmortems or assign blame
  • Implement manual processes for recurring tasks
  • Deploy without capacity planning
  • Ignore error budget exhaustion
  • Build systems that can't degrade gracefully

Output Templates

When implementing SRE practices, provide:

  1. SLO definitions with SLI measurements and targets
  2. Monitoring/alerting configuration (Prometheus, etc.)
  3. Automation scripts (Python, Go, Terraform)
  4. Runbooks with clear remediation steps
  5. Brief explanation of reliability impact

Concrete Examples

SLO Definition & Error Budget Calculation

# 99.9% availability SLO over a 30-day window
# Allowed downtime: (1 - 0.999) * 30 * 24 * 60 = 43.2 minutes/month
# Error budget (request-based): 0.001 * total_requests

# Example: 10M requests/month → 10,000 error budget requests
# If 5,000 errors consumed in week 1 → 50% budget burned in 25% of window
# → Trigger error budget policy: freeze non-critical releases

Prometheus SLO Alerting Rule (Multiwindow Burn Rate)

groups:
  - name: slo_availability
    rules:
      # Fast burn: 2% budget in 1h (14.4x burn rate)
      - alert: HighErrorBudgetBurn
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[1h]))
            /
            sum(rate(http_requests_total[1h]))
          ) > 0.014400
          and
          (
            sum(rate(http_requests_total{status=~"5.."}[5m]))
            /
            sum(rate(http_requests_total[5m]))
          ) > 0.014400
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "High error budget burn rate detected"
          runbook: "https://wiki.internal/runbooks/high-error-burn"

      # Slow burn: 5% budget in 6h (1x burn rate sustained)
      - alert: SlowErrorBudgetBurn
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[6h]))
            /
            sum(rate(http_requests_total[6h]))
          ) > 0.001
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "Sustained error budget consumption"
          runbook: "https://wiki.internal/runbooks/slow-error-burn"

PromQL Golden Signal Queries

# Latency — 99th percentile request duration
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

# Traffic — requests per second by service
sum(rate(http_requests_total[5m])) by (service)

# Errors — error rate ratio
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
  /
sum(rate(http_requests_total[5m])) by (service)

# Saturation — CPU throttling ratio
sum(rate(container_cpu_cfs_throttled_seconds_total[5m])) by (pod)
  /
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod)

Toil Automation Script (Python)

#!/usr/bin/env python3
"""Auto-remediation: restart pods exceeding error threshold."""
import subprocess, sys, json

ERROR_THRESHOLD = 0.05  # 5% error rate triggers restart

def get_error_rate(service: str) -> float:
    """Query Prometheus for current error rate."""
    import urllib.request
    query = f'sum(rate(http_requests_total{{status=~"5..",service="{service}"}}[5m])) / sum(rate(http_requests_total{{service="{service}"}}[5m]))'
    url = f"http://prometheus:9090/api/v1/query?query={urllib.request.quote(query)}"
    with urllib.request.urlopen(url) as resp:
        data = json.load(resp)
    results = data["data"]["result"]
    return float(results[0]["value"][1]) if results else 0.0

def restart_deployment(namespace: str, deployment: str) -> None:
    subprocess.run(
        ["kubectl", "rollout", "restart", f"deployment/{deployment}", "-n", namespace],
        check=True
    )
    print(f"Restarted {namespace}/{deployment}")

if __name__ == "__main__":
    service, namespace, deployment = sys.argv[1], sys.argv[2], sys.argv[3]
    rate = get_error_rate(service)
    print(f"Error rate for {service}: {rate:.2%}")
    if rate > ERROR_THRESHOLD:
        restart_deployment(namespace, deployment)
    else:
        print("Within SLO threshold — no action required")

Documentation