sre-engineer
jeffallan/claude-skills
Define SLOs, manage error budgets, and automate production reliability at scale.
What is sre-engineer?
SRE Engineer skill provides structured workflows for defining service level objectives (SLOs), calculating error budgets, designing incident response, and building monitoring and automation for production systems. Use it when establishing reliability targets, managing on-call operations, reducing toil, or implementing chaos engineering.
- Define quantitative SLOs with SLI measurements and error budget calculations
- Build golden signal monitoring (latency, traffic, errors, saturation) with Prometheus alerting rules
- Create automation scripts to reduce toil and enable auto-remediation
- Design incident response procedures and blameless postmortem processes
- Develop capacity models and chaos engineering test scenarios
- Balance reliability targets with feature velocity using error budget policies
How to install sre-engineer
npx skills add https://github.com/jeffallan/claude-skills --skill sre-engineerHow to use sre-engineer
- 1.Assess current reliability posture by reviewing architecture, existing SLOs, incident history, and toil levels
- 2.Define meaningful SLIs and set quantitative SLO targets (e.g., 99.9% availability) with user impact justification
- 3.Verify SLO targets align with business expectations before implementation
- 4.Implement golden signal dashboards and multiwindow burn rate alerting rules
- 5.Identify repetitive operational tasks and build automation scripts to reduce toil
- 6.Design and execute chaos engineering experiments to validate recovery meets RTO/RPO targets
Use cases
- Setting SLO targets for a microservice and calculating monthly error budgets
- Implementing multiwindow burn rate alerts to detect fast and slow error budget consumption
- Automating pod restarts when error rates exceed thresholds
- Designing chaos experiments to verify RTO/RPO targets before incidents occur
- Identifying and automating repetitive operational tasks to reduce on-call toil
- Site reliability engineers managing production systems
- DevOps engineers building monitoring and automation infrastructure
- Platform teams defining reliability standards across services
- On-call engineers designing incident response runbooks
- Engineering leaders balancing reliability with deployment velocity
sre-engineer FAQ
An error budget is the allowed amount of downtime or errors within an SLO window. For a 99.9% availability SLO over 30 days, the error budget is (1 - 0.999) × 30 × 24 × 60 = 43.2 minutes. When budget burns faster than expected, trigger policies like freezing non-critical releases.
Golden signals are latency, traffic, errors, and saturation—the four metrics that best indicate system health. Monitor them via PromQL queries and dashboards to detect problems before users are impacted.
Multiwindow alerts detect both fast burns (2% budget in 1 hour, triggering immediately) and slow burns (5% budget in 6 hours, sustained degradation). This prevents alert fatigue while catching real problems early.
Identify repetitive manual tasks (e.g., restarting failed pods), measure their frequency, then build automation scripts (Python, Go, Terraform) to handle them. Track toil metrics and aim to keep operational work below 50% of on-call time.
Blameless postmortems document the incident timeline, root causes, impact, and action items. Focus on systems and processes, not individuals. Use postmortems to improve monitoring, runbooks, and automation.
Full instructions (SKILL.md)
Source of truth, from jeffallan/claude-skills.
name: sre-engineer description: Defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems. Use when defining SLIs/SLOs, managing error budgets, building reliable systems at scale, incident management, chaos engineering, toil reduction, or capacity planning. license: MIT metadata: author: https://github.com/Jeffallan version: "1.1.0" domain: devops triggers: SRE, site reliability, SLO, SLI, error budget, incident management, chaos engineering, toil reduction, on-call, MTTR role: specialist scope: implementation output-format: code related-skills: devops-engineer, cloud-architect, kubernetes-specialist
SRE Engineer
Core Workflow
- Assess reliability - Review architecture, SLOs, incidents, toil levels
- Define SLOs - Identify meaningful SLIs and set appropriate targets
- Verify alignment - Confirm SLO targets reflect user expectations before proceeding
- Implement monitoring - Build golden signal dashboards and alerting
- Automate toil - Identify repetitive tasks and build automation
- Test resilience - Design and execute chaos experiments; verify recovery meets RTO/RPO targets before marking the experiment complete; validate recovery behavior end-to-end
Reference Guide
Load detailed guidance based on context:
| Topic | Reference | Load When |
|---|---|---|
| SLO/SLI | references/slo-sli-management.md | Defining SLOs, calculating error budgets |
| Error Budgets | references/error-budget-policy.md | Managing budgets, burn rates, policies |
| Monitoring | references/monitoring-alerting.md | Golden signals, alert design, dashboards |
| Automation | references/automation-toil.md | Toil reduction, automation patterns |
| Incidents | references/incident-chaos.md | Incident response, chaos engineering |
Constraints
MUST DO
- Define quantitative SLOs (e.g., 99.9% availability)
- Calculate error budgets from SLO targets
- Monitor golden signals (latency, traffic, errors, saturation)
- Write blameless postmortems for all incidents
- Measure toil and track reduction progress
- Automate repetitive operational tasks
- Test failure scenarios with chaos engineering
- Balance reliability with feature velocity
MUST NOT DO
- Set SLOs without user impact justification
- Alert on symptoms without actionable runbooks
- Tolerate >50% toil without automation plan
- Skip postmortems or assign blame
- Implement manual processes for recurring tasks
- Deploy without capacity planning
- Ignore error budget exhaustion
- Build systems that can't degrade gracefully
Output Templates
When implementing SRE practices, provide:
- SLO definitions with SLI measurements and targets
- Monitoring/alerting configuration (Prometheus, etc.)
- Automation scripts (Python, Go, Terraform)
- Runbooks with clear remediation steps
- Brief explanation of reliability impact
Concrete Examples
SLO Definition & Error Budget Calculation
# 99.9% availability SLO over a 30-day window
# Allowed downtime: (1 - 0.999) * 30 * 24 * 60 = 43.2 minutes/month
# Error budget (request-based): 0.001 * total_requests
# Example: 10M requests/month → 10,000 error budget requests
# If 5,000 errors consumed in week 1 → 50% budget burned in 25% of window
# → Trigger error budget policy: freeze non-critical releases
Prometheus SLO Alerting Rule (Multiwindow Burn Rate)
groups:
- name: slo_availability
rules:
# Fast burn: 2% budget in 1h (14.4x burn rate)
- alert: HighErrorBudgetBurn
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
) > 0.014400
and
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
) > 0.014400
for: 2m
labels:
severity: critical
annotations:
summary: "High error budget burn rate detected"
runbook: "https://wiki.internal/runbooks/high-error-burn"
# Slow burn: 5% budget in 6h (1x burn rate sustained)
- alert: SlowErrorBudgetBurn
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[6h]))
/
sum(rate(http_requests_total[6h]))
) > 0.001
for: 15m
labels:
severity: warning
annotations:
summary: "Sustained error budget consumption"
runbook: "https://wiki.internal/runbooks/slow-error-burn"
PromQL Golden Signal Queries
# Latency — 99th percentile request duration
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))
# Traffic — requests per second by service
sum(rate(http_requests_total[5m])) by (service)
# Errors — error rate ratio
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/
sum(rate(http_requests_total[5m])) by (service)
# Saturation — CPU throttling ratio
sum(rate(container_cpu_cfs_throttled_seconds_total[5m])) by (pod)
/
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod)
Toil Automation Script (Python)
#!/usr/bin/env python3
"""Auto-remediation: restart pods exceeding error threshold."""
import subprocess, sys, json
ERROR_THRESHOLD = 0.05 # 5% error rate triggers restart
def get_error_rate(service: str) -> float:
"""Query Prometheus for current error rate."""
import urllib.request
query = f'sum(rate(http_requests_total{{status=~"5..",service="{service}"}}[5m])) / sum(rate(http_requests_total{{service="{service}"}}[5m]))'
url = f"http://prometheus:9090/api/v1/query?query={urllib.request.quote(query)}"
with urllib.request.urlopen(url) as resp:
data = json.load(resp)
results = data["data"]["result"]
return float(results[0]["value"][1]) if results else 0.0
def restart_deployment(namespace: str, deployment: str) -> None:
subprocess.run(
["kubectl", "rollout", "restart", f"deployment/{deployment}", "-n", namespace],
check=True
)
print(f"Restarted {namespace}/{deployment}")
if __name__ == "__main__":
service, namespace, deployment = sys.argv[1], sys.argv[2], sys.argv[3]
rate = get_error_rate(service)
print(f"Error rate for {service}: {rate:.2%}")
if rate > ERROR_THRESHOLD:
restart_deployment(namespace, deployment)
else:
print("Within SLO threshold — no action required")
Related skills
More from jeffallan/claude-skills and the wider catalog.

swift-expert
Expert Swift development for iOS/macOS with SwiftUI, async/await, and protocol-oriented architecture.

terraform-engineer
Senior Terraform engineer for infrastructure as code across AWS, Azure, and GCP with modular design and state management.

test-master
Comprehensive testing specialist for unit, integration, E2E, performance, and security tests.

the-fool
Play devil's advocate with structured critical reasoning to stress-test ideas, plans, and decisions.

typescript-pro
Advanced TypeScript type systems, branded types, and tRPC end-to-end type safety for complex applications.

vue-expert
Vue 3 Composition API specialist for components, Nuxt SSR/SSG, Pinia state, and mobile apps.