slo-implementation
wshobson/agents
Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for measurable reliability targets.
What is slo-implementation?
Framework for establishing and tracking service reliability using SLIs, SLOs, and error budgets. Use this when defining reliability targets, implementing SRE practices, measuring service performance, or creating SLO-based alerting strategies.
- Define SLIs (availability, latency, durability) with Prometheus queries
- Set SLO targets and calculate error budgets from reliability goals
- Implement Prometheus recording rules to track SLI and SLO compliance
- Create multi-window burn-rate alerts (fast and slow burn detection)
- Build Grafana dashboards to visualize SLO compliance and error budget status
- Track error budget consumption and forecast budget exhaustion
How to install slo-implementation
npx skills add https://github.com/wshobson/agents --skill slo-implementation- Prometheus or compatible metrics system collecting service metrics
- Grafana or similar dashboard tool for visualization
- HTTP request metrics (status codes, latency) or equivalent service metrics
How to use slo-implementation
- 1.Define your SLIs by selecting metric types (availability, latency, durability) and writing Prometheus queries
- 2.Set SLO targets based on user expectations, business requirements, and current performance
- 3.Calculate error budgets using the formula: Error Budget = 1 - SLO Target
- 4.Create Prometheus recording rules for SLI calculations and SLO compliance tracking
- 5.Implement multi-window burn-rate alert rules (fast burn at 14.4x, slow burn at 6x)
- 6.Build a Grafana dashboard displaying current SLO compliance, error budget remaining, and burn-rate trends
- 7.Define an error budget policy to trigger actions (normal velocity, postpone risky changes, feature freeze) based on remaining budget
Use cases
- Establish 99.9% availability SLO for an API with automated burn-rate alerting
- Implement error budget policy to balance feature velocity with reliability
- Create latency SLOs (e.g., p95 < 500ms) and track compliance over 28-day windows
- Monitor error budget burn rates to trigger feature freezes when budget is low
- Dashboard showing current SLO status, remaining error budget, and trend analysis
- Site Reliability Engineers (SREs) implementing SRE practices
- Platform and infrastructure teams defining service reliability targets
- DevOps engineers setting up monitoring and alerting for services
- Engineering leaders balancing innovation velocity with reliability requirements
slo-implementation FAQ
SLA is a contract with customers; SLO is an internal reliability target; SLI is the actual measurement. SLIs feed into SLOs, which are bounded by SLAs.
Consider user expectations, business requirements, current performance, cost of reliability, and competitor benchmarks. Common targets range from 99% to 99.99% depending on service criticality.
Error budget is the allowable downtime: 1 - SLO Target. For 99.9% SLO, you have 43.2 minutes/month. Use it to decide when to freeze features (0% remaining) or proceed normally (100% remaining).
Burn-rate alerts detect both fast outages (14.4x rate over 1 hour) and slow degradation (6x rate over 6 hours), triggering appropriate responses before error budget is exhausted.
Review quarterly or when business requirements change. Adjust targets based on actual performance, user feedback, and cost-benefit analysis of increased reliability.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: slo-implementation description: Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. Use when establishing reliability targets, implementing SRE practices, or measuring service performance.
SLO Implementation
Framework for defining and implementing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
Purpose
Implement measurable reliability targets using SLIs, SLOs, and error budgets to balance reliability with innovation velocity.
When to Use
- Define service reliability targets
- Measure user-perceived reliability
- Implement error budgets
- Create SLO-based alerts
- Track reliability goals
SLI/SLO/SLA Hierarchy
SLA (Service Level Agreement)
↓ Contract with customers
SLO (Service Level Objective)
↓ Internal reliability target
SLI (Service Level Indicator)
↓ Actual measurement
Defining SLIs
Common SLI Types
1. Availability SLI
# Successful requests / Total requests
sum(rate(http_requests_total{status!~"5.."}[28d]))
/
sum(rate(http_requests_total[28d]))
2. Latency SLI
# Requests below latency threshold / Total requests
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
/
sum(rate(http_request_duration_seconds_count[28d]))
3. Durability SLI
# Successful writes / Total writes
sum(storage_writes_successful_total)
/
sum(storage_writes_total)
Setting SLO Targets
Availability SLO Examples
| SLO % | Downtime/Month | Downtime/Year |
|---|---|---|
| 99% | 7.2 hours | 3.65 days |
| 99.9% | 43.2 minutes | 8.76 hours |
| 99.95% | 21.6 minutes | 4.38 hours |
| 99.99% | 4.32 minutes | 52.56 minutes |
Choose Appropriate SLOs
Consider:
- User expectations
- Business requirements
- Current performance
- Cost of reliability
- Competitor benchmarks
Example SLOs:
slos:
- name: api_availability
target: 99.9
window: 28d
sli: |
sum(rate(http_requests_total{status!~"5.."}[28d]))
/
sum(rate(http_requests_total[28d]))
- name: api_latency_p95
target: 99
window: 28d
sli: |
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
/
sum(rate(http_request_duration_seconds_count[28d]))
Error Budget Calculation
Error Budget Formula
Error Budget = 1 - SLO Target
Example:
- SLO: 99.9% availability
- Error Budget: 0.1% = 43.2 minutes/month
- Current Error: 0.05% = 21.6 minutes/month
- Remaining Budget: 50%
Error Budget Policy
error_budget_policy:
- remaining_budget: 100%
action: Normal development velocity
- remaining_budget: 50%
action: Consider postponing risky changes
- remaining_budget: 10%
action: Freeze non-critical changes
- remaining_budget: 0%
action: Feature freeze, focus on reliability
SLO Implementation
Prometheus Recording Rules
# SLI Recording Rules
groups:
- name: sli_rules
interval: 30s
rules:
# Availability SLI
- record: sli:http_availability:ratio
expr: |
sum(rate(http_requests_total{status!~"5.."}[28d]))
/
sum(rate(http_requests_total[28d]))
# Latency SLI (requests < 500ms)
- record: sli:http_latency:ratio
expr: |
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
/
sum(rate(http_request_duration_seconds_count[28d]))
- name: slo_rules
interval: 5m
rules:
# SLO compliance (1 = meeting SLO, 0 = violating)
- record: slo:http_availability:compliance
expr: sli:http_availability:ratio >= bool 0.999
- record: slo:http_latency:compliance
expr: sli:http_latency:ratio >= bool 0.99
# Error budget remaining (percentage)
- record: slo:http_availability:error_budget_remaining
expr: |
(sli:http_availability:ratio - 0.999) / (1 - 0.999) * 100
# Error budget burn rate
- record: slo:http_availability:burn_rate_5m
expr: |
(1 - (
sum(rate(http_requests_total{status!~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
)) / (1 - 0.999)
SLO Alerting Rules
groups:
- name: slo_alerts
interval: 1m
rules:
# Fast burn: 14.4x rate, 1 hour window
# Consumes 2% error budget in 1 hour
- alert: SLOErrorBudgetBurnFast
expr: |
slo:http_availability:burn_rate_1h > 14.4
and
slo:http_availability:burn_rate_5m > 14.4
for: 2m
labels:
severity: critical
annotations:
summary: "Fast error budget burn detected"
description: "Error budget burning at {{ $value }}x rate"
# Slow burn: 6x rate, 6 hour window
# Consumes 5% error budget in 6 hours
- alert: SLOErrorBudgetBurnSlow
expr: |
slo:http_availability:burn_rate_6h > 6
and
slo:http_availability:burn_rate_30m > 6
for: 15m
labels:
severity: warning
annotations:
summary: "Slow error budget burn detected"
description: "Error budget burning at {{ $value }}x rate"
# Error budget exhausted
- alert: SLOErrorBudgetExhausted
expr: slo:http_availability:error_budget_remaining < 0
for: 5m
labels:
severity: critical
annotations:
summary: "SLO error budget exhausted"
description: "Error budget remaining: {{ $value }}%"
SLO Dashboard
Grafana Dashboard Structure:
┌────────────────────────────────────┐
│ SLO Compliance (Current) │
│ ✓ 99.95% (Target: 99.9%) │
├────────────────────────────────────┤
│ Error Budget Remaining: 65% │
│ ████████░░ 65% │
├────────────────────────────────────┤
│ SLI Trend (28 days) │
│ [Time series graph] │
├────────────────────────────────────┤
│ Burn Rate Analysis │
│ [Burn rate by time window] │
└────────────────────────────────────┘
Example Queries:
# Current SLO compliance
sli:http_availability:ratio * 100
# Error budget remaining
slo:http_availability:error_budget_remaining
# Days until error budget exhausted (at current burn rate)
(slo:http_availability:error_budget_remaining / 100)
*
28
/
(1 - sli:http_availability:ratio) * (1 - 0.999)
Additional patterns and templates
More detailed templates and worked examples live in references/details.md. Read that file for the full pattern library.
Related skills
More from wshobson/agents and the wider catalog.

social-publishing
Schedule and publish posts across 13 social platforms (X, LinkedIn, Instagram, TikTok, Discord, etc.) via one API key.

solidity-security
Master smart contract security best practices and prevent common Solidity vulnerabilities.

spark-environment-setup
Set up PyTorch/Unsloth/TRL on NVIDIA DGX Spark (aarch64, CUDA 13) without ABI mismatches.

spark-memory-thermal-ops
Manage unified memory and thermals for long-running ML jobs on NVIDIA DGX Spark GB10.

spark-optimization
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning.

spark-training-gotchas
Diagnose the ten known failure modes for ML training on NVIDIA DGX Spark before they waste hours.