PluginBench
Skill
Pass
Audit score 90

slo-implementation

wshobson/agents

Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for measurable reliability targets.

What is slo-implementation?

Framework for establishing and tracking service reliability using SLIs, SLOs, and error budgets. Use this when defining reliability targets, implementing SRE practices, measuring service performance, or creating SLO-based alerting strategies.

  • Define SLIs (availability, latency, durability) with Prometheus queries
  • Set SLO targets and calculate error budgets from reliability goals
  • Implement Prometheus recording rules to track SLI and SLO compliance
  • Create multi-window burn-rate alerts (fast and slow burn detection)
  • Build Grafana dashboards to visualize SLO compliance and error budget status
  • Track error budget consumption and forecast budget exhaustion

How to install slo-implementation

npx skills add https://github.com/wshobson/agents --skill slo-implementation
Prerequisites
  • Prometheus or compatible metrics system collecting service metrics
  • Grafana or similar dashboard tool for visualization
  • HTTP request metrics (status codes, latency) or equivalent service metrics
Claude Code
Cursor
Windsurf
Cline

How to use slo-implementation

  1. 1.Define your SLIs by selecting metric types (availability, latency, durability) and writing Prometheus queries
  2. 2.Set SLO targets based on user expectations, business requirements, and current performance
  3. 3.Calculate error budgets using the formula: Error Budget = 1 - SLO Target
  4. 4.Create Prometheus recording rules for SLI calculations and SLO compliance tracking
  5. 5.Implement multi-window burn-rate alert rules (fast burn at 14.4x, slow burn at 6x)
  6. 6.Build a Grafana dashboard displaying current SLO compliance, error budget remaining, and burn-rate trends
  7. 7.Define an error budget policy to trigger actions (normal velocity, postpone risky changes, feature freeze) based on remaining budget

Use cases

Good for
  • Establish 99.9% availability SLO for an API with automated burn-rate alerting
  • Implement error budget policy to balance feature velocity with reliability
  • Create latency SLOs (e.g., p95 < 500ms) and track compliance over 28-day windows
  • Monitor error budget burn rates to trigger feature freezes when budget is low
  • Dashboard showing current SLO status, remaining error budget, and trend analysis
Who it's for
  • Site Reliability Engineers (SREs) implementing SRE practices
  • Platform and infrastructure teams defining service reliability targets
  • DevOps engineers setting up monitoring and alerting for services
  • Engineering leaders balancing innovation velocity with reliability requirements

slo-implementation FAQ

What's the difference between SLA, SLO, and SLI?

SLA is a contract with customers; SLO is an internal reliability target; SLI is the actual measurement. SLIs feed into SLOs, which are bounded by SLAs.

How do I choose an appropriate SLO target?

Consider user expectations, business requirements, current performance, cost of reliability, and competitor benchmarks. Common targets range from 99% to 99.99% depending on service criticality.

What is error budget and how is it used?

Error budget is the allowable downtime: 1 - SLO Target. For 99.9% SLO, you have 43.2 minutes/month. Use it to decide when to freeze features (0% remaining) or proceed normally (100% remaining).

Why use multi-window burn-rate alerts instead of simple threshold alerts?

Burn-rate alerts detect both fast outages (14.4x rate over 1 hour) and slow degradation (6x rate over 6 hours), triggering appropriate responses before error budget is exhausted.

How often should I review and adjust SLO targets?

Review quarterly or when business requirements change. Adjust targets based on actual performance, user feedback, and cost-benefit analysis of increased reliability.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: slo-implementation description: Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. Use when establishing reliability targets, implementing SRE practices, or measuring service performance.

SLO Implementation

Framework for defining and implementing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

Purpose

Implement measurable reliability targets using SLIs, SLOs, and error budgets to balance reliability with innovation velocity.

When to Use

  • Define service reliability targets
  • Measure user-perceived reliability
  • Implement error budgets
  • Create SLO-based alerts
  • Track reliability goals

SLI/SLO/SLA Hierarchy

SLA (Service Level Agreement)
  ↓ Contract with customers
SLO (Service Level Objective)
  ↓ Internal reliability target
SLI (Service Level Indicator)
  ↓ Actual measurement

Defining SLIs

Common SLI Types

1. Availability SLI

# Successful requests / Total requests
sum(rate(http_requests_total{status!~"5.."}[28d]))
/
sum(rate(http_requests_total[28d]))

2. Latency SLI

# Requests below latency threshold / Total requests
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
/
sum(rate(http_request_duration_seconds_count[28d]))

3. Durability SLI

# Successful writes / Total writes
sum(storage_writes_successful_total)
/
sum(storage_writes_total)

Setting SLO Targets

Availability SLO Examples

SLO %Downtime/MonthDowntime/Year
99%7.2 hours3.65 days
99.9%43.2 minutes8.76 hours
99.95%21.6 minutes4.38 hours
99.99%4.32 minutes52.56 minutes

Choose Appropriate SLOs

Consider:

  • User expectations
  • Business requirements
  • Current performance
  • Cost of reliability
  • Competitor benchmarks

Example SLOs:

slos:
  - name: api_availability
    target: 99.9
    window: 28d
    sli: |
      sum(rate(http_requests_total{status!~"5.."}[28d]))
      /
      sum(rate(http_requests_total[28d]))

  - name: api_latency_p95
    target: 99
    window: 28d
    sli: |
      sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
      /
      sum(rate(http_request_duration_seconds_count[28d]))

Error Budget Calculation

Error Budget Formula

Error Budget = 1 - SLO Target

Example:

  • SLO: 99.9% availability
  • Error Budget: 0.1% = 43.2 minutes/month
  • Current Error: 0.05% = 21.6 minutes/month
  • Remaining Budget: 50%

Error Budget Policy

error_budget_policy:
  - remaining_budget: 100%
    action: Normal development velocity
  - remaining_budget: 50%
    action: Consider postponing risky changes
  - remaining_budget: 10%
    action: Freeze non-critical changes
  - remaining_budget: 0%
    action: Feature freeze, focus on reliability

SLO Implementation

Prometheus Recording Rules

# SLI Recording Rules
groups:
  - name: sli_rules
    interval: 30s
    rules:
      # Availability SLI
      - record: sli:http_availability:ratio
        expr: |
          sum(rate(http_requests_total{status!~"5.."}[28d]))
          /
          sum(rate(http_requests_total[28d]))

      # Latency SLI (requests < 500ms)
      - record: sli:http_latency:ratio
        expr: |
          sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
          /
          sum(rate(http_request_duration_seconds_count[28d]))

  - name: slo_rules
    interval: 5m
    rules:
      # SLO compliance (1 = meeting SLO, 0 = violating)
      - record: slo:http_availability:compliance
        expr: sli:http_availability:ratio >= bool 0.999

      - record: slo:http_latency:compliance
        expr: sli:http_latency:ratio >= bool 0.99

      # Error budget remaining (percentage)
      - record: slo:http_availability:error_budget_remaining
        expr: |
          (sli:http_availability:ratio - 0.999) / (1 - 0.999) * 100

      # Error budget burn rate
      - record: slo:http_availability:burn_rate_5m
        expr: |
          (1 - (
            sum(rate(http_requests_total{status!~"5.."}[5m]))
            /
            sum(rate(http_requests_total[5m]))
          )) / (1 - 0.999)

SLO Alerting Rules

groups:
  - name: slo_alerts
    interval: 1m
    rules:
      # Fast burn: 14.4x rate, 1 hour window
      # Consumes 2% error budget in 1 hour
      - alert: SLOErrorBudgetBurnFast
        expr: |
          slo:http_availability:burn_rate_1h > 14.4
          and
          slo:http_availability:burn_rate_5m > 14.4
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Fast error budget burn detected"
          description: "Error budget burning at {{ $value }}x rate"

      # Slow burn: 6x rate, 6 hour window
      # Consumes 5% error budget in 6 hours
      - alert: SLOErrorBudgetBurnSlow
        expr: |
          slo:http_availability:burn_rate_6h > 6
          and
          slo:http_availability:burn_rate_30m > 6
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "Slow error budget burn detected"
          description: "Error budget burning at {{ $value }}x rate"

      # Error budget exhausted
      - alert: SLOErrorBudgetExhausted
        expr: slo:http_availability:error_budget_remaining < 0
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "SLO error budget exhausted"
          description: "Error budget remaining: {{ $value }}%"

SLO Dashboard

Grafana Dashboard Structure:

┌────────────────────────────────────┐
│ SLO Compliance (Current)           │
│ ✓ 99.95% (Target: 99.9%)          │
├────────────────────────────────────┤
│ Error Budget Remaining: 65%        │
│ ████████░░ 65%                     │
├────────────────────────────────────┤
│ SLI Trend (28 days)                │
│ [Time series graph]                │
├────────────────────────────────────┤
│ Burn Rate Analysis                 │
│ [Burn rate by time window]         │
└────────────────────────────────────┘

Example Queries:

# Current SLO compliance
sli:http_availability:ratio * 100

# Error budget remaining
slo:http_availability:error_budget_remaining

# Days until error budget exhausted (at current burn rate)
(slo:http_availability:error_budget_remaining / 100)
*
28
/
(1 - sli:http_availability:ratio) * (1 - 0.999)

Additional patterns and templates

More detailed templates and worked examples live in references/details.md. Read that file for the full pattern library.