PluginBench
Skill
Fail
Audit score 45

python-resilience

wshobson/agents

Automatic retries, exponential backoff, and fault-tolerant decorators for resilient Python services.

What is python-resilience?

Python resilience patterns for handling transient failures, network issues, and service outages. Use this skill when adding retry logic to external calls, implementing timeouts, building fault-tolerant microservices, or handling rate limiting and backpressure.

  • Automatic retry logic with exponential backoff and jitter using tenacity decorators
  • Distinguish between transient errors (retry) and permanent errors (fail fast)
  • HTTP status code-based retries for handling 429, 5xx responses
  • Combined exception and status code retry strategies
  • Bounded retry attempts and total duration to prevent infinite loops
  • Logging and monitoring of retry behavior

How to install python-resilience

npx skills add https://github.com/wshobson/agents --skill python-resilience
Prerequisites
  • Python 3.7+
  • tenacity library (pip install tenacity)
  • httpx or requests library for HTTP calls
Claude Code
Cursor
Windsurf
Cline

How to use python-resilience

  1. 1.Import retry decorators from tenacity
  2. 2.Define which exceptions or HTTP status codes are retryable (transient errors)
  3. 3.Apply @retry decorator to functions with appropriate stop and wait strategies
  4. 4.Set timeouts on all network calls
  5. 5.Add logging to track retry attempts
  6. 6.Test retry behavior with simulated failures

Use cases

Good for
  • Adding retry logic to external API calls that may timeout or fail temporarily
  • Implementing timeouts and circuit breaker patterns for microservices
  • Handling rate limiting (HTTP 429) with exponential backoff
  • Building resilient database connection pools with automatic reconnection
  • Creating fault-tolerant webhook handlers that gracefully degrade
Who it's for
  • Backend engineers building microservices
  • DevOps engineers designing fault-tolerant infrastructure
  • Python developers handling unreliable external dependencies
  • Site reliability engineers implementing resilience patterns

python-resilience FAQ

Should I retry all exceptions?

No. Only retry transient errors like ConnectionError, TimeoutError, and specific HTTP 5xx codes. Never retry bugs (ValueError, TypeError) or permanent failures (invalid credentials, 4xx errors).

What's the difference between exponential backoff and jitter?

Exponential backoff increases wait time between retries (1s, 2s, 4s...). Jitter adds randomness to prevent thundering herd when many clients retry simultaneously at the same time.

How do I prevent infinite retry loops?

Use bounded retries with both attempt count and total duration limits: stop_after_attempt(5) | stop_after_delay(60)

Should I retry HTTP 429 (rate limit) responses?

Yes, 429 indicates transient overload. Retry with exponential backoff to respect rate limits. Also retry 502, 503, 504 (server errors) but not 4xx client errors.

How do I test retry behavior?

Mock the underlying function to raise transient errors or return retryable status codes, then verify the decorator retries the expected number of times before succeeding or failing.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: python-resilience description: Python resilience patterns including automatic retries, exponential backoff, timeouts, and fault-tolerant decorators. Use when adding retry logic, implementing timeouts, building fault-tolerant services, or handling transient failures.

Python Resilience Patterns

Build fault-tolerant Python applications that gracefully handle transient failures, network issues, and service outages. Resilience patterns keep systems running when dependencies are unreliable.

When to Use This Skill

  • Adding retry logic to external service calls
  • Implementing timeouts for network operations
  • Building fault-tolerant microservices
  • Handling rate limiting and backpressure
  • Creating infrastructure decorators
  • Designing circuit breakers

Core Concepts

1. Transient vs Permanent Failures

Retry transient errors (network timeouts, temporary service issues). Don't retry permanent errors (invalid credentials, bad requests).

2. Exponential Backoff

Increase wait time between retries to avoid overwhelming recovering services.

3. Jitter

Add randomness to backoff to prevent thundering herd when many clients retry simultaneously.

4. Bounded Retries

Cap both attempt count and total duration to prevent infinite retry loops.

Quick Start

from tenacity import retry, stop_after_attempt, wait_exponential_jitter

@retry(
    stop=stop_after_attempt(3),
    wait=wait_exponential_jitter(initial=1, max=10),
)
def call_external_service(request: dict) -> dict:
    return httpx.post("https://api.example.com", json=request).json()

Fundamental Patterns

Pattern 1: Basic Retry with Tenacity

Use the tenacity library for production-grade retry logic. For simpler cases, consider built-in retry functionality or a lightweight custom implementation.

from tenacity import (
    retry,
    stop_after_attempt,
    stop_after_delay,
    wait_exponential_jitter,
    retry_if_exception_type,
)

TRANSIENT_ERRORS = (ConnectionError, TimeoutError, OSError)

@retry(
    retry=retry_if_exception_type(TRANSIENT_ERRORS),
    stop=stop_after_attempt(5) | stop_after_delay(60),
    wait=wait_exponential_jitter(initial=1, max=30),
)
def fetch_data(url: str) -> dict:
    """Fetch data with automatic retry on transient failures."""
    response = httpx.get(url, timeout=30)
    response.raise_for_status()
    return response.json()

Pattern 2: Retry Only Appropriate Errors

Whitelist specific transient exceptions. Never retry:

  • ValueError, TypeError - These are bugs, not transient issues
  • AuthenticationError - Invalid credentials won't become valid
  • HTTP 4xx errors (except 429) - Client errors are permanent
from tenacity import retry, retry_if_exception_type
import httpx

# Define what's retryable
RETRYABLE_EXCEPTIONS = (
    ConnectionError,
    TimeoutError,
    httpx.ConnectTimeout,
    httpx.ReadTimeout,
)

@retry(
    retry=retry_if_exception_type(RETRYABLE_EXCEPTIONS),
    stop=stop_after_attempt(3),
    wait=wait_exponential_jitter(initial=1, max=10),
)
def resilient_api_call(endpoint: str) -> dict:
    """Make API call with retry on network issues."""
    return httpx.get(endpoint, timeout=10).json()

Pattern 3: HTTP Status Code Retries

Retry specific HTTP status codes that indicate transient issues.

from tenacity import retry, retry_if_result, stop_after_attempt
import httpx

RETRY_STATUS_CODES = {429, 502, 503, 504}

def should_retry_response(response: httpx.Response) -> bool:
    """Check if response indicates a retryable error."""
    return response.status_code in RETRY_STATUS_CODES

@retry(
    retry=retry_if_result(should_retry_response),
    stop=stop_after_attempt(3),
    wait=wait_exponential_jitter(initial=1, max=10),
)
def http_request(method: str, url: str, **kwargs) -> httpx.Response:
    """Make HTTP request with retry on transient status codes."""
    return httpx.request(method, url, timeout=30, **kwargs)

Pattern 4: Combined Exception and Status Retry

Handle both network exceptions and HTTP status codes.

from tenacity import (
    retry,
    retry_if_exception_type,
    retry_if_result,
    stop_after_attempt,
    wait_exponential_jitter,
    before_sleep_log,
)
import logging
import httpx

logger = logging.getLogger(__name__)

TRANSIENT_EXCEPTIONS = (
    ConnectionError,
    TimeoutError,
    httpx.ConnectError,
    httpx.ReadTimeout,
)
RETRY_STATUS_CODES = {429, 500, 502, 503, 504}

def is_retryable_response(response: httpx.Response) -> bool:
    return response.status_code in RETRY_STATUS_CODES

@retry(
    retry=(
        retry_if_exception_type(TRANSIENT_EXCEPTIONS) |
        retry_if_result(is_retryable_response)
    ),
    stop=stop_after_attempt(5),
    wait=wait_exponential_jitter(initial=1, max=30),
    before_sleep=before_sleep_log(logger, logging.WARNING),
)
def robust_http_call(
    method: str,
    url: str,
    **kwargs,
) -> httpx.Response:
    """HTTP call with comprehensive retry handling."""
    return httpx.request(method, url, timeout=30, **kwargs)

Detailed worked examples and patterns

Detailed sections (starting with ## Advanced Patterns) live in references/details.md. Read that file when the navigation summary above is insufficient.

Best Practices Summary

  1. Retry only transient errors - Don't retry bugs or authentication failures
  2. Use exponential backoff - Give services time to recover
  3. Add jitter - Prevent thundering herd from synchronized retries
  4. Cap total duration - stop_after_attempt(5) | stop_after_delay(60)
  5. Log every retry - Silent retries hide systemic problems
  6. Use decorators - Keep retry logic separate from business logic
  7. Inject dependencies - Make infrastructure testable
  8. Set timeouts everywhere - Every network call needs a timeout
  9. Fail gracefully - Return cached/default values for non-critical paths
  10. Monitor retry rates - High retry rates indicate underlying issues