PluginBench
Skill
Pass
Audit score 90

distributed-tracing

wshobson/agents

Track requests across microservices with Jaeger and Tempo to identify latency and bottlenecks.

What is distributed-tracing?

Implement distributed tracing to visualize request flows across microservices, understand dependencies, and pinpoint performance issues. Use this when debugging latency problems, analyzing service interactions, or building observability into distributed systems.

  • Track requests across service boundaries with trace propagation
  • Identify latency bottlenecks and slow service dependencies
  • Correlate logs with traces using trace IDs for unified debugging
  • Sample traces efficiently (1-10% in production) to minimize overhead
  • Tag spans with contextual metadata (user_id, request_id) for filtering and analysis
  • Monitor instrumentation overhead and set alerts for trace errors

How to install distributed-tracing

npx skills add https://github.com/wshobson/agents --skill distributed-tracing
Prerequisites
  • Jaeger or Tempo collector running and accessible
  • OpenTelemetry SDK installed in your application
  • Network connectivity from services to the tracing collector
  • Logging system configured to accept trace IDs for correlation
Claude Code
Cursor
Windsurf
Cline

How to use distributed-tracing

  1. 1.Install OpenTelemetry SDK and exporters for your language
  2. 2.Configure the tracer to point to your Jaeger or Tempo collector endpoint
  3. 3.Add instrumentation to track requests at service entry points
  4. 4.Propagate trace context across all service-to-service calls
  5. 5.Set sampling rate (1-10% for production) in exporter configuration
  6. 6.Add meaningful tags and span events to important operations
  7. 7.Correlate logs with traces by including trace_id in log output
  8. 8.Monitor collector performance and adjust batch processor settings if needed

Use cases

Good for
  • Debug why a multi-service request is slow by examining the full trace path
  • Understand how services depend on each other by analyzing request flows
  • Identify which microservice is causing error propagation in a distributed call chain
  • Correlate application logs with traces to investigate failures across services
  • Analyze request patterns and dependencies for capacity planning
Who it's for
  • Backend engineers debugging microservices
  • DevOps/SRE teams implementing observability
  • Architects designing distributed systems
  • Performance engineers analyzing latency issues

distributed-tracing FAQ

What sampling rate should I use in production?

Start with 1-10% sampling to balance observability with overhead. Adjust based on traffic volume and storage capacity. Use tail-based sampling for error traces.

How do I correlate logs with traces?

Extract the trace_id from the current span context and include it in all log messages. This allows you to jump between logs and traces in your visualization tool.

What's the typical CPU overhead of distributed tracing?

Well-configured tracing should add <1% CPU overhead. If higher, reduce sampling rate, use batch span processors, or optimize exporter settings.

How do I propagate trace context between services?

Use standard propagation formats (W3C Trace Context or Jaeger) in HTTP headers or message metadata. OpenTelemetry SDKs handle this automatically if configured correctly.

What should I do if no traces are appearing?

Verify the collector endpoint is reachable, check sampling configuration isn't set to 0%, review application logs for export errors, and confirm network connectivity between services and collector.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: distributed-tracing description: Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.

Distributed Tracing

Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices.

Purpose

Track requests across distributed systems to understand latency, dependencies, and failure points.

When to Use

  • Debug latency issues
  • Understand service dependencies
  • Identify bottlenecks
  • Trace error propagation
  • Analyze request paths

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

  1. Sample appropriately (1-10% in production)
  2. Add meaningful tags (user_id, request_id)
  3. Propagate context across all service boundaries
  4. Log exceptions in spans
  5. Use consistent naming for operations
  6. Monitor tracing overhead (<1% CPU impact)
  7. Set up alerts for trace errors
  8. Implement distributed context (baggage)
  9. Use span events for important milestones
  10. Document instrumentation standards

Integration with Logging

Correlated Logs

import logging
from opentelemetry import trace

logger = logging.getLogger(__name__)

def process_request():
    span = trace.get_current_span()
    trace_id = span.get_span_context().trace_id

    logger.info(
        "Processing request",
        extra={"trace_id": format(trace_id, '032x')}
    )

Troubleshooting

No traces appearing:

  • Check collector endpoint
  • Verify network connectivity
  • Check sampling configuration
  • Review application logs

High latency overhead:

  • Reduce sampling rate
  • Use batch span processor
  • Check exporter configuration

Related Skills

  • prometheus-configuration - For metrics
  • grafana-dashboards - For visualization
  • slo-implementation - For latency SLOs