distributed-tracing
wshobson/agents
Track requests across microservices with Jaeger and Tempo to identify latency and bottlenecks.
What is distributed-tracing?
Implement distributed tracing to visualize request flows across microservices, understand dependencies, and pinpoint performance issues. Use this when debugging latency problems, analyzing service interactions, or building observability into distributed systems.
- Track requests across service boundaries with trace propagation
- Identify latency bottlenecks and slow service dependencies
- Correlate logs with traces using trace IDs for unified debugging
- Sample traces efficiently (1-10% in production) to minimize overhead
- Tag spans with contextual metadata (user_id, request_id) for filtering and analysis
- Monitor instrumentation overhead and set alerts for trace errors
How to install distributed-tracing
npx skills add https://github.com/wshobson/agents --skill distributed-tracing- Jaeger or Tempo collector running and accessible
- OpenTelemetry SDK installed in your application
- Network connectivity from services to the tracing collector
- Logging system configured to accept trace IDs for correlation
How to use distributed-tracing
- 1.Install OpenTelemetry SDK and exporters for your language
- 2.Configure the tracer to point to your Jaeger or Tempo collector endpoint
- 3.Add instrumentation to track requests at service entry points
- 4.Propagate trace context across all service-to-service calls
- 5.Set sampling rate (1-10% for production) in exporter configuration
- 6.Add meaningful tags and span events to important operations
- 7.Correlate logs with traces by including trace_id in log output
- 8.Monitor collector performance and adjust batch processor settings if needed
Use cases
- Debug why a multi-service request is slow by examining the full trace path
- Understand how services depend on each other by analyzing request flows
- Identify which microservice is causing error propagation in a distributed call chain
- Correlate application logs with traces to investigate failures across services
- Analyze request patterns and dependencies for capacity planning
- Backend engineers debugging microservices
- DevOps/SRE teams implementing observability
- Architects designing distributed systems
- Performance engineers analyzing latency issues
distributed-tracing FAQ
Start with 1-10% sampling to balance observability with overhead. Adjust based on traffic volume and storage capacity. Use tail-based sampling for error traces.
Extract the trace_id from the current span context and include it in all log messages. This allows you to jump between logs and traces in your visualization tool.
Well-configured tracing should add <1% CPU overhead. If higher, reduce sampling rate, use batch span processors, or optimize exporter settings.
Use standard propagation formats (W3C Trace Context or Jaeger) in HTTP headers or message metadata. OpenTelemetry SDKs handle this automatically if configured correctly.
Verify the collector endpoint is reachable, check sampling configuration isn't set to 0%, review application logs for export errors, and confirm network connectivity between services and collector.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: distributed-tracing description: Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.
Distributed Tracing
Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices.
Purpose
Track requests across distributed systems to understand latency, dependencies, and failure points.
When to Use
- Debug latency issues
- Understand service dependencies
- Identify bottlenecks
- Trace error propagation
- Analyze request paths
Detailed patterns and worked examples
Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
Best Practices
- Sample appropriately (1-10% in production)
- Add meaningful tags (user_id, request_id)
- Propagate context across all service boundaries
- Log exceptions in spans
- Use consistent naming for operations
- Monitor tracing overhead (<1% CPU impact)
- Set up alerts for trace errors
- Implement distributed context (baggage)
- Use span events for important milestones
- Document instrumentation standards
Integration with Logging
Correlated Logs
import logging
from opentelemetry import trace
logger = logging.getLogger(__name__)
def process_request():
span = trace.get_current_span()
trace_id = span.get_span_context().trace_id
logger.info(
"Processing request",
extra={"trace_id": format(trace_id, '032x')}
)
Troubleshooting
No traces appearing:
- Check collector endpoint
- Verify network connectivity
- Check sampling configuration
- Review application logs
High latency overhead:
- Reduce sampling rate
- Use batch span processor
- Check exporter configuration
Related Skills
prometheus-configuration- For metricsgrafana-dashboards- For visualizationslo-implementation- For latency SLOs
Related skills
More from wshobson/agents and the wider catalog.

dotnet-backend-patterns
Master C#/.NET backend patterns for robust APIs, MCP servers, and enterprise applications.

e2e-testing-patterns
Master end-to-end testing with Playwright and Cypress to build reliable, maintainable test suites.

embedding-strategies
Select and optimize embedding models for semantic search and RAG applications.

employment-contract-templates
Create legally sound employment contracts, offer letters, and HR policy documents with templates and best practices.

error-handling-patterns
Master error handling patterns across languages to build resilient, fault-tolerant applications.

eval-harness-first
Build the evaluation harness that gates fine-tuning — goldens, graders, judge calibration, and baselines.