golang-performance
samber/cc-skills-golang
Go performance optimization patterns: identify bottlenecks via profiling, then apply the right optimization pattern.
What is golang-performance?
A performance engineering guide for Go that teaches iterative optimization methodology: profile first, diagnose the bottleneck type, apply one targeted optimization, re-measure, and repeat. Covers allocation reduction, CPU efficiency, memory layout, GC tuning, pooling, caching, and hot-path optimization. Use when profiling has identified a specific bottleneck or during performance code review.
- Provides decision tree to classify bottlenecks (allocations, CPU, GC, I/O, caching, algorithm) from pprof signals
- Teaches iterative cycle: define metric → baseline benchmark → diagnose → improve one thing → compare with benchstat
- Covers memory optimization (struct alignment, sync.Pool, backing array leaks) and CPU optimization (inlining, cache locality, reflection avoidance)
- Guides I/O and networking tuning (HTTP transport config, connection pools, streaming)
- Identifies common mistakes (optimizing without profiling, default http.Client limits, logging in hot loops, unsafe without proof)
- Distinguishes from related skills: golang-benchmark (measurement methodology) and golang-troubleshooting (debugging workflow)
How to install golang-performance
npx skills add https://github.com/samber/cc-skills-golang --skill golang-performance- Go toolchain installed
- benchstat: `go install golang.org/x/perf/cmd/benchstat@latest`
- Familiarity with `go test -bench` and `go tool pprof` for profiling
How to use golang-performance
- 1.Profile your code with pprof to identify the actual bottleneck (CPU, allocations, GC, I/O, or caching)
- 2.Consult the decision tree in the skill to classify your bottleneck type
- 3.Write an atomic benchmark isolating the function under test
- 4.Record baseline: `go test -bench=BenchmarkFunc -benchmem -count=6 ./pkg/... | tee /tmp/report-1.txt`
- 5.Apply ONE optimization at a time with an explanatory comment referencing the pattern used
- 6.Re-run benchmark and compare: `benchstat /tmp/report-1.txt /tmp/report-2.txt`
- 7.Commit with `perf(scope): summary` type and paste benchstat output in commit body
- 8.Repeat for next bottleneck, incrementing report number
Use cases
- After pprof identifies high allocation rate in heap profile, apply memory optimization patterns to reduce GC pressure
- When benchmarks show a function dominates CPU profile, use CPU optimization techniques like inlining and cache locality tuning
- During code review of a service, fan out three sub-agents to scan for structural anti-patterns (connection pools, unbounded goroutines, wrong data structures)
- When distributed tracing shows 90% of latency is external (DB query or API call), diagnose and optimize that component instead of Go code
- Comparing multiple candidate optimizations for the same bottleneck by implementing each in isolated worktree, then benchmarking serially
- Go backend engineers optimizing latency, throughput, or memory usage in production services
- Performance reviewers auditing packages or services for structural anti-patterns
- Teams performing iterative optimization cycles with statistical rigor (benchstat validation)
golang-performance FAQ
No. Intuition about bottlenecks is wrong ~80% of the time. Always profile with pprof first to find actual hot spots, or you risk optimizing the wrong thing.
Optimize the database component instead — query tuning, caching, connection pools, or circuit breakers. Reducing Go allocations won't help if the bottleneck is external I/O wait. Use fgprof to distinguish on-CPU vs off-CPU time.
Use benchstat to confirm statistical significance. Only use unsafe or complex patterns if profiling shows >10% improvement in a verified hot path. Document the optimization with benchmark numbers so future readers understand why it exists.
golang-benchmark teaches measurement methodology (writing good benchmarks, avoiding contamination). This skill teaches optimization patterns and the iterative cycle to apply after you've identified a bottleneck.
No. Concurrent benchmark runs on shared CPU contaminate results even if implementations were built in parallel. Run benchmarks serially and use separate worktrees to compare candidate optimizations.
Full instructions (SKILL.md)
Source of truth, from samber/cc-skills-golang.
name: golang-performance
description: "Golang performance optimization patterns and methodology - if X bottleneck, then apply Y. Covers allocation reduction, CPU efficiency, memory layout, GC tuning, pooling, caching, and hot-path optimization. Use when profiling or benchmarks have identified a bottleneck and you need the right optimization pattern to fix it. Also use when performing performance code review to suggest improvements or benchmarks that could help identify quick performance gains. Not for measurement methodology (→ See samber/cc-skills-golang@golang-benchmark skill) or debugging workflow (→ See samber/cc-skills-golang@golang-troubleshooting skill)."
user-invocable: true
license: MIT
compatibility: Designed for Claude Code, Codex or similar harness, and for projects using Golang.
metadata:
author: samber
version: "1.3.2"
openclaw:
emoji: "🏎"
homepage: https://github.com/samber/cc-skills-golang
requires:
bins:
- go
- benchstat
install:
- kind: go
package: golang.org/x/perf/cmd/benchstat@latest
bins: [benchstat]
allowed-tools: Read Edit Write Glob Grep Bash(go:) Bash(golangci-lint:) Bash(git:) Agent WebFetch Bash(benchstat:) Bash(fieldalignment:) Bash(staticcheck:) Bash(curl:) Bash(fgprof:) Bash(perf:*) WebSearch AskUserQuestion EnterWorktree ExitWorktree
paths:
- "**/*.go"
Persona: You are a Go performance engineer. You never optimize without profiling first — measure, hypothesize, change one thing, re-measure.
Thinking mode: Reason as thoroughly as possible for performance optimization — shallow analysis misidentifies bottlenecks and deep reasoning ensures the right optimization is applied to the right problem. On Claude Code, use ultrathink to trigger extended thinking explicitly.
Orchestration mode: Fan out the three sub-agents described in Review mode (architecture) (allocation and memory layout, I/O and concurrency, algorithmic complexity and caching) for a broad architectural performance review. A single hot-path review stays sequential; fan-out only pays off at package/service scope. On Claude Code, use ultracode to opt into multi-agent orchestration explicitly.
Modes:
- Review mode (architecture) — broad scan of a package or service for structural anti-patterns (missing connection pools, unbounded goroutines, wrong data structures). Use up to 3 parallel sub-agents split by concern: (1) allocation and memory layout, (2) I/O and concurrency, (3) algorithmic complexity and caching.
- Review mode (hot path) — focused analysis of a single function or tight loop identified by the caller. Work sequentially; one sub-agent is sufficient.
- Optimize mode — a bottleneck has been identified by profiling. Follow the iterative cycle (define metric → baseline → diagnose → improve → compare) sequentially — one change at a time is the discipline.
Dependencies:
- benchstat:
go install golang.org/x/perf/cmd/benchstat@latest
Go Performance Optimization
Core Philosophy
- Profile before optimizing — intuition about bottlenecks is wrong ~80% of the time. Use pprof to find actual hot spots (→ See
samber/cc-skills-golang@golang-troubleshootingskill) - Allocation reduction yields the biggest ROI — Go's GC is fast but not free. Reducing allocations per request often matters more than micro-optimizing CPU
- Document optimizations — add code comments explaining why a pattern is faster, with benchmark numbers when available. Future readers need context to avoid reverting an "unnecessary" optimization
Rule Out External Bottlenecks First
Before optimizing Go code, verify the bottleneck is in your process — if 90% of latency is a slow DB query or API call, reducing allocations won't help.
Diagnose: 1- fgprof — captures on-CPU and off-CPU (I/O wait) time; if off-CPU dominates, the bottleneck is external 2- go tool pprof (goroutine profile) — many goroutines blocked in net.(*conn).Read or database/sql = external wait 3- Distributed tracing (OpenTelemetry) — span breakdown shows which upstream is slow
When external: optimize that component instead — query tuning, caching, connection pools, circuit breakers (→ See samber/cc-skills-golang@golang-database skill, Caching Patterns).
Iterative Optimization Methodology
The cycle: Define Goals → Benchmark → Diagnose → Improve → Benchmark
- Define your metric — latency, throughput, memory, or CPU? Without a target, optimizations are random
- Write an atomic benchmark — isolate one function per benchmark to avoid result contamination (→ See
samber/cc-skills-golang@golang-benchmarkskill) - Measure baseline —
go test -bench=BenchmarkMyFunc -benchmem -count=6 ./pkg/... | tee /tmp/report-1.txt - Diagnose — use the Diagnose lines in each deep-dive section to pick the right tool
- Improve — apply ONE optimization at a time with an explanatory comment
- Compare —
benchstat /tmp/report-1.txt /tmp/report-2.txtto confirm statistical significance - Commit — paste the benchstat output in the commit body so reviewers and future readers see the exact improvement; follow the
perf(scope): summarycommit type - Repeat — increment report number, tackle next bottleneck
Refer to library documentation for known patterns before inventing custom solutions. Keep all /tmp/report-*.txt files as an audit trail.
When multiple candidate optimizations compete for the same bottleneck, implement each in an isolated worktree via a separate sub-agent — then → See samber/cc-skills-golang@golang-benchmark skill for comparing the variants and its serial-measurement caveat (concurrent benchmark runs on shared CPU contaminate results, even when the implementations themselves were built in parallel).
Decision Tree: Where Is Time Spent?
| Bottleneck | Signal (from pprof) | Action |
|---|---|---|
| Too many allocations | alloc_objects high in heap profile | Memory optimization |
| CPU-bound hot loop | function dominates CPU profile | CPU optimization |
| GC pauses / OOM | high GC%, container limits | Runtime tuning |
| Network / I/O latency | goroutines blocked on I/O | I/O & networking |
| Repeated expensive work | same computation/fetch multiple times | Caching patterns |
| Wrong algorithm | O(n²) where O(n) exists | Algorithmic complexity |
| Lock contention | mutex/block profile hot | → See samber/cc-skills-golang@golang-concurrency skill |
| Slow queries | DB time dominates traces | → See samber/cc-skills-golang@golang-database skill |
Common Mistakes
| Mistake | Fix |
|---|---|
| Optimizing without profiling | Profile with pprof first — intuition is wrong ~80% of the time |
Default http.Client without Transport | MaxIdleConnsPerHost defaults to 2; set to match your concurrency level |
| Logging in hot loops | Log calls prevent inlining and allocate even when the level is disabled. Use slog.LogAttrs |
panic/recover as control flow | panic allocates a stack trace and unwinds the stack; use error returns |
unsafe without benchmark proof | Only justified when profiling shows >10% improvement in a verified hot path |
| No GC tuning in containers | Set GOMEMLIMIT to 80-90% of container memory to prevent OOM kills |
reflect.DeepEqual in production | 50-200x slower than typed comparison; use slices.Equal, maps.Equal, bytes.Equal |
Deep Dives
- Memory Optimization — allocation patterns, backing array leaks, sync.Pool, struct alignment
- CPU Optimization — inlining, cache locality, false sharing, ILP, reflection avoidance
- I/O & Networking — HTTP transport config, streaming, JSON performance, cgo, batch operations
- Runtime Tuning — GOGC, GOMEMLIMIT, GC diagnostics, GOMAXPROCS, PGO
- Caching Patterns — algorithmic complexity, compiled patterns, singleflight, work avoidance
- Production Observability — Prometheus metrics, PromQL queries, continuous profiling, alerting rules
CI Regression Detection
Automate benchmark comparison in CI to catch regressions before they reach production. → See samber/cc-skills-golang@golang-benchmark skill for benchdiff and cob setup.
Cross-References
- → See
samber/cc-skills-golang@golang-benchmarkskill for benchmarking methodology,benchstat, andb.Loop()(Go 1.24+) - → See
samber/cc-skills-golang@golang-troubleshootingskill for pprof workflow, escape analysis diagnostics, and performance debugging - → See
samber/cc-skills-golang@golang-data-structuresskill for slice/map preallocation andstrings.Builder - → See
samber/cc-skills-golang@golang-concurrencyskill for worker pools,sync.PoolAPI, goroutine lifecycle, and lock contention - → See
samber/cc-skills-golang@golang-safetyskill for defer in loops, slice backing array aliasing - → See
samber/cc-skills-golang@golang-databaseskill for connection pool tuning and batch processing - → See
samber/cc-skills-golang@golang-observabilityskill for continuous profiling in production
Related skills
More from samber/cc-skills-golang and the wider catalog.

golang-pkg-go-dev
Query pkg.go.dev for Go package docs, symbols, versions, vulnerabilities, and importers via godig CLI or MCP server.

golang-popular-libraries
Vetted Go library and framework recommendations by category—web, database, testing, logging, messaging—with maturity signals and stdlib-first guidance.

golang-project-layout
Establish Go project structure with cmd/internal/pkg conventions, module naming, workspaces, and config files.

golang-refactoring
Safe, at-scale Go refactoring with coverage-adaptive safety nets and behavior-preserving transforms.

golang-safety
Defensive Go coding: prevent nil panics, slice aliasing, numeric truncation, and resource leaks.

golang-samber-do
Type-safe dependency injection for Go using samber/do with generics, scopes, and lifecycle management.