PluginBench
Skill
Pass
Audit score 90

performance-optimization

gamedev-skills/awesome-gamedev-agent-skills

Measure, find the bottleneck, apply the right fix—CPU pooling/allocation control or GPU batching/draw calls.

What is performance-optimization?

Performance optimization is a measurement discipline: profile with the engine's built-in profiler, identify whether CPU or GPU is the bottleneck, then apply targeted fixes like object pooling, draw-call batching, and allocation control. Use this when frame rate is low, uneven, or missing a target (60 FPS desktop, 30/60 mobile).

  • Profile frame time and identify CPU vs. GPU bottlenecks using the engine profiler
  • Calculate frame budgets (60 FPS = 16.67 ms) and allocate sub-budgets per subsystem
  • Implement object pooling to eliminate per-frame allocations and GC spikes
  • Reduce draw calls through atlasing, material sharing, GPU instancing, and static batching
  • Remove per-frame allocations, LINQ calls, and boxing in hot loops (C#/GDScript)
  • Set asset budgets (texture sizes, triangle counts, draw-call ceilings) to prevent regressions

How to install performance-optimization

npx skills add https://github.com/gamedev-skills/awesome-gamedev-agent-skills --skill performance-optimization
Claude Code
Cursor
Windsurf
Cline

How to use performance-optimization

  1. 1.Open your engine's profiler (Godot Debugger ▸ Profiler, Unity Profiler window, Unreal `stat unit`) and run a representative worst-case scene on target hardware in a release/optimized build
  2. 2.Read the frame time split: identify whether CPU (game logic, physics, scripts) or GPU (rendering, draw calls) dominates the frame budget
  3. 3.Calculate your target frame budget (e.g. 60 FPS = 16.67 ms) and allocate sub-budgets per subsystem (gameplay ~5 ms, physics ~3 ms, rendering ~4 ms, etc.)
  4. 4.Apply the matching fix: if GPU-bound, reduce draw calls via atlasing and batching; if CPU-bound, implement object pooling, remove per-frame allocations, or optimize algorithms
  5. 5.Re-measure on the same scene and hardware to confirm the fix moved the numbers; keep or revert based on data
  6. 6.Set asset budgets (texture sizes, triangle counts, draw-call ceilings) and add performance checks to verification to prevent regressions

Use cases

Good for
  • Game stutters or frame rate drops below target; profile to find whether CPU (scripts, physics) or GPU (rendering) is the bottleneck, then apply the matching fix
  • Bullets or particles spawn/destroy every frame, causing GC spikes; implement object pooling to reuse a fixed set
  • Thousands of draw calls from unbatched sprites or unique materials; atlas textures and share materials to batch into fewer calls
  • Per-frame allocations in C# Update loops feed the garbage collector; cache references and reuse buffers instead
  • Physics simulation or script logic consumes most of the frame budget; profile to confirm, then optimize the algorithm or reduce work frequency
Who it's for
  • Game developers optimizing for a target frame rate (60 FPS desktop, 30/60 mobile)
  • Programmers debugging frame drops, stutters, or hitches in existing games
  • Teams setting performance budgets and regression checks during development
  • Anyone using Godot, Unity, Unreal, or engine-agnostic profiling workflows

performance-optimization FAQ

Should I optimize before profiling?

No. Most performance fixes applied without profiling target the wrong thing and add complexity for no gain. Always profile a release build on target hardware first, find the single biggest cost, and fix that.

How do I know if the CPU or GPU is the bottleneck?

Use the engine profiler to read the frame split. If GPU time ≫ CPU time, attack draw calls, overdraw, and shaders. If CPU time dominates, attack scripts, physics, and allocations. Fixing the wrong side does nothing.

What is object pooling and when should I use it?

Object pooling reuses a fixed set of objects (bullets, particles, enemies) instead of instantiating and freeing them every frame. Use it whenever you spawn/destroy objects in hot loops to avoid memory fragmentation and garbage-collection spikes.

Why do per-frame allocations cause stutters?

Allocating memory every frame fills the managed heap (C#) or causes fragmentation (GDScript). The garbage collector then stalls the entire frame to clean up, causing a visible hitch. Cache references once and reuse buffers instead.

What is the most common GPU-side win?

Reducing draw calls. Each unique material, texture, or state change is roughly one draw call. Atlas textures, share materials, use GPU instancing, and mark static geometry as static to batch objects and cut the draw-call count.

Full instructions (SKILL.md)

Source of truth, from gamedev-skills/awesome-gamedev-agent-skills.


name: performance-optimization description: > Find and fix game performance problems methodically — measure with the engine profiler first, reason about the frame-time budget, locate the CPU-vs-GPU bottleneck, then apply the right fix: object pooling, draw-call batching, fewer allocations/GC spikes, and asset budgets. Engine- neutral method that pairs with each engine's profiler. Use when the user mentions performance, optimize, low/dropping FPS, frame drops, stutter, lag, profiler, frame budget, draw calls, batching, garbage collection/GC spikes, object pooling, or "the game runs slow".

Performance optimization

Performance work is a measurement discipline, not a bag of tricks. The method is always the same: profile → find the one bottleneck → fix that → measure again. This skill teaches that loop and the highest-leverage fixes (pooling, batching, allocation control, asset budgets), and points you at each engine's profiler. It pairs with physics-tuning for simulation cost.

When to use

  • Use when the frame rate is low or uneven, the game stutters/hitches, or it must hit a target (60 FPS desktop, 30/60 mobile) and currently doesn't.
  • Use to decide what to optimize: profile, read the frame budget, and identify whether the CPU or GPU is the bottleneck before changing any code.
  • Use to apply specific fixes: object pooling, draw-call/batch reduction, removing per-frame allocations and GC spikes, and setting asset budgets.

When not to use: for physics jitter/tunneling/timestep specifically, use physics-tuning. For the engine's concrete profiler UI and rendering settings, use that engine skill (godot-export covers some build settings; engine cores cover the rest). This skill is the cross-engine method and the shared fixes.

The golden rule: measure first, never guess

Most performance "fixes" applied without profiling target the wrong thing and add complexity for no gain. Do not optimize code you have not measured. Open the profiler, find the single biggest cost in a representative scene on representative hardware, and fix that. Re-measure to confirm the fix helped before moving on. Profile a release/optimized build where it matters — editor and debug builds lie (editor overhead, no compiler optimization).

Core workflow

  1. Define the target and reproduce. State the goal (e.g. 60 FPS = 16.67 ms/frame) and find a repeatable worst-case scene. "Sometimes slow" is unfixable; a reproducible spike is fixable.
  2. Profile before touching code. Run the engine profiler and read the frame: total frame time, and the split between CPU (game logic, physics, scripts) and GPU (rendering).
  3. Find the bottleneck — CPU or GPU. If GPU time ≫ CPU, attack draw calls/overdraw/shaders/ resolution. If CPU time dominates, attack scripts/physics/allocations. Fixing the wrong side does nothing.
  4. Fix the single biggest cost. Prefer an algorithmic win (do less work, cache, spatial partition, run less often) over micro-optimizing a hot line. Apply the matching shared fix (pooling, batching, allocation removal).
  5. Re-measure on the same scene/hardware. Confirm the number moved. Keep or revert based on data, not intuition.
  6. Set budgets so it stays fixed. Per-frame ms budgets per subsystem, plus asset budgets (texture sizes, triangle counts, draw-call ceilings); add a perf check to verification.
  7. Report measured numbers. State before/after frame time, the bottleneck found, and the fix — never "should be faster". If you could only measure in-editor, say so.

Patterns

1. Frame budget math (turn "feels slow" into a number)

target FPS → frame budget:   60 FPS = 16.67 ms   |   30 FPS = 33.3 ms   |   120 FPS = 8.33 ms
The WHOLE frame (CPU sim + render submit + GPU) must fit the budget; the GPU runs in parallel,
so the slower of CPU-frame and GPU-frame sets your FPS. Allocate sub-budgets, e.g. @60 FPS:
  gameplay/scripts ~5 ms · physics ~3 ms · rendering(CPU submit) ~4 ms · UI/other ~2 ms · slack.
If one subsystem blows its slice, that's your target — not whatever you assumed.

2. Measure with the engine profiler (do this before any fix)

Godot 4.7 : Debugger ▸ Profiler (script/physics time) and Monitors tab (FPS, draw calls, memory).
            In code: Performance.get_monitor(Performance.TIME_PROCESS) and
            Performance.get_monitor(Performance.RENDER_TOTAL_DRAW_CALLS_IN_FRAME).
Unity 6.3 LTS   : Profiler window (CPU/GPU/Memory/Rendering modules) + Frame Debugger for draw calls.
            In code: a ProfilerRecorder tracking "CPU Main Thread Frame Time" for a HUD/log.
Unreal 5  : `stat unit` (Frame/Game/Draw/GPU ms), `stat fps`, `stat scenerendering` (draw calls);
            Unreal Insights for deep traces.
# Read the split: is the Draw/GPU line the biggest, or the Game/CPU line? That decides the fix.

3. Object pooling (stop allocating/freeing in hot loops)

# Bullets, particles, enemies, damage numbers: reuse a fixed set instead of instantiate()/free()
# every frame — that thrashes memory and (in C#) feeds the GC.
var _pool: Array[Node] = []
func acquire() -> Node:
    var n: Node = _pool.pop_back() if not _pool.is_empty() else bullet_scene.instantiate()
    n.set_process(true); n.visible = true
    return n
func release(n: Node) -> void:
    n.set_process(false); n.visible = false       # disable + hide; DON'T free
    _pool.append(n)                                # back to the pool for reuse
# RIGHT: pre-warm the pool at load; reuse. WRONG: instantiate()/queue_free() per shot.

4. Cut draw calls (the most common GPU-side win)

Each unique material/texture/state change is roughly a draw call; thousands of them stall the GPU.
- Atlas textures and share materials so sprites/meshes batch into one call.
- Identical meshes → GPU instancing (Unity), MultiMesh / MultiMeshInstance (Godot), Instanced
  Static Mesh (Unreal).
- Static geometry → static batching / baking; mark non-moving objects static.
- Reduce overdraw: limit large overlapping transparent/particle layers (they re-shade pixels).
- Fewer real-time lights/shadows; bake lighting where it doesn't move.
Measure draw calls before and after — the count should drop, and so should GPU frame time.

5. Kill per-frame allocations (GC spikes = stutter)

// Unity 6.3 LTS (C#). Allocating every frame fills the managed heap; the GC then stalls a frame.
// WRONG (allocates each call): foreach (var e in FindObjectsOfType<Enemy>()) ...  // + LINQ, new[]
// RIGHT: cache references once, reuse buffers, avoid LINQ/boxing in Update.
void Update() {
    _hits = Physics.RaycastNonAlloc(ray, _hitBuffer);   // reuse a preallocated array
    for (int i = 0; i < _hits; i++) { /* ... */ }       // no per-frame allocation
}
// Godot/GDScript: avoid building new arrays/dictionaries every frame in _process; reuse them.

Pitfalls

  • Optimizing without profiling. The intuitive culprit is usually wrong. Measure first, every time.
  • Profiling the editor / a debug build. Editor overhead and unoptimized code mislead. Profile a release build on target hardware for real numbers.
  • Fixing the wrong side. Micro-optimizing CPU code when the GPU is the bottleneck (or vice versa) changes nothing. Check the CPU-vs-GPU split first.
  • Micro-optimizing over algorithm. Shaving a function when an O(n²) loop or a per-frame full-scene query is the real cost. Reduce the work, don't polish it.
  • Instantiate/free in hot loops. Spawning and destroying bullets/particles every frame causes fragmentation and GC spikes. Pool them.
  • Per-frame allocations / LINQ / boxing in Update (C#) feed the GC → periodic hitches. Cache and reuse.
  • Draw-call explosion from unique materials and unbatched sprites/meshes. Atlas, share materials, instance, batch.
  • Overdraw from stacked transparents/particles/full-screen effects re-shading pixels.
  • No budgets. Without per-subsystem ms and asset ceilings, performance silently regresses; enforce them in your build/CI checks.
  • Optimizing too early. Don't contort a prototype for performance before it's fun or measured.

References

  • For per-engine profiler walkthroughs, the CPU-vs-GPU triage flowchart, a complete pooling manager, batching/instancing rules per engine, allocation/GC guidance, LOD/culling, and asset budgets (texture sizes, triangle counts, audio, mobile thermals), read references/profiling-and-budgets.md.

Related skills

  • physics-tuning — simulation cost, fixed-step budget, sleeping bodies, broadphase layers.
  • godot-export — release/build settings that affect measured performance.
  • procedural-gen, game-ai — common CPU hotspots (generation, pathfinding) to budget and defer.
  • roguelike, tower-defense, survival-crafting — entity-heavy genres that need pooling/budgets.