spark-training-gotchas
wshobson/agents
Diagnose the ten known failure modes for ML training on NVIDIA DGX Spark before they waste hours.
What is spark-training-gotchas?
Preflight and diagnose ten recurring failure modes (G1–G10) across launch, memory, thermals, bandwidth, and precision on DGX Spark's GB10 chip. Use when a training run fails to start, OOMs despite headroom, slows mid-run, or before any multi-hour job.
- Identify CUDA 12/13 ABI mismatches causing import errors or segfaults
- Detect flash-attn backend conflicts and SDPA fallback issues
- Diagnose unified-memory OOMs that appear despite free capacity in nvidia-smi
- Spot thermal throttling and power-draw ceilings causing mid-run slowdowns
- Measure actual memory bandwidth (180–192 GB/s) versus spec figures
- Detect global UMA resource contention from uncapped competing processes
How to install spark-training-gotchas
npx skills add https://github.com/wshobson/agents --skill spark-training-gotchas- Access to an NVIDIA DGX Spark system with GB10 chip (aarch64, SM121, 128GB unified memory)
- Python 3 and torch installed (to run quick version checks)
- Root access to run `drop_caches` command if needed for G3 diagnosis
- Familiarity with nvidia-smi, /proc/meminfo, and basic Linux system monitoring
How to use spark-training-gotchas
- 1.Run the fast triage checks: `python3 -c "import torch; print(torch.version.cuda)"` (expect 13.x for G1), `torch.cuda.get_device_capability()` (expect (12, 1) for G7), and container detection.
- 2.Execute `assets/preflight.sh` to automatically check G1, G3, G4, G7, G9 and produce a one-line-per-gotcha report with PASS/FAIL/WARN/SKIP/INFO status.
- 3.For any FAIL or WARN, consult the corresponding gotcha section (G1–G10) in the skill for the symptom, cause, and fix.
- 4.If G3 (OOM despite headroom) is suspected, drop the page cache with `sync; echo 3 > /proc/sys/vm/drop_caches` (requires root) between runs.
- 5.For G6 (UMA contention), check `references/gotcha-checks.md` G6 for other GPU-resident processes and cap or stop uncapped workloads.
- 6.Before dual-Spark setups, verify the parallelism strategy is DDP or FSDP (G10), never tensor parallelism.
Use cases
- Before starting a multi-hour or multi-epoch training job on GB10, run preflight checks to catch G1, G3, G4, G7, G9 early.
- When a training run fails to start with an import error or segfault, check G1 (CUDA ABI) and G9 (container vs. bare pip).
- When a run OOMs while nvidia-smi shows headroom, drop the page cache (G3) and verify no uncapped competing processes (G6).
- When throughput degrades partway through a run, sample temperature and power draw to distinguish thermal throttling (G4) from bandwidth limits (G5).
- When wiring two Sparks together, verify the parallelism strategy is DDP or FSDP, never tensor parallelism (G10).
- ML engineers training models on NVIDIA DGX Spark GB10 hardware
- DevOps and platform teams setting up Spark environments for multi-hour runs
- Researchers debugging mysterious OOMs, slowdowns, or reboots on Spark
- Teams deploying inference workloads and choosing between FP8 and NVFP4 precision
spark-training-gotchas FAQ
G1 is a CUDA ABI mismatch (wheels built for cu12 on a cu13 system); G9 is environment drift in bare pip. G1 is a wheel problem; G9 is a dependency-management problem. Both can cause import errors, but G1 surfaces at first .cuda() call, while G9 emerges after unrelated pip installs.
Unified memory (UMA) double-counts pages during safetensors load and CUDA allocation. nvidia-smi does not reflect the true free memory. Check /proc/meminfo and free -g instead, and drop the page cache with `sync; echo 3 > /proc/sys/vm/drop_caches` before the run.
Only if both are small and capped (e.g., <4GB LoRA + vLLM at gpu-memory-utilization<=0.5). Uncapped or near-capacity workloads compete for the global UMA pool and evict each other silently (G6). For multi-hour runs, use one heavy job per Spark.
Stay on FP8 unless your build explicitly targets sm_121a. SM121 lacks the cvt.e2m1x2 instruction, so NVFP4 runs ~32% slower without it (G7). Check torch.cuda.get_device_capability() and your build target.
Use DDP or FSDP only; never tensor parallelism (G10). ConnectX-7 is fast enough for gradient/parameter sync but too thin for TP's fine-grained traffic. Tensor parallelism is single-node only on Spark.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: spark-training-gotchas description: Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark. Use when a training run on DGX Spark fails to start, OOMs below the 128GB limit, slows down mid-run, or before any multi-hour training job on GB10.
Spark Training Gotchas
DGX Spark's GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these checks. Read this before a long run, not after hour six.
When to Use This Skill
- A training run fails to start, with an import error or a segfault that doesn't point at the real cause.
- A run OOMs while
nvidia-smistill shows headroom. - Throughput degrades partway through a run that started fine.
- Before any multi-hour or multi-epoch job on GB10.
- Wiring two Sparks together, before picking a parallelism strategy.
- Choosing between FP8 and NVFP4 for a Spark-hosted run.
Common Issues Quick Reference
| # | Symptom | Fix |
|---|---|---|
| G1 | undefined symbol / segfault | cu130 wheel or container |
| G2 | flash-attn wrong backend used | skip pip build; monkeypatch on NGC |
| G3 | OOM despite headroom | drop page cache |
| G4 | throughput drop / reboot | expect ~100W sustained cap |
| G5 | memory-bound step slow | budget 180–192 GB/s |
| G6 | cache evicted mid-run | one GPU server at a time |
| G7 | NVFP4 slower than FP8 | stay FP8 unless sm_121a |
| G8 | playbook fails outright | check upstream issues |
| G9 | env breaks after install | use a container |
| G10 | 2-Spark TP hangs | DDP/FSDP only, never TP |
The Ten Gotchas
G1: CUDA 12/13 ABI Mismatch
- SYMPTOM:
ImportError: undefined symbolnaming a CUDA function, or a segfault on the first.cuda()call. - CAUSE: most PyPI wheels link
libcudart.so.12; Spark ships CUDA 13. pip never checks CUDA ABI, so it surfaces only at import or first kernel launch. - CHECK:
references/gotcha-checks.mdG1 — the wheel's CUDA build tag. - FIX: reinstall from
download.pytorch.org/whl/cu130or use a matched container.
G2: flash-attn — Skip the pip Build, Watch Unsloth's Auto-Detect
- SYMPTOM:
pip install flash-attnstill fails/hangs. Unsloth may also silently train flash-attn over an explicitly requested SDPA. - CAUSE: no aarch64/sm_121 wheel for bare pip — but NGC
containers ship a working SM121 flash-attn, and Unsloth
auto-prefers it, dropping
attn_implementation="sdpa". - CHECK:
references/gotcha-checks.mdG2 — is flash-attn already present and working. - FIX: bare pip — skip flash-attn, use SDPA (unchanged). On
NGC — the only reliable override is the monkeypatch in
references/gotcha-checks.mdG2.
G3: UMA OOM Below 128GB
- SYMPTOM: OOM during model load/training while
nvidia-smistill reports free memory under the 128GB cap — or, on some setups,[N/A]outright instead of a number. - CAUSE: mmap and the CUDA allocator double-count pages during safetensors load; QLoRA can OOM earlier than bf16 since dequantization adds transient allocs.
- CHECK:
references/gotcha-checks.mdG3 — readfree -gand/proc/meminfo, notnvidia-smi. - FIX: drop the page cache with
sync; echo 3 > /proc/sys/vm/drop_caches— needs root, a between-run reset, not a mid-training step.
G4: Thermal Throttling
- SYMPTOM: throughput drops partway through a multi-hour run, or the box spontaneously reboots under sustained load.
- CAUSE: sustained power draw caps around 100W versus the 240W rated figure; long runs push into that ceiling and throttle or, sometimes, reboot.
- CHECK:
references/gotcha-checks.mdG4 — samplenvidia-smi --query-gpu=temperature.gpu,power.draw. - FIX: if power plateaus under 240W while temperature climbs, treat throttling as the cause; improve cooling or cap run length.
G5: Bandwidth Ceiling
- SYMPTOM: memory-bound workloads, decode-heavy RL loops especially, plateau well below expected throughput.
- CAUSE: 273 GB/s is a spec ceiling, not sustained; measured bandwidth runs 180–192 GB/s.
- CHECK:
references/gotcha-checks.mdG5 — observed step time vs. the measured range, not spec. - FIX: budget throughput from 180–192 GB/s; revise a plan built on the 273 GB/s figure.
G6: Global UMA Resource Contention
- SYMPTOM: a process's KV cache/weights get evicted mid-run silently, no OOM in its own logs.
- CAUSE: unified memory is
one global pool; an uncapped
or near-capacity process
competes with anything else
and can evict it. A small,
bounded workload doesn't — a
<4GB LoRA coexists fine
alongside vLLM capped at
gpu-memory-utilization<=0.5. - CHECK:
references/gotcha-checks.mdG6 — other GPU-resident processes and whether capped. - FIX: the one-heavy-job rule applies to uncapped or near-capacity workloads — cap or stop unrelated servers first. A small, capped workload need not stop.
G7: NVFP4 Slower Than FP8 on SM121
- SYMPTOM: switching an inference workload from FP8 to NVFP4 on Spark makes it slower, not faster.
- CAUSE: SM121 lacks
cvt.e2m1x2unless kernels targetsm_121a; NVFP4 runs ~32% slower without it. - CHECK:
references/gotcha-checks.mdG7 — capability reports(12, 1); does the build targetsm_121a? - FIX: stay on FP8 unless the build targets
sm_121a.
G8: Stale Official Playbooks
- SYMPTOM: following an official DGX Spark playbook still fails, with no local misconfiguration explaining it.
- CAUSE: official playbooks have shipped broken before; the stack moves faster than the docs.
- CHECK:
references/gotcha-checks.mdG8 — the playbook repo's recent issues. - FIX: check
github.com/NVIDIA/dgx-spark-playbooksissues before trusting a recipe for an expensive run.
G9: Container-First, Not Bare Pip
- SYMPTOM: a bare-pip environment that worked yesterday
breaks after an unrelated
pip install, or two "identical" environments behave differently. - CAUSE: bare pip lets Triton, xformers, and transformers drift independently; nothing pins them to GB10's SM121 target.
- CHECK:
references/gotcha-checks.mdG9 — container or bare pip? - FIX: prefer an NGC container (see
spark-environment-setupfor tag guidance) or Unsloth's container. If bare pip is unavoidable, follow the NVIDIA install order, including--no-depson Unsloth.
G10: Dual-Spark Is DDP/FSDP Only
- SYMPTOM: a tensor-parallel launch across two Sparks hangs, runs far slower than single-Spark, or errors out.
- CAUSE: ConnectX-7 is fast enough for gradient/parameter sync (DDP, FSDP) but too thin for TP's fine-grained traffic.
- CHECK:
references/gotcha-checks.mdG10 — the configured parallelism strategy. - FIX: on a two-Spark setup, choose DDP or FSDP, never tensor parallelism — TP is single-node only here.
Fast Triage
The cheapest checks to run before anything else:
python3 -c "import torch; print(torch.version.cuda)" # expect 13.x (G1); NGC builds have no +cu130 tag — that's not a failure
import torch; print(torch.cuda.get_device_capability()) # expect (12, 1) (G7)
{ [ -f /.dockerenv -o -f /run/.containerenv ] || grep -qE 'docker|containerd' /proc/1/cgroup; } 2>/dev/null && echo container || echo unknown # G9
assets/preflight.sh runs G1, G3, G4, G7, G9 and produces one
output line per gotcha in a fixed format: G-number first, then
PASS/FAIL/WARN where automatable, SKIP when unavailable, or
INFO: for a raw reading (G3, G4). Full commands:
references/gotcha-checks.md. See also
spark-environment-setup for the environment assumed working.
Related skills
More from wshobson/agents and the wider catalog.

sql-optimization-patterns
Master SQL query optimization, indexing strategies, and EXPLAIN analysis to eliminate slow queries.

startup-financial-modeling
Build 3-5 year financial models with revenue projections, cost structures, cash flow analysis, and scenario planning for early-stage startups.

startup-metrics-framework
Track and optimize key startup metrics from seed through Series A across SaaS, marketplace, consumer, and B2B models.

stride-analysis-patterns
Apply STRIDE methodology to systematically identify threats across authentication, integrity, confidentiality, availability, and authorization.

stripe-integration
Implement Stripe payment processing with checkout, subscriptions, and webhooks for PCI-compliant payment flows.

tailwind-design-system
Build scalable design systems with Tailwind CSS v4, design tokens, and responsive component patterns.