PluginBench
Skill
Pass
Audit score 90

safe-debug

lllllllama/rigorpilot-skills

Conservative diagnosis and minimal patching for deep learning training failures and errors.

What is safe-debug?

safe-debug is a rigor-focused debugging skill for deep learning research. Use it when you have a concrete error (traceback, CUDA OOM, shape mismatch, NaN loss, checkpoint failure) and want systematic diagnosis with human-approved fixes that are clearly separated from research changes.

  • Diagnoses root causes of training and inference failures from tracebacks and error symptoms
  • Proposes minimal, conservative patches with explicit human approval required before changes
  • Separates debug fixes from research contributions to preserve experiment integrity
  • Generates structured debug outputs (DIAGNOSIS.md, PATCH_PLAN.md, status.json)
  • Escalates savepoint or branch creation for medium- and high-risk changes
  • Follows research rigor principles to avoid introducing confounds

How to install safe-debug

npx skills add https://github.com/lllllllama/rigorpilot-skills --skill safe-debug
Claude Code
Cursor
Windsurf
Cline

How to use safe-debug

  1. 1.Paste the full traceback or describe the concrete failure symptom
  2. 2.Let the skill diagnose and generate DIAGNOSIS.md with root-cause analysis
  3. 3.Review the proposed PATCH_PLAN.md with minimal fix suggestions
  4. 4.Approve or request changes before any code modifications are applied
  5. 5.Check status.json for change risk level and savepoint recommendations

Use cases

Good for
  • Debugging CUDA out-of-memory errors during model training
  • Investigating shape mismatches in tensor operations
  • Diagnosing NaN loss or training instability with minimal code changes
  • Resolving checkpoint loading failures before continuing experiments
  • Narrowing root causes of inference failures in production models
Who it's for
  • Deep learning researchers and ML engineers
  • Practitioners running reproducible experiments who need audit trails
  • Teams requiring conservative debugging with explicit change approval
  • Anyone needing to separate bug fixes from research contributions

safe-debug FAQ

When should I use safe-debug vs. general code editing?

Use safe-debug when you have an active, concrete failure (error message, traceback, or symptom) and want conservative diagnosis with human approval gates. Use general editing for refactoring, exploration, or tasks without a specific error.

Does safe-debug automatically fix my code?

No. It diagnoses, proposes minimal patches, and requires explicit human approval before making any changes to your repository.

How does safe-debug handle research integrity?

It explicitly flags whether a fix changes experiment meaning or comparability, and separates debug patches from research contributions in the output.

What if the patch is risky?

The skill escalates medium- and high-risk changes, recommending savepoint or branch creation before proceeding.

What output files does safe-debug create?

It generates debug_outputs/DIAGNOSIS.md (root-cause analysis), debug_outputs/PATCH_PLAN.md (minimal fix proposals), and debug_outputs/status.json (change metadata).

Full instructions (SKILL.md)

Source of truth, from lllllllama/rigorpilot-skills.


name: safe-debug description: Rigor Debug / Rigor Audit skill for deep learning research work. Use when the user pastes a traceback, terminal error, CUDA OOM, checkpoint load failure, shape mismatch, NaN loss symptom, or training failure and wants conservative diagnosis before any patching, with debug fixes clearly separated from research contributions. Do not use for broad refactoring, speculative adaptation, automatic exploratory patching, or general repository familiarization.

safe-debug

Use this as the Rigor Debug / Rigor Audit skill. The installed slug remains safe-debug for compatibility.

Use the shared operating principles in ../ai-research-reproduction/references/agent-operating-principles.md; this skill should guide conservative diagnosis without blocking the model from finding the local root cause.

When to apply

  • The user provides a traceback, terminal error, or concrete training or inference failure symptom.
  • The user wants diagnosis, root-cause narrowing, and minimal patch suggestions before code is changed.
  • The user wants a safe debug flow with explicit human approval before mutation.

When not to apply

  • When the user wants a broad repository walkthrough without an active failure.
  • When the task is speculative experimentation or code adaptation.
  • When the user is asking for a large refactor or readability rewrite.

Clear boundaries

  • Diagnose first.
  • Do not modify repository code by default.
  • If a patch is needed, propose the smallest fix and require explicit approval first.
  • Escalate savepoint or branch creation before medium-risk or high-risk changes.
  • A debug fix is not automatically a research contribution; if it changes experiment meaning or comparability, say so explicitly.

Output expectations

  • debug_outputs/DIAGNOSIS.md
  • debug_outputs/PATCH_PLAN.md
  • debug_outputs/status.json

Notes

Use references/debug-policy.md, ../ai-research-reproduction/references/research-rigor-principles.md, and the shared ../ai-research-reproduction/references/research-pitfall-checklist.md.