PluginBench
Skill
Review
Audit score 70

run-train

lllllllama/rigorpilot-skills

Execute and document deep learning training runs with standardized evidence capture for reproducibility.

What is run-train?

The run-train skill executes a selected training command in deep learning research repositories and captures structured evidence (logs, checkpoints, configs, metrics, status) to a standardized `train_outputs/` directory. Use it when you have a specific training command ready and need conservative execution with full reproducibility tracking—not for environment setup, exploratory sweeps, or orchestration.

  • Executes a selected training command with startup, short-run, full, or resume modes
  • Captures training logs, checkpoints, configurations, and random seeds to standardized outputs
  • Generates structured reports: SUMMARY.md, COMMANDS.md, LOG.md, SCIENTIFIC_CHANGELOG.md, COMPARABILITY_REPORT.md, and status.json
  • Records partial, blocked, resumed, and completed training states clearly
  • Preserves reproducibility context including runtime assumptions and metric evidence

How to install run-train

npx skills add https://github.com/lllllllama/rigorpilot-skills --skill run-train
Prerequisites
  • A selected, runnable training command ready to execute
  • Environment and assets already set up (this skill does not handle setup)
  • Access to the repository with training scripts and configuration files
Claude Code
Cursor
Windsurf
Cline

How to use run-train

  1. 1.Identify the training command to run and the desired mode (startup verification, short-run, full kickoff, or resume)
  2. 2.Provide the command, relevant config file paths, and any seed or checkpoint information
  3. 3.Run the skill with the selected training goal and mode
  4. 4.Monitor the execution and review the generated `train_outputs/` directory for logs, status, and metrics
  5. 5.Check SUMMARY.md and COMPARABILITY_REPORT.md to verify reproducibility and results

Use cases

Good for
  • Verify a training setup works before committing to a full run
  • Resume an interrupted training job from the last checkpoint
  • Execute a documented training command and capture all evidence for later analysis or reproduction
  • Run a short verification pass to confirm hyperparameters and data pipeline before full training
  • Document the exact command, environment, and results for a research paper or reproducibility report
Who it's for
  • Deep learning researchers running experiments in established repositories
  • ML engineers verifying training pipelines before production deployment
  • Teams needing standardized training evidence for reproducibility and collaboration
  • Researchers documenting experimental runs for publication or peer review

run-train FAQ

When should I use run-train vs. just running the command manually?

Use run-train when you need standardized evidence capture, reproducibility tracking, and structured reporting. It's especially valuable for research runs that must be documented, resumed, or compared later.

Can this skill set up my environment or download datasets?

No. run-train assumes your environment and assets are already in place. Use it only after setup is complete and your training command is ready to execute.

What happens if the training run is interrupted?

run-train records the partial state and can resume from the last checkpoint. The status.json and logs will show the interrupted state clearly.

Does this skill choose what to train or explore different hyperparameters?

No. You must select the training command and goal beforehand. run-train executes that specific command conservatively and captures evidence; it does not own exploratory branching or sweeps.

What output files should I expect?

You'll get SUMMARY.md (overview), COMMANDS.md (exact command run), LOG.md (training logs), SCIENTIFIC_CHANGELOG.md (changes), COMPARABILITY_REPORT.md (reproducibility details), and status.json (structured status).

Full instructions (SKILL.md)

Source of truth, from lllllllama/rigorpilot-skills.


name: run-train description: Rigor Train skill for deep learning research repositories. Use when a documented or selected training command should be run conservatively for startup verification, short-run verification, full kickoff, or resume, with command, config, seed, log, checkpoint, status, and metric evidence written to standardized train_outputs/. Do not use for environment setup, exploratory sweeps, speculative idea implementation, or end-to-end orchestration.

run-train

Use this as the Rigor Train skill. The installed slug remains run-train for compatibility.

Use the shared operating principles in ../ai-research-reproduction/references/agent-operating-principles.md; this skill should keep training evidence bounded while leaving repository-specific monitoring details to the model.

When to apply

  • When the training command has already been selected and should be executed conservatively.
  • When the researcher wants startup verification, short-run verification, full training kickoff, or resume handling.
  • When the run needs structured training status, checkpoint, and metric reporting.

When not to apply

  • When the main task is environment setup or asset download.
  • When the researcher wants inference-only or evaluation-only execution.
  • When the task is speculative exploration, multi-variant sweeps, or autonomous idea implementation.
  • When the user still needs repository intake or paper gap resolution.

Clear boundaries

  • This skill executes a selected training command and normalizes the resulting evidence.
  • It does not choose the overall research goal on its own.
  • It does not own exploratory branching or speculative code adaptation.
  • It should record partial, blocked, resumed, and kicked-off states clearly.
  • It should preserve reproducibility context such as configs, seeds, checkpoints, logs, metrics, and runtime assumptions when available.

Input expectations

  • selected training goal
  • runnable training command
  • environment and asset assumptions
  • run mode such as startup verification, short-run verification, full kickoff, or resume

Output expectations

  • train_outputs/SUMMARY.md
  • train_outputs/COMMANDS.md
  • train_outputs/LOG.md
  • train_outputs/SCIENTIFIC_CHANGELOG.md
  • train_outputs/COMPARABILITY_REPORT.md
  • train_outputs/status.json

Notes

Use references/training-policy.md, ../ai-research-reproduction/references/deep-learning-experiment-principles.md, scripts/run_training.py, and scripts/write_outputs.py.