PluginBench
Skill
Review
Audit score 70

explore-run

lllllllama/rigorpilot-skills

Plan and execute bounded exploratory runs for deep learning research with fair-comparison tracking.

What is explore-run?

explore-run is a leaf skill for authorized exploratory evidence in deep learning repositories. Use it for small-subset validation, batch sweeps, short-cycle probes, and transfer-learning trials when the researcher explicitly requests exploratory work—not for trusted baseline execution or SOTA claims.

  • Plan candidate runs using cost, success_rate, and expected_gain scoring
  • Execute small-subset and short-cycle exploratory checks before heavier runs
  • Isolate experiment state from trusted baseline to preserve reproducibility
  • Generate fair-comparison caveats and no-overclaim summaries in explore_outputs/
  • Rank executed runs by real evidence rather than heuristic prediction
  • Hand off execution to minimal-run-and-audit or run-train as needed

How to install explore-run

npx skills add https://github.com/lllllllama/rigorpilot-skills --skill explore-run
Prerequisites
  • Explicit researcher authorization for exploratory work
  • Access to the deep learning repository with training infrastructure
  • Optional: variant specification using variant_axes, subset_sizes, and short_run_steps
Claude Code
Cursor
Windsurf
Cline

How to use explore-run

  1. 1.Define exploratory task scope and obtain explicit researcher authorization
  2. 2.Specify candidate dimensions using variant_axes or accept defaults
  3. 3.Set subset_sizes and short_run_steps to bound exploratory scale
  4. 4.Optionally provide selection_weights to rebalance cost, success_rate, and expected_gain
  5. 5.Run the skill to generate candidate ranking and execution plan
  6. 6.Review explore_outputs/ for CHANGESET.md, COMPARABILITY_REPORT.md, and TOP_RUNS.md
  7. 7.Execute top candidates via minimal-run-and-audit or run-train
  8. 8.Use real execution results to update ranking rather than relying on heuristics

Use cases

Good for
  • Validate hyperparameter changes on a small data subset before full training
  • Run batch sweeps across model variants to identify promising directions
  • Perform quick transfer-learning trials on idle GPU resources
  • Conduct short-cycle guess-and-check probes to guide next steps
  • Compare candidate approaches with explicit bounded-evidence labeling
Who it's for
  • Deep learning researchers conducting exploratory validation
  • ML engineers testing hypotheses before committing to full training runs
  • Teams needing to rank candidate experiments by real execution data
  • Researchers who want to preserve trusted baselines while exploring variants

explore-run FAQ

When should I use explore-run vs. run-train?

Use explore-run when the researcher explicitly authorizes bounded exploratory work (small subsets, batch sweeps, short cycles). Use run-train for trusted baseline execution and conservative verification.

Does explore-run execute training directly?

No; explore-run plans and ranks candidates, then hands off actual execution to minimal-run-and-audit or run-train while keeping experiment state isolated from the trusted baseline.

How does ranking work?

Pre-execution ranking uses cost, success_rate, and expected_gain with conservative default weights. After execution, ranking switches to real evidence. You can customize weights via selection_weights.

What output files should I expect?

explore_outputs/ contains CHANGESET.md, SCIENTIFIC_CHANGELOG.md, COMPARABILITY_REPORT.md, TOP_RUNS.md, and status.json with fair-comparison caveats and no-overclaim summaries.

Can I use this for SOTA claims?

No; this skill is for bounded exploratory evidence only. Do not use it for verified SOTA claims or end-to-end orchestration on top of current_research.

Full instructions (SKILL.md)

Source of truth, from lllllllama/rigorpilot-skills.


name: explore-run description: Rigor Improve / Rigor Explore run leaf skill for bounded exploratory evidence in deep learning research repositories. Use when the researcher explicitly authorizes exploratory runs such as small-subset validation, short-cycle guess-and-check, batch sweeps, idle-GPU search, or quick transfer-learning trials, with fair-comparison caveats and no-overclaim summaries in explore_outputs/. Do not use for end-to-end exploration orchestration on top of current_research, trusted baseline execution, conservative training verification, default routing, verified SOTA claims, or implicit experimentation.

explore-run

Use this as the Rigor Improve / Rigor Explore run leaf skill. The installed slug remains explore-run for compatibility.

Use the shared operating principles in ../ai-research-reproduction/references/agent-operating-principles.md; this skill should guide candidate run planning while preserving model judgment about the active repo.

When to apply

  • When the researcher explicitly authorizes exploratory runs.
  • When the task is a small-subset validation, short-cycle training probe, batch sweep, idle-GPU search, or quick transfer-learning trial.
  • When the output should rank candidate runs rather than certify trusted success.

When not to apply

  • When the user wants trusted training execution or conservative verification.
  • When there is no explicit exploratory authorization.
  • When the task is repository setup, intake, or debugging.

Clear boundaries

  • This skill owns exploratory execution planning and summary only.
  • Use ai-research-explore instead when the task spans both current_research coordination and exploratory code changes.
  • It may hand off actual command execution to minimal-run-and-audit or run-train.
  • It should keep experiment state isolated from the trusted baseline.
  • It should prefer small-subset and short-cycle checks before heavier exploratory runs.
  • It should label run results as bounded evidence and explain when a comparison is not directly fair.

Ranking Semantics

  • Pre-execution candidate selection uses three factors: cost, success_rate, and expected_gain.
  • Default weights should stay conservative unless the researcher explicitly provides selection_weights.
  • Budget pruning still applies after scoring through max_variants and max_short_cycle_runs.
  • If runs are executed later, downstream ranking should switch to real execution evidence, not stay purely heuristic.

Variant Spec Hints

  • Use variant_axes to define the candidate dimension grid.
  • Use subset_sizes and short_run_steps to express exploratory run scale.
  • Use selection_weights to rebalance cost, success_rate, and expected_gain.
  • Use primary_metric and metric_goal so downstream ranking can order executed candidates consistently.

Output expectations

  • explore_outputs/CHANGESET.md
  • explore_outputs/SCIENTIFIC_CHANGELOG.md
  • explore_outputs/COMPARABILITY_REPORT.md
  • explore_outputs/TOP_RUNS.md
  • explore_outputs/status.json

Notes

Use references/execution-policy.md, ../ai-research-reproduction/references/explore-variant-spec.md, ../ai-research-reproduction/references/deep-learning-experiment-principles.md, scripts/plan_variants.py, and scripts/write_outputs.py.