minimal-run-and-audit
lllllllama/rigorpilot-skills
Execute and audit deep learning repo smoke tests with standardized evidence capture and scientific changelog tracking.
What is minimal-run-and-audit?
Minimal-run-and-audit is a Rigor Run skill for README-first deep learning repository reproduction. Use it after a reproduction target and setup plan exist to execute documented smoke tests, inference runs, or evaluation commands and generate standardized repro_outputs/ files with patch notes and scientific comparability reports.
- Executes selected smoke tests, inference runs, or evaluation commands with full evidence capture
- Generates standardized repro_outputs/ directory structure with normalized execution results
- Creates SCIENTIFIC_CHANGELOG.md documenting any changes that alter evaluation, preprocessing, or metrics
- Produces COMPARABILITY_REPORT.md comparing against README, paper, and baseline specifications
- Tracks PATCHES.md when repository files are modified during execution
- Distinguishes between verified, partial, and blocked execution states
How to install minimal-run-and-audit
npx skills add https://github.com/lllllllama/rigorpilot-skills --skill minimal-run-and-audit- Selected reproduction target and documented command to execute
- Environment and asset assumptions already defined
- Repository already cloned and initial setup completed
How to use minimal-run-and-audit
- 1.Provide the selected reproduction goal and the exact command to run (smoke test, inference, or evaluation)
- 2.Specify any environment variables, asset paths, or assumptions needed for execution
- 3.Execute the command via the skill and capture all stdout, stderr, and output artifacts
- 4.Review generated repro_outputs/ files, SCIENTIFIC_CHANGELOG.md, and COMPARABILITY_REPORT.md
- 5.Verify that any repository file changes are documented in PATCHES.md with scientific impact notes
Use cases
- Capturing evidence from a documented inference command after environment setup is complete
- Running a smoke test to verify model evaluation pipeline matches paper specifications
- Normalizing outputs from multiple inference runs into standardized repro_outputs/ format
- Documenting and auditing any code changes made during reproduction with scientific impact assessment
- Generating comparability reports to confirm reproduced results align with baseline claims
- ML researchers reproducing deep learning papers
- Research engineers validating model inference and evaluation pipelines
- Teams conducting systematic repository audits and reproducibility assessments
- Scientists needing standardized evidence capture for research rigor documentation
minimal-run-and-audit FAQ
No. This skill is for smoke tests, inference runs, and evaluation commands only. Do not use it for training startup, resume, or long-running training state management.
The skill generates PATCHES.md documenting all changes and flags any modifications that alter scientific meaning (evaluation logic, preprocessing, metrics, checkpoints) in SCIENTIFIC_CHANGELOG.md.
Apply it after a reproduction target and setup plan exist and the user knows what command should be attempted. Do not use during initial repo scanning or when the environment is still undefined.
SCIENTIFIC_CHANGELOG.md tracks changes that alter evaluation or preprocessing meaning and evidence status. COMPARABILITY_REPORT.md compares the reproduced results against README, paper, and baseline specifications.
No. The skill receives the selected reproduction goal and runnable commands as input. It does not choose the overall target or perform broad paper analysis.
Full instructions (SKILL.md)
Source of truth, from lllllllama/rigorpilot-skills.
name: minimal-run-and-audit
description: Rigor Run skill for README-first deep learning repo reproduction. Use when the task is specifically to capture or normalize evidence from the selected smoke test or documented inference or evaluation command and write standardized repro_outputs/ files, including patch notes when repository files changed. Do not use for training execution, initial repo intake, generic environment setup, paper lookup, target selection, hidden scientific-meaning changes, or end-to-end orchestration by itself.
minimal-run-and-audit
Use this as the Rigor Run skill. The installed slug remains
minimal-run-and-audit for compatibility.
Use the shared operating principles in
../ai-research-reproduction/references/agent-operating-principles.md; this skill should make run
evidence auditable without turning every command into a rigid protocol.
When to apply
- After a reproduction target and setup plan exist.
- When the main skill needs execution evidence and normalized outputs.
- When a smoke test, documented inference run, documented evaluation run, or other short non-training verification is appropriate.
- When the user already knows what command should be attempted and wants execution plus reporting only.
When not to apply
- During initial repo scanning.
- When environment or assets are still undefined enough to make execution meaningless.
- When the task is a literature lookup rather than repository execution.
- When the user is still deciding which reproduction target should count as the main run.
Clear boundaries
- This skill owns normalized reporting for an attempted command.
- It may receive execution evidence from the main skill or a thin helper.
- It does not choose the overall target on its own.
- It does not perform broad paper analysis.
- It does not own training startup, resume, or long-running training state.
- It should not normalize risky code edits into acceptable practice.
- It must not hide changes that alter evaluation, preprocessing, checkpoints, metrics, or other scientific meaning.
Input expectations
- selected reproduction goal
- runnable commands or smoke commands
- environment and asset assumptions
- optional patch metadata
Output expectations
- execution result summary
- standardized
repro_outputs/files SCIENTIFIC_CHANGELOG.mdfor changed scientific meaning and evidence statusCOMPARABILITY_REPORT.mdfor README/paper/baseline comparability- clear distinction between verified, partial, and blocked states
PATCHES.mdwhen repo files changed
Notes
Use references/reporting-policy.md, ../ai-research-reproduction/references/research-rigor-principles.md, scripts/run_command.py, and scripts/write_outputs.py.
Related skills
More from lllllllama/rigorpilot-skills and the wider catalog.

paper-context-resolver
Resolve reproduction-critical paper details when README and repo files leave gaps.

repo-intake-and-plan
README-first repository scanner for deep learning reproduction planning.

run-train
Execute and document deep learning training runs with standardized evidence capture for reproducibility.

safe-debug
Conservative diagnosis and minimal patching for deep learning training failures and errors.

drizzle
Drizzle ORM schema and query patterns for PostgreSQL databases in LobeChat.

hotkey
Add or edit LobeHub keyboard shortcuts with proper scoping, conflict detection, and i18n support.