verifying-before-done
riekelt/principal-engineer
Verify every completion claim by driving the change at its surface and running the verify command before declaring done.
What is verifying-before-done?
This skill encodes verification as the definition of done: before claiming a change is complete, fixed, passing, or shipped, you must drive the changed code at its surface in the running system, run the verify command, and report the actual output faithfully. Use this discipline for every completion claim, regardless of how small or obviously correct the change appears.
- Drive the smallest path that executes changed code in the running system and capture what it produced
- Run the verify command from the pre-change checkpoint and paste its actual output as proof
- Report results faithfully in both directions: failing tests with output, skipped steps as skipped, verified work plainly
- Distrust green suites by confirming tests actually executed and can fail when the code breaks
- Distinguish implementation (in repo) from deployment (live) from external verification (in external system)
- Sweep the diff before declaring complete: no unrelated changes, no planning residue, docs updated, all acceptance criteria met
How to install verifying-before-done
npx skills add https://github.com/riekelt/principal-engineer --skill verifying-before-done- The principal-engineering skill must be installed first
- Access to run the verify command from the pre-change checkpoint
- Ability to drive the changed code at its surface in a running system
How to use verifying-before-done
- 1.Before declaring done, fixed, passing, or shipped, identify the smallest path that executes your changed code in the running system
- 2.Run that path and capture what the system produced (response body, terminal output, rendered screen)
- 3.Run the verify command from the pre-change checkpoint and paste its actual output
- 4.Report results faithfully: include failing test output, name any skipped steps, state verified work plainly without hedging
- 5.Confirm tests actually executed by running them alone and breaking the code to watch them fail, then unbreak it
- 6.Sweep the diff for unrelated changes, planning residue, missing documentation updates, and unmet acceptance criteria
- 7.For critical changes, ensure an independent verifier (not the author) performed the verification, or name the weaker substitute used
Use cases
- Before claiming a bug fix is complete, run the minimal path that triggers the fixed error and show the corrected output
- When a test suite passes but the change seems too small for the result, run the test alone and break the code to confirm it can fail
- Before shipping a config or migration change, verify it in an environment matching production rather than relying on local test results
- When independent verification is unavailable for a critical change, name the weaker substitute verification you used instead of silently downgrading
- After implementing a feature flag, drive it with the flag set and probe edge cases like empty values and conflicting options
- Principal engineers and senior developers responsible for critical changes
- Teams working on money, sales, stored data, safety, or other systems that must never get wrong
- Code reviewers verifying changes before merge
- Anyone shipping irreversible migrations or top-tier changes requiring independent verification
verifying-before-done FAQ
Run the smallest path that executes the changed code in the running system where its caller meets it: the command line that runs it, the endpoint that serves it, the screen that renders it, or the consumer that imports it. Do not verify internal functions in isolation; observe where they are called.
A green suite over code that cannot work means the suite did not run, did not cover the path, or cannot fail. Confirm the test executed by running it alone and confirm it can fail by breaking the code and watching it go red.
Attribute the origin first, then fix the test regardless of whose it is. 'Pre-existing' is a footnote in the report, never an excuse. The only exception is a pre-existing red on main that blocks an unrelated green fix, which should be surfaced with an offer to merge the green fix anyway.
Say 'no runtime surface' as the verification for changes that produce no behavior. Never run substitute gate commands to fill the space.
Implementation means the code is in the repo; deployment means it is live; external verification means it has been checked in the external system. Never claim a later tier from evidence of an earlier one.
Full instructions (SKILL.md)
Source of truth, from riekelt/principal-engineer.
name: verifying-before-done description: "Use when about to say "done", "fixed", "passing", or "shipped", or when reporting the outcome of any change. Encodes verification as the definition of done: drive the change at its surface, run the verify command, report faithfully, distrust green suites, own failing gates. Use before every completion claim, even when the change was small and obviously correct, which is when this is skipped."
Verifying before done
REQUIRED BACKGROUND: the principal-engineering skill.
Overview
Done means verified, and verified names what was checked. Careful work is not a check.
The discipline
- Drive the change at its surface, capture what it did. Run the smallest path that executes the changed code in the running system; "Verify at the surface" below expands this.
- Run the verify command, paste its output. The verify command comes from the pre-change checkpoint; it proves the gates hold, and its actual output backs that part of the claim. Redact secret values from output before quoting it (
operating-safelyowns secrets hygiene). "Verified" is always "PASS (checked X and Y)", never a bare checkmark. - Report faithfully, both directions. Report failing tests with their output, name skipped steps as skipped, and state verified work plainly without hedging. Underclaiming verified work wastes the reader's re-verification exactly like overclaiming wastes their trust.
- Distrust green. A green suite over code that cannot work means the suite does not run, does not cover, or cannot fail. When a result seems too clean for the change's size, confirm the test executed (run it alone, watch it appear) and confirm it can fail (break the code, watch it go red, unbreak it). A test that never ran and a gate that never fires produce confident wrong "done"s.
- Distinguish the tiers. Implemented (in the repo) is not deployed (live) is not externally verified (checked in the external system). Never claim a later tier from evidence of an earlier one.
- Lookback before declaring complete. Sweep the diff: no unrelated changes, no planning residue in code or comments, docs updated in the same change, every acceptance criterion actually met rather than approximately met.
- Independent verification for top-tier changes. The author of a change is the worst-placed person to verify it. For the project's declared critical paths (money, sales, stored data, safety, whatever the system must never get wrong) and for irreversible migrations, the verifier is someone or something that did not write the code. When no independent verifier is reachable in time, use the nearest substitute and name it as the weaker form it is; downgrading the check silently is the failure, downgrading it visibly is a decision.
Verify at the surface
The gates (the verify command, the suite, the build) prove the repository holds; the change is proven where its caller meets it: the command line that runs it, the endpoint that serves it, the screen that renders it, or the consumer that imports it.
- Drive the smallest path that executes the changed code, in the running system. Run a changed flag with the flag set, send a changed handler its request, trigger the error of a changed error path. An internal function is not a surface: its caller ends at a surface, so observe there.
- Read tests as the author's evidence, not the verification. A test says what to drive; the gates re-run it.
- Probe beside the change. The happy path confirms the claim; the neighbors test it: the empty value, the repeated call, the conflicting option, the adjacent error the change did not touch. One probe past the claim is the minimum; a probe that holds is still reported, because it says what was covered.
- Quote what the system produced. The response body, the terminal output, and the rendered screen back the claim. Redact before quoting: tokens, connection strings, and auth headers. Treat ambiguous output as a failure with the redacted raw capture attached, never interpret it into a pass.
- Say "no runtime surface" when none exists. Docs, comments, and type declarations that produce no behavior get that as the verification, never a substitute gate run to fill the space.
- Never drive destructive paths live. Where the changed code deletes or writes beyond the workspace and no safe target exists, verify around it and name the unexercised path (
operating-safelyowns the guards).
Failing gates you own
Attribute the origin first, then fix the test regardless of whose it is. "Pre-existing" is a footnote in the report, never an excuse in the gate. The one exception is procedural: a pre-existing red on the main branch that blocks an unrelated green fix gets surfaced with an offer to merge the green fix anyway, decided by the operator.
Common mistakes
- Declaring done from the diff looking right. The diff is the hypothesis; the run at the surface is the experiment.
- Running the whole suite instead of the targeted verify command, and reading "no new failures" as "my change works". A suite that never covered the path cannot vouch for it.
- Re-running the gates and calling it verification. Green gates plus an undriven surface claim "CI works", not "the change works".
- Verifying the happy path of a change whose risk is in the failure path.
- "Tests pass locally" as the terminal claim for a change whose risk is environmental (config, migrations, permissions, prod data shape).
- Fixing the test instead of the code when red is inconvenient. The test was the messenger.
Related skills
More from riekelt/principal-engineer and the wider catalog.

writing-unit-tests
Behavior-first unit testing: one claim per test, deterministic setup, mocks only at external boundaries.

adding-dependencies
Systematic framework for adding, vetting, updating, and removing dependencies with cost-aware decision-making.

grounding-before-coding
Map real code and data before writing—ground every claim in evidence before any change.

guarding-architecture
Encode structural invariants as named, enforced contracts to prevent architectural decay.

diagramming-processes
Diagram business processes, workflows, and system interactions as maintainable source code.

documenting-contracts
Document HTTP APIs, message contracts, and file formats with exhaustive wire-level detail at the right abstraction level.