PluginBench
Skill
Pass
Audit score 90

principle-prove-it-works

cursor/plugins

Verify task completion by checking real artifacts, not proxies or self-reports.

What is principle-prove-it-works?

A verification discipline that requires direct observation of actual outputs before declaring work done. Apply this after completing any task to confirm the feature runs, values are correct, and diffs match expectations—not inferred from compilation status, cached state, or agent claims.

  • Check process liveness and behavior directly, not through derived state or logs
  • Read actual values and outputs instead of cached or inferred representations
  • Inspect real diffs and file changes to confirm modifications
  • Write deterministic verification scripts that can be re-run and audited
  • Distinguish between observation method failures and actual system failures

How to install principle-prove-it-works

npx skills add https://github.com/cursor/plugins --skill principle-prove-it-works
Claude Code
Cursor
Windsurf
Cline

How to use principle-prove-it-works

  1. 1.After completing a task, identify the real artifact (running process, output file, API response, etc.)
  2. 2.Directly observe or test the artifact instead of relying on logs, compilation status, or agent self-reports
  3. 3.When possible, write a deterministic script that reproduces the verification check
  4. 4.Run the script and keep its output visible for review
  5. 5.If verification fails, check your observation method before assuming the system is broken

Use cases

Good for
  • After implementing a feature, run it end-to-end to confirm it works as intended
  • Before merging code, verify the actual diff matches the intended changes
  • When a task claims success, independently check the real artifact instead of trusting the report
  • Create reusable verification scripts for complex changes that need auditable proof
  • Catch silent failures where compilation succeeds but runtime behavior is wrong
Who it's for
  • Developers completing features or fixes
  • Code reviewers verifying pull requests
  • Teams requiring auditable work trails for migrations or large refactors
  • Anyone working in environments where unverified changes cause downstream costs

principle-prove-it-works FAQ

What counts as 'the real artifact'?

The actual running process, output value, file content, API response, or user-visible result—whatever the task was supposed to produce. Not logs, mtimes, compilation output, or cached representations.

Should I commit verification scripts?

For small tasks, keep the script output visible but don't commit it. For large work like migrations or ports where an auditable trail is needed, commit the script so reviewers can re-run it.

What if I can't directly observe the artifact?

Write a deterministic script that captures the observable state and can be re-run. This is stronger than a one-time eyeball check and gives reviewers something to verify against.

How does this differ from testing?

Testing checks expected behavior against a spec. This checks that the real output matches what you actually see, catching cases where the spec was wrong or the implementation diverged silently.

What if verification reveals a failure?

First suspect your observation method—wrong environment, stale cache, wrong file, wrong process. Only after ruling that out should you suspect the system itself.

Full instructions (SKILL.md)

Source of truth, from cursor/plugins.


name: principle-prove-it-works description: "Apply after completing a task, before declaring done. Verify against the real artifact (run the feature, read the actual value, inspect the diff), not a proxy, self-report, or 'it compiles.'" disable-model-invocation: true

Prove It Works

Verify every task output by checking the real thing directly. Do not infer from proxies, self-reports, or "it compiles."

Why: Unverified work has unknown correctness. Indirect verification (file mtimes, output freshness, agent self-reports, cached screenshots) feels cheaper than direct observation. Acting on a wrong inference costs far more than checking the source.

Check the real thing, not a proxy:

  • Check process liveness directly, not indirectly through derived state
  • Read the actual value, not a cached or derived representation
  • When verification fails, suspect the observation method before suspecting the system

Script the check when you can

The strongest proof is a deterministic script that re-runs the same comparison, not a one-time eyeball. Write the script, run it, and keep its output as an artifact a reviewer can re-run instead of trusting your word.

Keep the artifact visible for the human. Commit it only for large or complex work where the trail has to be auditable later, like a big port or migration (the show-me-your-work skill).