PluginBench
Skill
Pass
Audit score 90

verify-this

cursor/plugins

Verify claims with repeatable local evidence: baseline vs. treatment comparison with VERIFIED, NOT VERIFIED, or INCONCLUSIVE verdict.

What is verify-this?

Verify-this proves or disproves a specific claim by capturing baseline and treatment artifacts, then comparing them objectively. Use it when you need to confirm a bug fix works, measure performance or UI changes, or validate that a test actually solves the user-visible problem.

  • Restates claims in falsifiable form with measurable conditions and thresholds
  • Captures baseline state from the original code or broken repro
  • Captures treatment state from the changed code with identical test conditions
  • Compares raw artifacts: timings, screenshots, terminal output, HTTP responses, profiles, or heap snapshots
  • Returns one of three verdicts: VERIFIED, NOT VERIFIED, or INCONCLUSIVE
  • Detects confounds and flags inconclusive results when signal is noisy or environment differs

How to install verify-this

npx skills add https://github.com/cursor/plugins --skill verify-this
Claude Code
Cursor
Windsurf
Cline

How to use verify-this

  1. 1.Restate the claim in falsifiable form: specify the condition, metric, and threshold
  2. 2.Identify the smallest local surface that can disprove it (code behavior, CLI output, UI, API response, performance timing, or memory usage)
  3. 3.Capture a baseline artifact from the old state (merge base, parent commit, or current broken repro) using the same command and environment
  4. 4.Capture a treatment artifact from the changed state using identical test conditions, warmup, and data
  5. 5.Compare the raw artifacts side-by-side and note the delta
  6. 6.Return the verdict with evidence and reasoning, naming any confounds

Use cases

Good for
  • Confirm a bug fix actually resolves the reported issue before merging
  • Measure CLI or API performance improvement with before/after timings on the same machine
  • Validate UI changes with screenshot or accessibility snapshot comparison
  • Verify memory leak fix by comparing heap snapshots before and after the suspected operation
  • Prove a test passes and user-visible behavior matches the expected outcome
Who it's for
  • Developers verifying their own fixes and changes
  • Code reviewers who need proof that a claim holds before approval
  • QA engineers confirming performance or behavior claims
  • Teams using test-driven development who want to validate beyond unit tests

verify-this FAQ

What counts as a valid claim?

A falsifiable claim with a measurable condition and threshold: 'response time drops below 100ms', 'error count goes to zero', 'screenshot matches baseline'. Vague claims like 'the code is cleaner' are not valid; ask the user to specify the metric first.

What if the measurement is noisy or the environment differs between baseline and treatment?

Return INCONCLUSIVE and explain why: unstable timing, different machine load, cache state, or other confounds that invalidate the comparison.

Can I store artifacts to disk?

Yes, if safe. Use /tmp/verify-this/<claim-slug>/ with subdirectories for claim.md, baseline/, treatment/, diff/, and verdict.md. If artifacts contain sensitive code, prompts, or HTTP bodies, keep only minimal inline evidence unless the user agrees.

What if the test passes but user-visible behavior is still broken?

This skill is designed for exactly that case. Capture the user-visible behavior (screenshot, CLI output, API response) as the artifact, not just the test result.

Should I soften a negative result?

No. A clear NOT VERIFIED is more useful than hedging. State exactly what changed (or didn't) and by how much relative to the threshold.

Full instructions (SKILL.md)

Source of truth, from cursor/plugins.


name: verify-this description: "Verify a claim with fresh local evidence: restate it falsifiably, capture baseline and treatment, compare artifacts, and return VERIFIED, NOT VERIFIED, or INCONCLUSIVE."

Verify This

Verification is not a recap. It proves or disproves a specific claim with repeatable evidence.

When To Use

  • The user asks "verify this", "prove it works", "did this fix it", or "show me the evidence".
  • A bug fix needs a before/after repro.
  • A UI, CLI, API, performance, or memory claim needs measurement.
  • A test passes but the user-visible behavior still needs confirmation.

Do not use this for vague claims like "the code is cleaner". Ask for a measurable claim first.

Workflow

  1. Restate the claim in falsifiable form: condition, metric, and threshold.
  2. Pick the smallest local surface that can disprove it.
  3. Capture a baseline from the old state: merge base, parent commit, failing branch, or current broken repro.
  4. Capture treatment from the changed state with the same command, data, warmup, and environment.
  5. Compare raw artifacts: numbers, screenshots, terminal transcripts, HTTP responses, profiles, heap snapshots, or test output.
  6. Return exactly one verdict: VERIFIED, NOT VERIFIED, or INCONCLUSIVE.

Local Surfaces

  • Code behavior: focused unit/integration tests or a minimal repro script.
  • CLI/TUI behavior: control-cli, terminal transcript, or demo recording.
  • UI behavior: control-ui, screenshots, accessibility snapshots, or browser traces.
  • API behavior: local HTTP/RPC request and response diff.
  • Performance: same-machine baseline/treatment timings or CPU profiles.
  • Memory: heap snapshots before and after the suspected operation.

Artifact Layout

When safe to write artifacts:

/tmp/verify-this/<claim-slug>/
├── claim.md
├── timeline.md
├── baseline/
├── treatment/
├── diff/
└── verdict.md

If artifacts may contain sensitive code, prompts, screenshots, HTTP bodies, or heap data, keep only the minimal inline evidence unless the user agrees to disk storage.

Verdict Rules

  • VERIFIED: baseline and treatment differ in the predicted direction, by the claimed threshold, with no obvious confound.
  • NOT VERIFIED: the behavior is unchanged, moves the wrong way, or misses the threshold.
  • INCONCLUSIVE: no valid baseline, noisy signal, failed measurement, or an environment difference invalidates the comparison.

Output

Use this shape:

VERIFIED | NOT VERIFIED | INCONCLUSIVE
Claim: <falsifiable claim>

Evidence:
<metric/artifact>: baseline=<...>, treatment=<...>, delta=<...>, threshold=<...>

Reasoning:
<one tight paragraph naming the evidence and any confounds>

Do not soften a negative result. A clear NOT VERIFIED is useful.