create-verification-skill
cursor/plugins
Generate a project-local verification skill that automates testing your app the way a user does.
What is create-verification-skill?
This skill creates a scripted verification harness tailored to your repository that launches your app, exercises features programmatically, and captures evidence of correct behavior. Use it when your project lacks automated tests or needs a user-centric verification approach across any language, framework, or platform.
- Interviews your codebase to identify the app's user-facing surface (web UI, CLI, API, desktop, mobile, or library)
- Generates a project-local skill with launch, doctor, drive, evidence, and cleanup sections tailored to your repo
- Creates a feature map documenting top user-facing features and how to verify each one
- Produces harness recipes using existing test infrastructure (Playwright, Cypress, PTY helpers, HTTP) or generic approaches
- Captures observable proof: screenshots, terminal transcripts, response bodies, logs, exit codes, and database state
How to install create-verification-skill
npx skills add https://github.com/cursor/plugins --skill create-verification-skill- A working local checkout of the project that builds and starts as documented
- Knowledge of how the app is launched (dev command, Makefile, README quickstart)
- Ability to identify the primary user-facing surface (web UI, CLI, API, etc.)
How to use create-verification-skill
- 1.Run the skill and let it interview your codebase to identify the app's surface, startup command, and interaction points
- 2.Review the generated `.cursor/skills/verify-<app>/SKILL.md` with Launch, Doctor, Drive, Evidence, and Cleanup sections
- 3.Review the generated feature map in `.cursor/skills/verify-<app>/features/` documenting top user-facing features
- 4.Execute the generated skill end-to-end once: launch the app, run the doctor check, drive one mapped feature, capture evidence, and clean up
- 5.Verify that evidence artifacts persist after cleanup, then use `/maintain-verification-skill` to keep the map current as the app evolves
Use cases
- Verify a web app's login flow by driving a browser through the UI and capturing screenshots of each step
- Test a CLI tool by launching it in an isolated PTY session and validating output and exit codes
- Confirm an API service responds correctly by making HTTP requests and inspecting response bodies and side effects
- Validate a desktop or Electron app by controlling it via CDP and observing window state and file changes
- Document and verify a library's public API by calling its functions and checking return values and state changes
- Developers building projects without existing test automation
- Teams needing user-centric verification that mirrors real workflows
- Projects spanning multiple languages or frameworks requiring a unified verification approach
- Maintainers who want to document feature behavior for future agents or contributors
create-verification-skill FAQ
The skill will generate a generic harness recipe: browser/CDP for web and Electron apps, a tmux/PTY harness for CLI/TUI, or plain HTTP for services. It adapts to what the codebase already provides.
The skill checks whether two instances can run side by side (different ports, data dirs, profiles). If not, it documents this limitation so agents know not to attempt parallel runs.
Exercise the real user path (not internal setters or test-only endpoints), capture both the action and resulting state, verify side effects (files, database rows, network calls) alongside visible output, and avoid mocks unless the production system already isolates external boundaries.
Fix the build or startup first before generating the skill. If an irrelevant missing asset blocks startup, the skill may create it (clearly marked as verification scaffolding) and remove it during cleanup.
Use `/maintain-verification-skill` to update the feature map and harness as the app evolves. The skill points you to that command and suggests a cadence if you ask.
Full instructions (SKILL.md)
Source of truth, from cursor/plugins.
name: create-verification-skill description: "Generate a project-local verification skill that drives your app the way a user does — any language, framework, or platform. Use for /create-verification-skill, "make a control skill for this repo", or when a project has no scripted way to prove UI/CLI/service behavior." disable-model-invocation: true
Create a verification skill
Every serious project needs a scripted way to drive the real app and prove behavior: launch it, exercise a feature the way a user would, and capture evidence. This skill generates that as a project-local skill (.cursor/skills/verify-<app>/) tailored to the repo. You write the generator's output for the next agent, not for a human: it will be read cold, mid-task, by an agent that has never seen the app.
1. Interview the repo, not the user
Answer these from the codebase and only ask the user what you cannot observe:
- Surface: what does a user actually touch? A web UI, a CLI/TUI, a desktop app, an API, a mobile app, a library? A repo can have several; pick the primary one and note the rest.
- Run: how does the app start locally? Prefer the repo's own documented dev command (package scripts, Makefile, README quickstart). Note ports, env vars, seed data, auth.
- Drive: how can an agent interact with it programmatically? Existing harnesses first — Playwright/Cypress specs, expect scripts, PTY helpers, curl-able endpoints, a debug port. Only then pick a generic recipe: browser/CDP for web and Electron, a tmux/PTY harness for CLI/TUI, plain HTTP for services.
- Observe: what evidence can be captured? Screenshots, terminal transcripts, response bodies, logs, exit codes, DB state.
- Isolate: can two instances run side by side (ports, data dirs, profiles)? If not, say so in the generated skill: refusing to double-drive a shared instance beats corrupting the user's session.
If the checkout doesn't build or start as-is, fix that first (or report it precisely) before generating; a skill written against a broken base teaches wrong steps. When an irrelevant missing asset blocks startup (a static dir the API never serves, a sample config), the generated skill may create it, clearly marked as verification scaffolding, and remove it in cleanup.
2. Generate the skill
Write .cursor/skills/verify-<app>/SKILL.md with YAML frontmatter (name: verify-<app> and a description that names the app, the surface, and when to reach for it — without frontmatter the skill never registers) and these sections, each grounded in what the interview actually found (no placeholders left):
- Launch: the exact command that starts the app for verification, and how to tell it's ready (a log line, a port answering, a prompt). Include teardown. For a short-lived CLI or TUI there is no server to keep alive: launch means build the binary (or install deps) once, then start each drive in its own isolated PTY or tmux session.
- Doctor: one read-only check that answers "is this instance worth driving?" — process up, right version/build, port owned by us, auth valid. An agent runs this first whenever anything looks off.
- Drive: the harness recipe with real selectors/commands from this repo, not examples. Prefer stable handles (ARIA labels, data attributes, prompt strings, route paths) over coordinates and tab order.
- Evidence: what to capture for a proof and where it goes. State the proof standards: exercise the real user path, not internal setters or test-only endpoints; capture the action and the resulting state, not just the final screen; verify side effects (files written, rows inserted, messages sent) alongside what's visible; mocks only where a production boundary already isolates the external system. When the safe path is a dry-run or test mode, verify what it actually skips by observing (files, network, git refs) rather than trusting its name: some dry-runs still touch the network or open a browser.
- Cleanup: how to tear down instances the run created. Never kill by process name; kill what you started. Cleanup removes instances and scratch state, never the evidence: proof artifacts survive the teardown, in a location the skill names.
- Helpers: any script the skill ships is executable and its invocation is shown in the skill body. A helper the reader has to reverse-engineer is not a helper.
3. Seed the feature map
Create .cursor/skills/verify-<app>/features/README.md plus one file per user-facing feature you can identify (aim for the top 3-5 to start, from routes, commands, menus, or docs). Follow the shape in references/feature-map-example/, with a README index and one file per feature. Each file answers, from the user's point of view: what the feature is, how to reach it, how to drive it with the harness, and what observable end state proves it works. The four H2s are Sub-features, How to get to it (user POV), Driving it with <harness>, and Gotchas. The map is the repo's maintained verification source; a proof that drives one convenient entry point is incomplete when the map lists others.
4. Prove the generated skill before handing it over
Run its own instructions end to end once: launch, doctor, drive ONE mapped feature (one is enough; the map exists so later runs can cover the rest), capture evidence, clean up. After cleanup, confirm the evidence still exists at the named location — a cleanup that eats the proof fails this step. Fix what fails, and run the generated cleanup after every failed iteration too, so broken attempts don't strand processes and ports. A generated skill that was never executed is a draft, not a deliverable.
5. Offer the maintenance loop
Point the user at /maintain-verification-skill for keeping the map honest as the app changes. Suggest a cadence only if they ask.
Related skills
More from cursor/plugins and the wider catalog.

deslop
Remove AI-generated code slop and clean up code style to match your codebase standards.

figure-it-out
Design auditable playbooks for complex tasks when no narrower skill fits.

fix-ci
Diagnose and fix failing PR checks with iterative, focused solutions.

fix-merge-conflicts
Resolve merge conflicts non-interactively with validation and testing.

get-pr-comments
Fetch and summarize review comments from your active pull request

how
Understand subsystem architecture and code flow before making changes.