PluginBench
Skill
Pass
Audit score 90

figure-it-out

cursor/plugins

Design auditable playbooks for complex tasks when no narrower skill fits.

What is figure-it-out?

Figure-it-out creates a structured, hypothesis-driven workflow for large migrations, multi-part changes, and work requiring human review. It scales rigor to task complexity, runs a scientific method loop, and logs decisions for auditability.

  • Frames the task with falsifiable success criteria, quantified scope, and appropriate rigor level before committing to work
  • Decomposes work into atomic, independently-landable units sequenced by risk and unknowns
  • Builds verification harnesses before implementation to establish baseline measurements
  • Runs a hypothesis loop: state prediction, make minimal change, measure against predicate, keep or revert
  • Logs decision trails via show-me-your-work for human review and audit after stepping away
  • Encodes lessons learned as gates, lint rules, checks, or scripts to prevent recurrence

How to install figure-it-out

npx skills add https://github.com/cursor/plugins --skill figure-it-out
Prerequisites
  • Familiarity with the poteto-mode skill's Principles section
  • Understanding of the prove-it-works, never-block-on-the-human, and show-me-your-work principle skills
Claude Code
Cursor
Windsurf
Cline

How to use figure-it-out

  1. 1.Open a todolist and read the Principles section of poteto-mode skill
  2. 2.Add the five phases (Frame, Design, Run, Log, Verify) as top-level todos
  3. 3.In Phase A, define done as a falsifiable predicate, quantify scope and blockers, and set rigor level
  4. 4.In Phase B, decompose into atomic units, build verification harness first, sequence by risk, and document the phase list for human review
  5. 5.In Phase C, execute each unit as an experiment: state hypothesis, make minimal change, verify against predicate, keep or revert
  6. 6.In Phase D, log the audit trail using show-me-your-work as work lands, not after completion
  7. 7.In Phase E, verify the whole against Phase A predicate and encode lessons as structural improvements

Use cases

Good for
  • Planning and executing a large codebase migration with multiple interdependent phases
  • Designing a multi-part architectural change where decisions need human sign-off at checkpoints
  • Implementing a complex feature requiring reversibility and rollback capability
  • Coordinating parallel work across team members with clear verification gates between steps
  • Documenting decision rationale for work that will be reviewed after the implementer steps away
Who it's for
  • Engineering leads planning large-scope changes
  • Individual contributors tackling ambitious multi-phase projects
  • Teams coordinating migrations or architectural refactors
  • Anyone needing to create an auditable record of complex technical decisions

figure-it-out FAQ

When should I use figure-it-out instead of a narrower skill?

Use figure-it-out when the task matches no existing playbook: large migrations, ambitious multi-part changes, or work a human will review after you step away. If a narrower skill exists, use that instead.

What does 'rigor level' mean and how do I choose it?

Rigor level is the gates and artifacts required, not effort. Bias high for one-way doors and high blast radius; use less for reversible low-stakes steps. The Phase A framing should justify your choice.

How do I know if my verification harness is correct?

Build it against the pre-change state so the check reads as 'old value vs new value'. If something passes too easily, suspect the observation method before the system.

What should I do if a unit doesn't verify?

Revert the change and investigate. A verdict is VERIFIED, NOT VERIFIED, or INCONCLUSIVE—inconclusive is not a pass. Don't hide negatives or route around a bad gate; fix the gate itself in its own change.

How do I parallelize work safely?

Decompose across seams only, give each worker its own worktree or branch, and pair delegated work with a judge. Don't over-fan; sequence verification before starting the next unit.

Full instructions (SKILL.md)

Source of truth, from cursor/plugins.


name: figure-it-out description: "Design an auditable playbook when no narrower one fits: a large migration, an ambitious multi-part change, or work a human reviews after stepping away. Scales rigor to the task, runs a hypothesis loop, and logs decisions via show-me-your-work. Use for /figure-it-out, 'figure it out', a large migration, or when no narrower playbook applies." disable-model-invocation: true

Figure it out

When the task matches no playbook, design one. The deliverable before any code is the workflow itself: a sequence of phases that scales rigor to the task, runs the scientific method, and leaves a decision trail a human can audit after stepping away.

Start

Open a todolist whose first item is to read the Principles section of the poteto-mode skill. Then add the phases below as todos.

Phase A: Frame

Ground first, then commit. Don't start the run until you can state:

  • The definition of done as a falsifiable predicate (the prove-it-works principle skill).
  • Scope, quantified: rough units and effort, plus the blockers grounding surfaced.
  • The rigor level, biased high. One-way doors and high blast radius get more. Reversible low-stakes steps get less. Rigor is gates and artifacts, not "try harder".

Present the framing and tradeoffs before committing to a long run. Reversible work proceeds (the never-block-on-the-human principle skill), but a multi-hour run earns one checkpoint.

Phase B: Design the workflow

Decompose into atomic, independently-landable units. Sequence riskiest-unknown-first. Scaffold and verification come before features (the foundational-thinking principle skill).

  • Build the verification harness before the work, with the baseline captured from the pre-change state, so the check reads as "old value vs new value".
  • For one-way-door design decisions, run the architect skill (it runs arena). Skip it for mechanical work whose shape is already concrete. A second arena over a settled design is over-engineering (the laziness-protocol principle skill).
  • Decide what fans out. Parallelize only across seams, and give each worker its own worktree or branch (the separate-before-serializing-shared-state principle skill). Don't over-fan.
  • Write the designed phase list down. That list is what the human reviews.

Then execute the design. Add its steps to the todolist as concrete items, after the Phase C entry and before Phase D. Run each under the Phase C loop discipline, and weave the Phase D log through them, a row as each step lands, rather than saving the whole trail for the end.

Phase C: Run the loop

Each unit is an experiment. State the hypothesis, make the smallest change, measure against the predicate on the real artifact, keep it if it advanced, revert it if it didn't. Apply the sequence-verifiable-units principle skill, verifying each unit before starting the next instead of batching checks at the end.

  • Verify by inspecting the artifact, never a self-report. When something passes too easily, suspect the observation method before the system.
  • Pair delegated work with a judge. If a worker games the gate, reset and harden the contract. If the gate itself is wrong, fix the gate in its own change rather than routing around it.
  • A verdict is VERIFIED, NOT VERIFIED, or INCONCLUSIVE. Inconclusive is not a pass. Don't hide a negative.

Phase D: Keep the audit trail

Log the run via the show-me-your-work skill. figure-it-out's work is usually ambitious enough to commit the trail so the reviewer can read it in the PR. The trail plus the diff is what lets the human come back and trust the work.

Phase E: Verify and hand back

Check the whole against the Phase A predicate on the real product, not just the harness. Encode any recurring correction as a gate, a lint rule, a check, or a script (the encode-lessons-in-structure principle skill).

Reply: the playbook you designed, the rigor level and why, the decision-trail path, what's verified against the predicate, and what's still open.