PluginBench
Skill
Pass
Audit score 90

build-loop-codex

buildgreatproducts/builder-os

Disciplined build→review→test→fix loop for feature work with OpenAI Codex CLI.

What is build-loop-codex?

Automates quality-gated feature implementation by running code through Codex's review tool, end-to-end testing, and iterative fixes until all plan tasks pass. Use when you need structured, verified feature delivery rather than "it compiles" code.

  • Executes tasks from a plan file (or direct feature prompts) in order
  • Runs Codex `/review` on uncommitted changes and fixes all findings
  • Tests features end-to-end and re-runs full test suite to prevent regressions
  • Iterates fix→review→test until each task passes completely
  • Marks plan tasks complete and reports progress with verification details
  • Catches security issues, edge cases, and design-system violations before shipping

How to install build-loop-codex

npx skills add https://github.com/buildgreatproducts/builder-os --skill build-loop-codex
Prerequisites
  • OpenAI Codex CLI installed and configured
  • A plan file (roadmap, task list, or refactor plan with checkboxes) or a clear feature prompt
  • Existing test suite and app that can be run locally
Claude Code
Cursor
Windsurf
Cline

How to use build-loop-codex

  1. 1.Trigger the skill with "run the build loop", "build the next task", "continue the plan", or "build this feature properly"
  2. 2.If no plan exists, restate the feature as a goal with 2–4 success criteria and confirm scope
  3. 3.The skill builds the first unchecked task (or your prompt), runs Codex `/review`, fixes all findings, tests end-to-end, and iterates until passing
  4. 4.Review findings on security-sensitive code (auth, payments, input) get a second pass
  5. 5.Mark tasks complete as they pass; the skill loops to the next task until all requested work is done
  6. 6.Review the final report: what was built, findings fixed, how it was verified, and what needs your attention

Use cases

Good for
  • Building features from a roadmap or task-list plan file with checkbox tracking
  • Implementing a feature prompt with verifiable success criteria when no plan exists
  • Refactoring work that requires review and test gates to prevent breakage
  • Ensuring new code passes security review (auth, payments, user input) before merge
  • Walking through real user flows to verify loading, empty, and error states work
Who it's for
  • Teams using OpenAI Codex CLI and wanting automated quality gates
  • Developers building from structured plans or roadmaps
  • Projects requiring security review on sensitive code changes
  • Anyone prioritizing verified, tested work over rapid iteration

build-loop-codex FAQ

What if a task seems wrong or unclear?

Ask one specific question rather than guessing. Don't relitigate plan decisions or silently expand scope.

What counts as 'end-to-end testing'?

Run the task's verification step, the full test suite (everything that passed before must still pass), and walk the real user flow including empty, loading, and error states.

What if Codex review finds something that contradicts the task spec?

The spec wins. Flag the disagreement in your report and fix according to spec, not the review finding.

Can I skip review or testing to move faster?

No. Skipped review or untested work is unfinished work. The loop doesn't advance until every step passes.

What if the project has a design system or theme config?

Check UI changes against it. No hardcoded colors, type, or spacing that bypass design tokens.

Full instructions (SKILL.md)

Source of truth, from buildgreatproducts/builder-os.


name: build-loop-codex description: Use when building features with Codex (OpenAI Codex CLI) in any codebase and the work should go through a disciplined build → review → test → fix loop. Triggers on "run the build loop", "build the next task", "continue the plan", "build this feature properly", or any request to implement work from a plan file or a direct feature prompt. Builds from the plan (or the prompt if no plan exists), runs Codex's /review on uncommitted changes and fixes every issue found, tests and verifies the feature end to end, fixes anything testing surfaces, and reports back once complete. Repeats until all plan tasks are checked off. license: MIT metadata: author: BuilderOS version: "1.0"

Codex Build Loop

Quality-gated feature work: nothing ships on "it compiles" — every increment is built, reviewed, tested end to end, and fixed before the user hears "done."

Source of work

  • A plan file exists (roadmap, refactor plan, or task list with - [ ] checkboxes — search the repo): work the first unchecked task. Tasks are ordered intentionally — never skip ahead. If the plan references spec docs, read only the sections relevant to the current task.
  • No plan (or the request is outside it): build from the user's prompt. Restate it as a verifiable goal with 2–4 success criteria and confirm scope in one message before building.

The loop

Run per task (or per prompted feature). Do not advance until every step passes.

  1. Build. Implement exactly what the task specifies. Simplest implementation that satisfies it, surgical changes, no speculative scope. Match existing project conventions.

  2. Review. Run /review and select "Review uncommitted changes". If the change touches auth, payments, user input, or data access, run a second pass via "Custom review instructions" (e.g. "Focus on security vulnerabilities and unvalidated input"). Fix all findings in scope — bugs, security issues, edge cases, performance, style in files you touched. If the project has a design system spec (design tokens file, DESIGN.md, theme config), check UI changes against it — no hardcoded colors, type, or spacing that bypass tokens. Note pre-existing issues in untouched code for the report instead of fixing silently. Re-run /review until clean. If a finding contradicts the task or spec, the spec wins — flag the disagreement.

  3. Test end to end. Run the task's verification step (or the success criteria). Run the full test suite — everything that passed before must still pass. Add tests for new logic. Then exercise the feature as a user would: run the app, walk the real flow including empty, loading, and error states.

  4. Fix. Anything testing finds goes back through the loop: fix → /review → re-test. Never mark a failing task complete; never start the next task with the app broken.

  5. Continue. Mark the task - [x], update any progress/status line in the plan, and loop to the next task until the requested scope is complete.

  6. Report. When done, tell the user: what was built and plan progress, review findings fixed and anything deferred, how it was verified (tests + flow walked), and what needs their attention next. Be honest about anything flaky or partially verified.

Rules

  • Skipped review or untested work = unfinished work.
  • Don't relitigate plan decisions; if a task seems wrong, ask one specific question rather than guessing.
  • Discovered work no task covers? Surface it and propose a task — never silently expand scope.