PluginBench
Skill
Pass
Audit score 90

convex-launch-readiness

get-convex/agent-skills

Unified backend readiness score with prioritized fix plan—run all Convex audits, dedupe findings, and dispatch fixes.

What is convex-launch-readiness?

Aggregates findings from Convex's authz, reviewer, advisor, and insights audits into a single scored readiness report. Use it to get a reproducible, actionable snapshot of your backend's production readiness before deployment, with an ordered fix plan ranked by severity and confidence.

  • Runs authz, reviewer, advisor, and insights audits concurrently and normalizes their findings into one deduplicated report
  • Computes an auditable readiness score (0–100) based on confirmed findings by severity, with the exact formula printed
  • Deduplicates findings across code and deployment loci so the same defect isn't double-counted
  • Ranks findings by severity (blockers/should-fix/nice-to-have) and confidence (confirmed vs. plausible)
  • Generates an ordered fix plan that dispatches to the appropriate fixer capability for each finding
  • Re-runs affected passes after fixes and shows the score delta to verify improvements

How to install convex-launch-readiness

npx skills add https://github.com/get-convex/agent-skills --skill convex-launch-readiness
Prerequisites
  • A Convex project (convex/ directory) or a deployed Convex deployment
  • Node.js and npm to run the skill
  • Optional: a deployed deployment with traffic for advisor and insights passes; code-only runs are supported but report a code-only score
Claude Code
Cursor
Windsurf
Cline

How to use convex-launch-readiness

  1. 1.Install the skill: npx skills add https://github.com/get-convex/agent-skills --skill convex-launch-readiness
  2. 2.Run the readiness report: invoke the skill to classify your target (local/dev/preview/prod), run all enabled passes, and generate the scored report
  3. 3.Review the report: read the score, severity-grouped findings, and the ordered fix plan; note which passes were skipped and why
  4. 4.Accept fixes: for each finding you want to address, approve the fix dispatch to the appropriate fixer capability (authz, reviewer, migrate, etc.)
  5. 5.Re-run and verify: after fixes are applied, re-run the readiness report to see the score delta and confirm improvements

Use cases

Good for
  • Pre-deployment readiness check: run the full audit suite once to get a single score and prioritized punch list instead of four separate reports
  • Iterative hardening: fix findings in priority order, re-run, and watch the score improve as you address data-loss and authz risks first, then scale, then idiom
  • Code-only assessment: run on local code without a deployed instance to get a code-readiness baseline and identify which passes require live traffic or a deployment
  • Deployment-locus debugging: when advisor or insights flags a live issue (read-limit, OCC, recent failure), trace it back to the code locus and dispatch the appropriate fixer
  • Coverage transparency: see which passes ran, which were skipped and why (no traffic, no deployment, no convex/ dir), so you know what the score does and doesn't cover
Who it's for
  • Backend engineers preparing for production launch or major deployments
  • Teams using Convex who want a single, reproducible readiness metric instead of juggling multiple audit reports
  • DevOps and platform engineers verifying backend health before releasing to users
  • Developers iterating on schema, auth rules, and performance to see real-time score improvements

convex-launch-readiness FAQ

What's the difference between a 'confirmed' and 'plausible' finding?

Confirmed findings have strong evidence (e.g., code audit found a missing index, or live traffic hit a read-limit). Plausible findings are candidates flagged by heuristics but lack definitive proof. Only confirmed findings move the score; plausible ones are listed as candidates for investigation.

Why does my code-only score differ from my deployment score?

A code-only run skips advisor (live read-limit/OCC evidence) and insights (recent failures from logs) because they require a deployed instance with traffic. The report header explicitly states which passes ran and which were skipped, so you know the code-only score measures code readiness, not production-verified readiness.

Can I fix findings one at a time and re-run?

Yes. After each fix, re-run the readiness report to see the score delta. The report dispatches fixes to the appropriate fixer capability and re-runs affected passes, so you can iterate and watch the score improve as you address findings in priority order.

What if a pass errors or is skipped?

The report lists which passes ran, which were skipped and why (e.g., 'no traffic for advisor'), and any errors. A skipped pass is not a clean pass; it's a stated coverage gap. The score reflects only the passes that ran.

How is the readiness score calculated?

Start at 100; subtract per confirmed finding by severity (high −15, medium −5, low −1), floor at 0. The report prints the exact formula and per-class breakdown so the number is reproducible and auditable, not a vibe.

Full instructions (SKILL.md)

Source of truth, from get-convex/agent-skills.


name: convex-launch-readiness description: "Run every Convex audit (authz, reviewer, advisor, insights) into one scored, deduped readiness report with an ordered fix plan — Lighthouse for your backend."

<!-- GENERATED from convex-agents content/capabilities/launch-readiness.json — do not edit by hand. -->

Launch-readiness report

Readiness is not one check — it's the union of the checks, deduped, ranked, and scored. This capability is pure composition over the findings bus (specs/finding.schema.json): it runs each audit capability, normalizes their outputs into one report (specs/finding-report.schema.json), computes an auditable score, and — because every finding names a fixCapability — hands the user a prioritized, actionable punch list instead of four separate reports. It fixes nothing itself; it decides WHAT to fix and in what order, then dispatches to the fixers.

Workflow

  1. GUARD + SCOPE: deploy-guard classifies the target (local-anonymous / dev / preview / prod); announce it. Detect what's assessable — is there a convex/ dir, a deployed deployment with traffic, an auth foundation? Skip passes whose preconditions aren't met and SAY which were skipped (a skipped pass is not a pass).
  2. RUN THE PASSES, each emitting findings on the bus:
    • convex-authz — the authz scan (identity-from-arg, missing ownership, PII leak, parent-ref-on-write). Always runnable on code.
    • convex-reviewer — validators, indexes-not-filter, idiom, error handling. Always runnable on code.
    • convex-advisor — live read-limit / OCC evidence (only if a deployment with traffic exists; else record 'skipped: no traffic').
    • convex-insights — recent failures from logs (only if a deployment exists). Run independent passes concurrently; each returns findings, not fixes.
  3. NORMALIZE + DEDUPE: collect all findings into one report. Set each finding's identity field to a normalized function/table key (e.g. messages:list) that is the SAME whether the pass reported a code-locus or a deployment-locus for that function — so the SAME defect seen from two loci (reviewer flags a missing index at code-locus, advisor flags its read-limit symptom at deployment-locus) collapses to ONE via the bus's (class, identity) dedup and isn't double-counted in the score. Keep the higher-confidence source. Drop nothing silently; a pass that errored/was skipped is a stated coverage gap, not a clean result.
  4. SCORE, auditable: start at 100; subtract per CONFIRMED finding by severity (high −15, med −5, low −1), floor at 0; print the exact formula and the per-class breakdown so the number is reproducible, not a vibe. plausible-only findings are listed as candidates but do NOT move the score (evidence-not-vibes). A deployment/traffic-less run reports a code-only score and says so.
  5. REPORT: the score, then findings ranked by severity, each with its evidence, its locus, and the fixCapability + a one-line fix note. Group by 'blockers' (high) / 'should-fix' (med) / 'nice-to-have' (low). End with the ordered fix plan: which capability to run next, in what order (authz/data-loss first, then perf/scale, then idiom/observability).
  6. DISPATCH on request: for each finding the user accepts, invoke its fixCapability (convex-authz, convex-reviewer's fixers, migrate-rehearse for schema changes, suggest for component swaps). After fixes, RE-RUN the affected passes and show the score delta — the readiness number is only meaningful if it moves when you fix things.
  7. Never claim more coverage than was run: the report header lists which passes ran, which were skipped and why. A green score on a code-only run is 'code looks ready', not 'production-verified'.

Rules

  • Compose, don't re-implement: run the existing audit capabilities and aggregate their bus findings — never re-derive an authz or perf check inline.
  • The score counts CONFIRMED findings only, by severity, with the formula printed; plausible findings are candidates that don't move the number.
  • Normalize each finding's locus to a function/table identity before dedup (map deployment functionId ↔ code file:line) so one defect seen from two loci collapses to one and isn't double-scored; keep the higher-confidence source; drop nothing silently.
  • Every finding carries its fixCapability; the report ends with an ORDERED fix plan (data-loss/authz first, then scale, then idiom/observability).
  • Re-run affected passes after fixes and show the score delta — a readiness number that doesn't move when you fix things is theater.
  • Never claim more than was run: header lists ran/skipped passes; a code-only run yields a code-only score, explicitly labeled.
  • This is a read + aggregate + dispatch pass; fixes happen in the fixer capabilities, gated by their own consent/deploy-target rules.