benchmark
affaan-m/ecc
Measure performance baselines, detect regressions, and compare stack alternatives before/after changes.
What is benchmark?
Benchmark measures real browser metrics (Core Web Vitals, resource sizes), API latency (p50/p95/p99), and build performance (cold build, HMR, test suite). Use it before/after PRs to catch regressions, set performance targets, and validate that your stack meets SLA requirements.
- Measure Core Web Vitals (LCP, CLS, INP, FCP, TTFB) and resource sizes (JS, CSS, images, third-party scripts)
- Benchmark API endpoints with p50/p95/p99 latency under load (100 hits, 10 concurrent requests)
- Track build performance: cold build time, hot reload (HMR), test suite duration, TypeScript checks, lint time
- Compare before/after metrics with delta reporting and verdict (BETTER/WARNING)
- Store baselines in `.ecc/benchmarks/` as JSON, git-tracked for team sharing
How to install benchmark
npx skills add null --skill benchmark- Browser MCP available for page performance mode (Core Web Vitals measurement)
- Network access to target URLs/endpoints for benchmarking
- Baseline saved via `/benchmark baseline` before running `/benchmark compare`
How to use benchmark
- 1.Run `/benchmark baseline` to capture current performance metrics and save to `.ecc/benchmarks/`
- 2.Make your changes (code, dependencies, configuration)
- 3.Run `/benchmark compare` to measure impact against the saved baseline
- 4.Review the delta table (Before/After/Delta/Verdict) for regressions or improvements
- 5.Integrate into CI: run `/benchmark compare` on every PR to catch performance regressions
- 6.Pair with `/canary-watch` for post-deploy monitoring and `/browser-qa` for pre-ship validation
Use cases
- Detect performance regressions in PRs by running `/benchmark compare` against saved baseline
- Measure Core Web Vitals before launch to ensure LCP < 2.5s, CLS < 0.1, INP < 200ms targets are met
- Compare API response times (p95/p99) against SLA targets and identify slow endpoints
- Track development feedback loop (build, HMR, test suite) to catch slowdowns in tooling
- Validate bundle size reductions after code-splitting or dependency optimization
- Backend engineers optimizing API performance and latency
- Frontend engineers measuring Core Web Vitals and bundle size impact
- DevOps/platform engineers tracking build and CI performance
- Tech leads setting performance baselines and SLAs for projects
- Teams running performance checks in CI/CD pipelines
benchmark FAQ
Core Web Vitals: LCP < 2.5s, CLS < 0.1, INP < 200ms, FCP < 1.8s, TTFB < 800ms. Resources: total page < 1MB, JS bundle < 200KB gzipped.
API endpoints are hit 100 times with p50/p95/p99 latency measured, plus 10 concurrent requests to test under load.
Baselines are stored in `.ecc/benchmarks/` as JSON files and are git-tracked so the team shares the same performance targets.
Yes, run `/benchmark compare` on every PR to automatically detect regressions before merge.
Cold build time, hot reload (HMR) time, test suite duration, TypeScript check time, lint time, and Docker build time.
Full instructions (SKILL.md)
Source of truth, from affaan-m/ecc.
name: benchmark description: Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives. metadata: origin: ECC
Benchmark — Performance Baseline & Regression Detection
When to Use
- Before and after a PR to measure performance impact
- Setting up performance baselines for a project
- When users report "it feels slow"
- Before a launch — ensure you meet performance targets
- Comparing your stack against alternatives
How It Works
Mode 1: Page Performance
Measures real browser metrics via browser MCP:
1. Navigate to each target URL
2. Measure Core Web Vitals:
- LCP (Largest Contentful Paint) — target < 2.5s
- CLS (Cumulative Layout Shift) — target < 0.1
- INP (Interaction to Next Paint) — target < 200ms
- FCP (First Contentful Paint) — target < 1.8s
- TTFB (Time to First Byte) — target < 800ms
3. Measure resource sizes:
- Total page weight (target < 1MB)
- JS bundle size (target < 200KB gzipped)
- CSS size
- Image weight
- Third-party script weight
4. Count network requests
5. Check for render-blocking resources
Mode 2: API Performance
Benchmarks API endpoints:
1. Hit each endpoint 100 times
2. Measure: p50, p95, p99 latency
3. Track: response size, status codes
4. Test under load: 10 concurrent requests
5. Compare against SLA targets
Mode 3: Build Performance
Measures development feedback loop:
1. Cold build time
2. Hot reload time (HMR)
3. Test suite duration
4. TypeScript check time
5. Lint time
6. Docker build time
Mode 4: Before/After Comparison
Run before and after a change to measure impact:
/benchmark baseline # saves current metrics
# ... make changes ...
/benchmark compare # compares against baseline
Output:
| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP | 1.2s | 1.4s | +200ms | WARNING: WARN |
| Bundle | 180KB | 175KB | -5KB | ✓ BETTER |
| Build | 12s | 14s | +2s | WARNING: WARN |
Output
Stores baselines in .ecc/benchmarks/ as JSON. Git-tracked so the team shares baselines.
Integration
- CI: run
/benchmark compareon every PR - Pair with
/canary-watchfor post-deploy monitoring - Pair with
/browser-qafor full pre-ship checklist
Related skills
More from affaan-m/ecc and the wider catalog.
benchmark-methodology
Agent skill from affaan-m/ecc.
benchmark-optimization-loop
Systematically benchmark and optimize code performance through measured iteration and variant testing.
blender-motion-state-inspection
Inspect Blender character rigs, poses, and animations with structured facts instead of screenshots alone.
blueprint
Turn complex objectives into step-by-step multi-session engineering plans with dependency graphs and adversarial review.
brand-discovery
Agent skill from affaan-m/ecc.
brand-voice
Extract and reuse a consistent writing voice from real source material across all content and outreach.