agentic-engineering
affaan-m/ecc
Operate as an agentic engineer using eval-first execution, decomposition, and cost-aware model routing.
What is agentic-engineering?
A framework for engineering workflows where AI agents perform implementation work under human quality and risk control. Use this skill to define completion criteria upfront, decompose work into verifiable units, route tasks to appropriate model tiers, and measure progress with evals and regression checks.
- Define capability and regression evals before execution to establish baselines and measure deltas
- Decompose work into 15-minute units with single dominant risks and clear done conditions
- Route tasks to model tiers (Haiku for classification/boilerplate, Sonnet for implementation, Opus for architecture)
- Track cost discipline per task including model, tokens, retries, time, and success/failure outcomes
- Conduct focused code review on invariants, edge cases, error boundaries, and security assumptions rather than style
How to install agentic-engineering
npx skills add null --skill agentic-engineeringHow to use agentic-engineering
- 1.Define your completion criteria and success metrics before starting any task
- 2.Write a capability eval that tests the core functionality you need
- 3.Write a regression eval that catches failures you've seen before
- 4.Run both evals on the current codebase to establish a baseline
- 5.Decompose your work into units following the 15-minute rule with single dominant risks
- 6.Route each unit to the appropriate model tier based on complexity
- 7.Execute the implementation and re-run evals to measure deltas
- 8.Review AI-generated code focusing on invariants, edge cases, error boundaries, and security rather than style
Use cases
- Large refactoring projects where you need to verify correctness across multiple files before and after changes
- Multi-step feature implementation where each step can be independently verified and tested
- Architecture decisions requiring root-cause analysis and evaluation of trade-offs
- Regression testing and quality gates for AI-generated code before human review
- Cost-optimized task execution by routing simple work to cheaper models and complex work to capable ones
- Engineering teams using AI agents for implementation work
- Technical leads managing AI-assisted development workflows
- Developers building systems that require high confidence in correctness
- Teams with budget constraints needing cost-aware model selection
agentic-engineering FAQ
Each decomposed unit should be independently verifiable, have a single dominant risk, and expose a clear done condition. This makes units small enough for focused execution and review.
Continue the session for closely-coupled units. Start fresh after major phase transitions. Compact after milestone completion, not during active debugging.
Use Haiku for classification and boilerplate transforms, Sonnet for implementation and refactors, and Opus for architecture decisions and multi-file invariants. Escalate only when a lower tier fails with a clear reasoning gap.
Prioritize invariants and edge cases, error boundaries, security and auth assumptions, and hidden coupling. Skip style-only disagreements if automated formatting already enforces style.
Use eval-first loops: define evals before execution, capture baseline failures, execute, then re-run evals to measure deltas. Track cost per task including model, tokens, retries, time, and success/failure.
Full instructions (SKILL.md)
Source of truth, from affaan-m/ecc.
name: agentic-engineering description: Operate as an agentic engineer using eval-first execution, decomposition, and cost-aware model routing. metadata: origin: ECC
Agentic Engineering
Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls.
Operating Principles
- Define completion criteria before execution.
- Decompose work into agent-sized units.
- Route model tiers by task complexity.
- Measure with evals and regression checks.
Eval-First Loop
- Define capability eval and regression eval.
- Run baseline and capture failure signatures.
- Execute implementation.
- Re-run evals and compare deltas.
Task Decomposition
Apply the 15-minute unit rule:
- each unit should be independently verifiable
- each unit should have a single dominant risk
- each unit should expose a clear done condition
Model Routing
- Haiku: classification, boilerplate transforms, narrow edits
- Sonnet: implementation and refactors
- Opus: architecture, root-cause analysis, multi-file invariants
Session Strategy
- Continue session for closely-coupled units.
- Start fresh session after major phase transitions.
- Compact after milestone completion, not during active debugging.
Review Focus for AI-Generated Code
Prioritize:
- invariants and edge cases
- error boundaries
- security and auth assumptions
- hidden coupling and rollout risk
Do not waste review cycles on style-only disagreements when automated format/lint already enforce style.
Cost Discipline
Track per task:
- model
- token estimate
- retries
- wall-clock time
- success/failure
Escalate model tier only when lower tier fails with a clear reasoning gap.
Related skills
More from affaan-m/ecc and the wider catalog.
agentic-os
Build persistent multi-agent operating systems on Claude Code with kernel routing, specialist agents, and file-based memory.
ai-first-engineering
Engineering operating model for teams shipping with AI-assisted code generation.
ai-regression-testing
Regression testing patterns for AI-assisted development to catch systematic blind spots.
android-clean-architecture
Clean Architecture patterns for Android and Kotlin Multiplatform projects with layered module structure and dependency rules.
angular-developer
Generates Angular code and provides architectural guidance for components, services, reactivity, forms, routing, and more.
api-connector-builder
Build API connectors matching your repo's existing integration pattern exactly.