enterprise-agent-ops
affaan-m/ecc
Operate long-lived agent workloads with observability, security boundaries, and lifecycle management.
What is enterprise-agent-ops?
This skill provides operational controls for cloud-hosted or continuously running agent systems. Use it when you need runtime lifecycle management, observability, safety controls, and change management beyond single CLI sessions.
- Manage runtime lifecycle (start, pause, stop, restart) for agent workloads
- Implement observability through logs, metrics, and traces
- Enforce safety controls including scopes, permissions, and kill switches
- Execute change management with rollout, rollback, and audit capabilities
- Deploy immutable artifacts with least-privilege credentials and environment-level secret injection
- Track key metrics: success rate, mean retries, time to recovery, cost per task, and failure distribution
How to install enterprise-agent-ops
npx skills add null --skill enterprise-agent-opsHow to use enterprise-agent-ops
- 1.Set up immutable deployment artifacts and configure least-privilege credentials for your agent workload
- 2.Inject environment-level secrets and establish hard timeout and retry budgets
- 3.Configure observability collection for logs, metrics, and traces from your running agents
- 4.Establish baseline safety controls including scopes, permissions, and kill switches
- 5.When failures occur, freeze new rollouts and capture representative traces to isolate the failing route
- 6.Apply the smallest safe patch, run regression and security checks, then resume rollout gradually
Use cases
- Operating production agent systems that run continuously across multiple environments
- Responding to failure spikes by freezing rollouts, capturing traces, and executing safe patches
- Managing multi-tenant agent deployments with strict permission boundaries and audit requirements
- Coordinating agent updates with regression and security checks before gradual resume
- Monitoring cost and performance metrics for long-running agent tasks
- DevOps engineers managing production agent infrastructure
- Platform teams building agent hosting and orchestration systems
- SREs responsible for agent reliability and incident response
- Engineering leads overseeing agent deployment and change control
enterprise-agent-ops FAQ
It pairs with PM2 workflows, systemd services, container orchestrators, and CI/CD gates.
Monitor success rate, mean retries per task, time to recovery, cost per successful task, and failure class distribution.
Freeze new rollouts, capture representative traces, isolate the failing route, patch with the smallest safe change, run regression and security checks, then resume gradually.
It enforces least-privilege credentials with environment-level secret injection rather than embedding secrets in deployment artifacts.
Full instructions (SKILL.md)
Source of truth, from affaan-m/ecc.
name: enterprise-agent-ops description: Operate long-lived agent workloads with observability, security boundaries, and lifecycle management. metadata: origin: ECC
Enterprise Agent Ops
Use this skill for cloud-hosted or continuously running agent systems that need operational controls beyond single CLI sessions.
Operational Domains
- runtime lifecycle (start, pause, stop, restart)
- observability (logs, metrics, traces)
- safety controls (scopes, permissions, kill switches)
- change management (rollout, rollback, audit)
Baseline Controls
- immutable deployment artifacts
- least-privilege credentials
- environment-level secret injection
- hard timeout and retry budgets
- audit log for high-risk actions
Metrics to Track
- success rate
- mean retries per task
- time to recovery
- cost per successful task
- failure class distribution
Incident Pattern
When failure spikes:
- freeze new rollout
- capture representative traces
- isolate failing route
- patch with smallest safe change
- run regression + security checks
- resume gradually
Deployment Integrations
This skill pairs with:
- PM2 workflows
- systemd services
- container orchestrators
- CI/CD gates
Related skills
More from affaan-m/ecc and the wider catalog.
error-handling
Patterns for robust error handling across TypeScript, Python, and Go with typed errors, retries, and circuit breakers.
eval-harness
Formal evaluation framework for Claude Code sessions using eval-driven development (EDD) principles
everything-claude-code
Development conventions and patterns for the everything-claude-code JavaScript project.
everything-claude-code-conventions
Development conventions and patterns for the everything-claude-code JavaScript project.
evm-token-decimals
Prevent silent decimal mismatch bugs across EVM chains with runtime lookup and chain-aware caching.
exa-search
Neural search via Exa MCP for web, code, companies, and people research.