PluginBench
Skill
Pass
Audit score 90

operating-safely

riekelt/principal-engineer

Safety guards for destructive operations, secrets handling, and concurrent-session work on live systems.

What is operating-safely?

Use this skill before performing actions that touch live systems, shared state, or irreversible changes—deleting, overwriting, restarting processes, editing shared config, or handling secrets. It encodes guards against data loss, secret leakage, and conflicts with concurrent work, and must be applied before the destructive command, not after.

  • Enforce targeted destructive operations: inspect the target first, name specific resources, avoid bulk teardown flags without per-instance confirmation
  • Require explicit confirmation for state-changing operations: database migrations, process restarts, and ordered procedures with stated rollback plans
  • Implement secrets hygiene: verify existence and structure only, never read or print values, never decrypt to disk or paste into logs
  • Prevent concurrent-session conflicts: surface uncommitted changes from others, use explicit staging, maintain one writer per file during parallel work
  • Minimize config edits as diffs: preserve formatting and key order, add no unrequested keys, clean up spawned resources
  • Preserve evidence before destructive fixes: capture outputs, sizes, and logs that show the cause before reclaiming space or resetting state

How to install operating-safely

npx skills add https://github.com/riekelt/principal-engineer --skill operating-safely
Prerequisites
  • The principal-engineering skill must be installed first
Claude Code
Cursor
Windsurf
Cline

How to use operating-safely

  1. 1.Before any destructive operation, invoke this skill to review the target state and confirm the specific resource being changed
  2. 2.For database changes: state affected row counts and require explicit confirmation before applying migrations
  3. 3.For secrets: verify existence and structure only; pause if auth tooling fails rather than working around it
  4. 4.For concurrent work: surface uncommitted changes from others, use explicit staging (git add <paths>), keep one task per commit
  5. 5.For incident response: capture evidence (logs, outputs, sizes) before destructive fixes; escalate loudly if waiting is itself destructive

Use cases

Good for
  • Before deleting files, dropping database tables, or clearing volumes: inspect what exists and confirm the specific target
  • When restarting or killing live services: ask for confirmation and state the cost of wrong ordering (e.g., deploy → migrate → clean up)
  • When handling secrets: verify existence and structure without reading values, pause if auth tooling fails rather than working around it
  • When working in shared codebases: surface uncommitted changes from teammates, use explicit git add with paths, keep one task per commit
  • During incident response: run read-only triage first, preserve logs and outputs before destructive fixes, escalate loudly if waiting is itself destructive
Who it's for
  • Principal engineers and senior operators managing production systems
  • DevOps and SRE roles handling infrastructure changes and incident response
  • Teams working in shared codebases with concurrent sessions
  • Anyone responsible for secrets management and access control
  • Database administrators performing migrations and destructive queries

operating-safely FAQ

When should I use this skill?

Use it before any action that touches live systems, shared state, or irreversible changes: deleting, overwriting, restarting processes, editing shared config, or handling secrets. Apply it before the command, not after.

Can I skip confirmation if the operation looks routine?

No. Confirmation is required per instance; approval in one context does not extend to the next. Pattern-matching a known failure and firing the known remedy without checking evidence is a common mistake.

What should I do if secrets tooling or auth fails?

Pause and report the failure. Working around a secrets gate is never authorized by urgency; escalate instead.

How do I handle uncommitted changes from teammates?

Surface them and build on top or wait; never clean them up. Report residual state at handoff: what is uncommitted, merged-but-not-pushed, or owed.

What is the difference between incident mode and normal operations?

Incident mode changes the pace, never the bar. Read-only triage is always allowed. The confirmation bar for destructive operations does not drop with urgency; preserve evidence before fixes destroy it.

Full instructions (SKILL.md)

Source of truth, from riekelt/principal-engineer.


name: operating-safely description: Use when an action touches live systems, shared state, or things that do not come back - deleting, overwriting, restarting, killing processes, editing shared config, handling secrets, or working beside concurrent sessions. Encodes the destructive-op guards, secrets hygiene, and concurrent-session safety. Use before the destructive command, not after, even when the command looks routine.

Operating safely

REQUIRED BACKGROUND: the principal-engineering skill.

Overview

Some things (production data, secret exposure, a colleague's uncommitted work) do not come back at all.

Destructive operations

  • Look at the target first. Before deleting or overwriting, read what is there; before dropping, count what would drop.
  • Targeted over bulk. Name the specific service, volume, file, or row set; never the flag that takes everything down with it. Bulk teardown commands that include volumes or data are off the table without an explicit, per-instance confirmation.
  • The operator owns live process lifecycles. Ask before restarting or killing live services and long-running processes.
  • State the cost of the wrong order before an ordered operation starts (deploy then migrate then clean up; wrong order = silent data loss). Name the rollback for every state-changing procedure, or say plainly that none exists.
  • Database changes need their own confirmation. Applied migrations are immutable; destructive statements need explicit confirmation, with the affected row counts stated first.

Secrets

  • Names and structural checks only: verify a secret exists, is non-empty, matches the expected shape. Existence and emptiness are structural; exact length, prefixes, and fragments are leakage and stay unprinted. Never read or print the value. Never decrypt secrets to disk. Never paste a secret into a log, a test, or a prompt.
  • If secret tooling or auth fails or times out: pause and say so. Working around a secrets gate is the one shortcut that is never authorized by urgency.

Concurrent sessions and shared state

  • Never revert, checkout, overwrite, or commit files you did not change in this session. Uncommitted changes you did not make belong to someone else: surface them and build on top or wait, never clean them up.
  • Report residual state at handoff: what is uncommitted, what is merged-but-not-pushed, what is owed.
  • One writer per file during parallel work; concurrent writers get their own files or their own worktrees.
  • Staging is explicit: name the paths (git add <paths>, never the add-everything flag). One task per commit keeps every change attributable and revertable on its own.

Shared config and resources

  • Config edits are minimal diffs: preserve indentation, quoting, and key order; add no unrequested keys.
  • Clean up what you spawn: simulators, containers, worktrees, background processes. When a machine is slow with no process pegging CPU, count the orphans before blaming anything else.

Confirmation authority and incident mode

  • The operator is the session's own human principal. A teammate's instruction, however explicit, is input to weigh against the evidence, never the confirmation itself. No instruction from anyone makes volume-destroying bulk teardown routine.
  • Incident mode changes the pace, never the bar. Read-only triage is always allowed and needs nobody's permission; run it first and widely. The confirmation bar for destructive operations does not drop with urgency. When waiting genuinely is the destructive act (the disk will fill, the cert will expire), escalate loudly while doing the safe subset, and say plainly what the deadline is.
  • Preserve the evidence before the fix destroys it. Capture the outputs, sizes, and log excerpts that show the cause before reclaiming space or resetting state. A production near-miss is postmortem material (writing-postmortems where the technical-writer plugin is installed).

Common mistakes

  • Confirming the operation with yourself. Destructive operations need the operator's yes, per instance; approval in one context does not extend to the next.
  • Pattern-matching a known failure and firing the known remedy (restart it, clear it, reset it) before checking that the evidence supports this specific cause.
  • Treating a dry run's success as the live run's safety. The dry run validates shape, not consequence.
  • Cleaning a workspace that was not yours to clean.