PluginBench
Skill
Pass
Audit score 90

debug-my-harness

zernie/vigiles

Diagnose harness misbehavior by analyzing the flight-recorder ledger (.vigiles/runs.jsonl).

What is debug-my-harness?

Reads the local append-only ledger that vigiles writes during harness execution to diagnose why skills fired or didn't fire, why hooks allowed or blocked incorrectly, and whether subagents violated their tool contracts. Use this when debugging unexpected harness behavior; it's evidence-based diagnosis, not rule-writing.

  • Reads and parses the flight-recorder ledger (.vigiles/runs.jsonl) to extract skill activations, hook decisions, agent contract violations, and trigger-rate metrics
  • Identifies skill selection collisions by comparing fire rates and descriptions when the wrong skill runs or one stops firing
  • Detects hook misconfigurations: rules in observe mode (shadow, never blocking), missing gates, or incorrect allow/deny decisions
  • Flags subagent tool-contract violations (allowed: false records) to pinpoint where agents reached outside their declared scope
  • Compares eval metrics across runs to detect performance drift in recall and precision

How to install debug-my-harness

npx skills add https://github.com/zernie/vigiles --skill debug-my-harness
Prerequisites
  • vigiles installed and configured in the project
  • At least one harness run completed (to populate .vigiles/runs.jsonl)
Claude Code
Cursor
Windsurf
Cline

How to use debug-my-harness

  1. 1.Run the harness or `vigiles audit` to generate flight-recorder ledger entries in .vigiles/runs.jsonl
  2. 2.Ask the skill a specific question: why a skill stopped firing, why a hook didn't block, or whether a subagent misbehaved
  3. 3.The skill reads the ledger, counts skill fires, checks hook decisions, lists agent violations, and compares metrics
  4. 4.Review the evidence-based diagnosis and the specific ledger records cited
  5. 5.Hand off to `edit-spec` for spec changes, `strengthen` for new linter rules, or `test-harness` to measure behavior after fixes

Use cases

Good for
  • Debugging why a skill stopped firing or why a different skill runs instead
  • Investigating why a security hook didn't block a disallowed action
  • Diagnosing subagent tool-contract violations and determining whether to tighten or widen the contract
  • Measuring trigger-rate changes and detecting harness drift after model or spec upgrades
  • Analyzing a series of runs to identify patterns in misbehavior
Who it's for
  • Agent harness maintainers troubleshooting skill or hook behavior
  • Security engineers validating that gates and contracts are enforced correctly
  • Developers investigating why a harness change caused unexpected behavior

debug-my-harness FAQ

What if .vigiles/runs.jsonl doesn't exist or is empty?

The skill will report that no ledger exists and suggest running the harness or `vigiles audit` first to generate records.

Can this skill write new rules or edit the harness spec?

No. This skill diagnoses only. Use `strengthen` to promote guidance to linter rules, or `edit-spec` to modify the spec.

How do I know if a hook is in observe mode vs. enforce mode?

The ledger records the `mode` field for each hook decision. Observe mode means the hook never actually blocks; it only logs.

What does a trigger-rate metric tell me?

Eval records show recall and precision for skill activation. A downward trend across runs indicates drift, often after a model or harness upgrade.

What should I do if I find a subagent contract violation?

The skill will list the agent, tool, and reason. Decide whether to tighten the tool-contract (restrict the agent) or widen it (allow the tool), then use `edit-spec` to update it.

Full instructions (SKILL.md)

Source of truth, from zernie/vigiles.


name: debug-my-harness allowed-tools: Read, Glob, Grep description: Diagnose why an agent harness misbehaved by reading the local flight-recorder ledger (.vigiles/runs.jsonl) — which skills fired or got hijacked, which hooks blocked or wrongly allowed, which subagent tool-contract violations happened, and how a skill's trigger rate moved. Use when asked why a skill stopped firing, why a hook didn't block, why the wrong skill ran, or to debug/investigate what the harness actually did. NOT for writing new rules (use strengthen) or editing the spec (use edit-spec).

Diagnose harness misbehavior from the flight recorder — the local, append-only ledger at .vigiles/runs.jsonl that vigiles writes as your harness runs. It records what actually happened, so you debug from evidence instead of guessing.

What's in the ledger

One JSON record per line, each with a kind:

  • hook — a compiled-hook gate decision: {event, decision: allow|deny|ask, mode: enforce|observe, rule, cmd, reason}.
  • agent — a subagent tool-contract decision: {name, tool, allowed, reason} (a false = the agent went outside its lane).
  • skill — a skill activation: {name, fired}.
  • eval — a measured metric: {name, metric, value} (e.g. trigger-rate recall/precision).
  • capability-diff — a blast-radius change: {pr, added, removed, widened}.

Instructions

Step 1: Read the ledger

Read .vigiles/runs.jsonl (JSONL — one record per line; tolerate a torn last line). If it's absent or empty, say so — there's nothing recorded yet; suggest running the harness (or vigiles audit) first. Do NOT fabricate records.

Step 2: Answer the specific question, evidence-first

Match the user's question to the ledger:

  • "Why did skill X stop firing / why does the wrong one run?" — count skill fires by name over time. If X's fire-rate dropped, look for a sibling that fired on the same kinds of prompts (a selection collision) and check their descriptions for overlap. Recommend differentiating or merging the descriptions.
  • "Why didn't my hook block that?" — find hook records for the event. A decision: allow on something that should be denied, or mode: observe (shadow, never blocks), or the absence of any record, tells you which. Recommend flipping observe→enforce or fixing the gate logic.
  • "Did a subagent misbehave?" — list agent records with allowed: false: the agent reached for a tool outside its declared contract. Point at the contract to tighten or widen.
  • "Is it getting worse?" — compare eval metric values (recall/precision) across runs; a downward trend is drift (often after a harness/model upgrade).

Step 3: Recommend a fix, tied to the evidence

Prefer promoting an ignored-but-decidable rule from prose to a deterministic gate: a repeated agent violation or a rule the agent keeps breaking → a compiled hook or a tighter tool-contract (the strengthen skill can help). A description collision → differentiate the skill descriptions. Always cite the specific records you based the diagnosis on.

Step 4: Offer the next step

If the fix is a spec change, hand off to edit-spec. If it's promoting guidance to a linter rule, hand off to strengthen. If a behavioral claim needs measuring (does the skill fire now?), hand off to test-harness (measureTriggerRate).