investigate-first
juliusbrussee/caveman
Diagnose ambiguous failures by gathering evidence before editing code.
What is investigate-first?
A systematic investigation methodology for unknown causes, intermittent behavior, performance regressions, and other ambiguous failures. Use this skill to separate observed symptoms from inferred causes, trace execution paths, and rank hypotheses by evidence before making any code changes.
- Separate observed symptom from inferred cause
- Trace inputs, state transitions, ownership boundaries, and failure output
- Rank hypotheses by evidence and cheap falsification value
- Avoid editing until one credible mechanism explains the evidence
- Stop exploration when evidence is sufficient to name the cause or exact blocker
How to install investigate-first
npx skills add https://github.com/juliusbrussee/caveman --skill investigate-firstHow to use investigate-first
- 1.Describe the observed symptom clearly without assumptions about cause
- 2.Trace the execution path: inputs, state changes, boundaries crossed, and failure output
- 3.List all plausible hypotheses that could explain the evidence
- 4.Rank hypotheses by how much evidence supports each and how cheaply you can test them
- 5.Test or gather evidence for the highest-ranked hypothesis first
- 6.Stop when you have sufficient evidence to name the cause or identify the exact blocker
- 7.Report the cause and proof; do not edit code unless the task explicitly authorizes implementation
Use cases
- Diagnosing intermittent bugs with unclear root causes
- Investigating performance regressions to identify the actual bottleneck
- Tracing failures across system boundaries to find ownership and responsibility
- Evaluating multiple competing hypotheses for a single symptom
- Gathering proof before proposing a fix to stakeholders
- Backend engineers debugging production issues
- Full-stack developers investigating cross-layer failures
- Performance engineers analyzing regressions
- QA engineers documenting failure mechanisms
- Technical leads reviewing root-cause analyses
investigate-first FAQ
Use it when the cause is unclear, the failure is intermittent, multiple hypotheses seem plausible, or the fix could have unintended side effects. Skip it only when the cause is obvious and the fix is low-risk.
Stop when you can name the specific cause or identify the exact blocker that explains all observed symptoms. You do not need to implement the fix—just prove what is broken and why.
It means testing hypotheses that are quick and easy to rule out first, before spending time on expensive or time-consuming tests. Prioritize tests that can quickly eliminate wrong answers.
Yes. Report the cause and proof so that the task owner or another engineer can decide whether to implement a fix. Your job is diagnosis, not necessarily implementation.
Follow the data and control flow across service calls, API boundaries, database queries, and inter-process communication. Document which component owns each state change and where failures occur relative to those boundaries.
Full instructions (SKILL.md)
Source of truth, from juliusbrussee/caveman.
name: investigate-first description: Diagnose ambiguous failures before editing. Use for unknown causes, intermittent behavior, performance regressions, or investigations needing evidence-ranked hypotheses.
Investigate first
Gather evidence before changing product code.
- Separate observed symptom from inferred cause.
- Trace inputs, state transitions, ownership boundaries, and failure output.
- Rank hypotheses by evidence and cheap falsification value.
- Do not edit until one credible mechanism explains evidence.
- Stop exploration when evidence is sufficient to name cause or exact blocker.
Report cause and proof. Make no fix unless task authorizes implementation.
Related skills
More from juliusbrussee/caveman and the wider catalog.

lean-build
Build focused features with strict scope and explicit stop conditions to avoid overbuilding.

migration
Implement reversible, compatibility-safe transitions for schema, data, API, and dependency migrations with rollback.

safe-refactor
Restructure code while preserving behavior through bracketed verification.

surgical-patch
Fix bugs at the narrowest responsible layer with regression proof and preserved surrounding behavior.

verify-and-stop
Prove existing work meets acceptance conditions without expanding scope.

context-canary
Install a per-turn canary signal to detect silent context degradation and trigger recovery when it fails.