diagnosing-bugs
vinvcn/mattpocock-skills-zh-cn
Structured diagnostic loop for tricky bugs and performance regressions.
What is diagnosing-bugs?
A disciplined debugging methodology for when code is broken, throwing, failing, or slow. Guides you through building a tight feedback loop, reproducing and minimizing the issue, generating ranked hypotheses, instrumenting strategically, fixing with regression tests, and cleaning up—emphasizing that a good feedback loop solves 90% of the problem.
- Build tight, red-capable feedback loops (failing tests, curl scripts, CLI invocations, headless browser scripts, replay harnesses, or fuzz loops) that reliably trigger the exact bug symptom
- Reproduce and minimize bugs to the smallest scenario that still fails, reducing hypothesis space for Phase 3
- Generate and rank 3–5 falsifiable hypotheses with explicit predictions before testing any assumption
- Instrument code with targeted, tagged debug logs and debugger breakpoints mapped to specific predictions
- Write regression tests at the correct seam before applying fixes, then verify against the original scenario
- Clean up all debug instrumentation, throwaway prototypes, and document the correct hypothesis in commit messages
How to install diagnosing-bugs
npx skills add https://github.com/vinvcn/mattpocock-skills-zh-cn --skill diagnosing-bugsHow to use diagnosing-bugs
- 1.Read CONTEXT.md (if it exists) and ADRs for the relevant modules to build a mental model of the codebase
- 2.Phase 1: Build a feedback loop—construct a tight, red-capable command (test, curl script, CLI, headless browser, replay harness, or fuzz loop) that reliably triggers the exact symptom; tighten it for speed, signal clarity, and determinism; for flaky bugs, raise reproduction rate until debuggable
- 3.Phase 2: Run the loop, confirm it reproduces the user's exact failure mode, then minimize the repro by removing non-load-bearing inputs, config, and steps until each remaining element is essential
- 4.Phase 3: Generate 3–5 ranked, falsifiable hypotheses with explicit predictions; share the ranked list with the user for domain-knowledge input before testing
- 5.Phase 4: Instrument strategically—use debugger breakpoints first, then targeted logs with unique prefixes (e.g., [DEBUG-a4f2]); for performance bugs, establish baseline measurements and bisect rather than logging
- 6.Phase 5: Write a regression test at the correct seam (where the real bug pattern triggers), apply the fix, verify the test passes, then re-run the original Phase 1 loop against the unfixed scenario
- 7.Phase 6: Remove all [DEBUG-...] tagged instrumentation, delete throwaway prototypes, verify the original repro no longer occurs, and document the correct hypothesis in the commit message
Use cases
- A user reports 'diagnose this' or 'debug this' with a failing test, error message, or performance regression
- Intermittent or flaky bugs where you need to raise reproduction rate to make them debuggable
- Performance regressions where you must establish baseline measurements and bisect before hypothesizing
- Complex multi-component failures where a tight feedback loop is the only way to isolate the root cause
- Production issues where you have captured artifacts (HAR files, logs, traces) but need to replay them in isolation
- Backend and full-stack engineers debugging production or staging issues
- Frontend developers troubleshooting UI bugs or performance regressions
- QA engineers or developers working with intermittent or hard-to-reproduce failures
- Anyone inheriting unfamiliar codebases and needing to diagnose existing bugs systematically
diagnosing-bugs FAQ
Stop and explicitly state what you've tried. Ask the user for (a) access to a reproducible environment, (b) redacted captured artifacts (HAR files, logs, core dumps, screen recordings with timestamps), or (c) permission to add temporary production instrumentation. Do not hypothesize without a loop.
The goal is not a clean repro but a higher reproduction rate. Run the loop 100x, parallelize, add stress, shrink timing windows, or inject sleeps. A 50%-flaky bug is debuggable; 1%-flaky is not. Keep raising the reproduction rate until it's deterministic enough to debug.
Yes, always. Replace secrets with <REDACTED>. Keep credentials in environment variables rather than in displayed output. If redacted output is insufficient to diagnose the bug, say so and ask the user for more information.
That's a valid finding—document it. It means the codebase architecture prevents you from locking down the bug at a good testing layer. Record this as a discovery for the next debugger and note it in your cleanup.
You have a single command (script path, test invocation, or curl) that is red-capable (catches the exact bug symptom), deterministic (same verdict each run), fast (seconds, not minutes), and agent-runnable (no human interaction needed). You've run it at least once and shown the invocation and redacted output.
Full instructions (SKILL.md)
Source of truth, from vinvcn/mattpocock-skills-zh-cn.
name: diagnosing-bugs description: 面向棘手缺陷和性能回退的诊断循环。适用于用户说 “diagnose” / “debug this”,或报告某些东西 broken、throwing、failing、slow 时。
Diagnosing Bugs
面向棘手 bugs 的纪律。只有在明确说明理由时才跳过阶段。
探索 codebase 时,先读取 CONTEXT.md(如果存在),建立相关 modules 的清晰 mental model,并检查你将触碰区域的 ADRs。
Redact
这个 skill 会要求你展示 commands、outputs 和捕获的 artifacts。先 redact 掉每个 secret——用 <REDACTED> 替换。Build loops 要针对 env vars 进行,让 credential 留在 environment 里而不是你展示的内容中。捕获的 artifacts 带有 auth headers:只引用携带 signal 的那些行。
如果 redact 后的 output 不足以诊断 bug,就说明情况并询问用户。
Phase 1 - Build a feedback loop
这就是这个 skill 的核心。 其他所有内容都是机械步骤。如果你拥有一个针对该 bug 的 tight pass/fail signal,即它会在 这个 bug 上变红,你就能找到原因;bisection、hypothesis-testing 和 instrumentation 都只是消费这个 signal。没有它,盯着代码看多久都救不了你。
在这里投入不成比例的精力。要强硬、要有创造力、拒绝放弃。
Ways to construct one - try them in roughly this order
- Failing test,放在能触达 bug 的 seam 上:unit、integration、e2e 都可以。
- Curl / HTTP script,打到运行中的 dev server。
- CLI invocation,使用 fixture input,并把 stdout 与 known-good snapshot diff。
- Headless browser script(Playwright / Puppeteer),驱动 UI,并断言 DOM/console/network。
- Replay a captured trace. 把真实 network request / payload / event log 保存到磁盘,并在隔离环境中 replay 到代码路径。
- Throwaway harness. 启动系统的最小子集(一个 service、mocked deps),用一次 function call 触发 bug code path。
- Property / fuzz loop. 如果 bug 是 "sometimes wrong output",运行 1000 个 random inputs 并寻找 failure mode。
- Bisection harness. 如果 bug 出现在两个已知状态之间(commit、dataset、version),自动化 "boot at state X, check, repeat",这样可以
git bisect run。 - Differential loop. 用同一 input 跑 old-version vs new-version(或两个 configs),然后 diff outputs。
- HITL bash script. 最后手段。如果必须由人点击,就用
scripts/hitl-loop.template.sh驱动 人,让 loop 仍保持结构化。捕获的输出反馈给你。
构建正确的 feedback loop,bug 就修好了 90%。
Tighten the loop
把 loop 当作产品。只要有了 一个 loop,就继续 tighten 它:
- 我能让它更快吗?(Cache setup、跳过无关 init、缩小 test scope。)
- 我能让 signal 更尖锐吗?(断言具体 symptom,而不是 "didn't crash"。)
- 我能让它更 deterministic 吗?(Pin time、seed RNG、isolate filesystem、freeze network。)
一个 30 秒且 flaky 的 loop 几乎不比没有 loop 好;一个 2 秒 deterministic loop 才是 tight 的调试超能力。
Non-deterministic bugs
目标不是 clean repro,而是 higher reproduction rate。循环触发 100x、parallelise、加 stress、缩小 timing windows、注入 sleeps。50%-flake bug 可以调试;1% 不行。持续提高复现率,直到它可调试。
When you genuinely cannot build a loop
停下来并明确说明。列出你尝试过什么。向用户请求:(a) 能复现的环境访问权限,(b) 经 redact 的 captured artifact(HAR file、log dump、core dump、带 timestamps 的 screen recording),或 (c) 添加临时 production instrumentation 的许可。不要 在没有 loop 时继续 hypothesise。
Completion criterion - a tight loop that goes red
Phase 1 完成条件:loop tight 且 red-capable。你能指出 一个 command(script path、test invocation、curl),并且你已经至少运行过一次(展示 invocation 和 output,已 redact),且它满足:
- Red-capable - 它驱动真实 bug code path,并断言 用户的 exact symptom,因此能在该 bug 上变红、修复后变绿。不是 "runs without erroring",而是必须能 catch this specific bug。
- Deterministic - 每次运行 verdict 相同(flaky bugs:按上文固定到高复现率)。
- Fast - 秒级,而不是分钟级。
- Agent-runnable - 你可以无人值守运行;human in the loop 只能通过
scripts/hitl-loop.template.sh。
如果你发现自己在 command 存在前就读代码构建理论,停下;直接跳到 hypothesis 正是这个 skill 要防止的失败。 没有 red-capable command,就没有 Phase 2。
Phase 2 - Reproduce + minimise
运行 loop。看它变红,也就是 bug 出现。
确认:
- Loop 产出的 failure mode 是 用户 描述的那个,而不是附近另一个失败。Wrong bug = wrong fix。
- Failure 能在多次运行中复现(或对于 non-deterministic bugs,复现率足够高,能用来调试)。
- 你已捕获 exact symptom(error message、wrong output、slow timing),后续阶段可以验证 fix 确实解决它。
Minimise
一旦变红,就把 repro 缩到 仍会变红的最小场景。逐个削减 inputs、callers、config、data 和 steps,每次削减后重新运行 loop;只保留 failure 的 load-bearing 部分。
原因:minimal repro 会缩小 Phase 3 的 hypothesis space(可怀疑的 moving parts 更少),并成为 Phase 5 中干净的 regression test。
完成条件:每个剩余元素都是 load-bearing,移除任意一个都会让 loop 变绿。
在 reproduce 并 minimise 之前不要继续。
Phase 3 - Hypothesise
在测试任何假设前,生成 3-5 个 ranked hypotheses。单假设会锚定在第一个看似合理的想法上。
每个 hypothesis 必须 falsifiable:说明它会做出什么 prediction。
Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
如果无法说明 prediction,这就是 vibe;丢弃或打磨它。
测试前把 ranked list 展示给用户。 用户常常有 domain knowledge,可以立即重排("we just deployed a change to #3"),或知道哪些 hypotheses 已被排除。便宜 checkpoint,大幅省时。不要因此阻塞;如果用户 AFK,就按你的排序继续。
Phase 4 - Instrument
每个 probe 都必须映射到 Phase 3 的某个具体 prediction。一次只改变一个变量。
Tool preference:
- Debugger / REPL inspection,如果环境支持。一个 breakpoint 胜过十条 logs。
- Targeted logs,放在能区分 hypotheses 的 boundaries。
- 永远不要 "log everything and grep"。
给每条 debug log 加唯一 prefix,例如 [DEBUG-a4f2]。最后 cleanup 就能一次 grep。未打 tag 的 logs 会存活;带 tag 的 logs 要删除。
Perf branch。 对 performance regressions,logs 通常不对。改为先建立 baseline measurement(timing harness、performance.now()、profiler、query plan),然后 bisect。先 measure,再 fix。
Phase 5 - Fix + regression test
在 fix 前写 regression test,但前提是存在 correct seam。
Correct seam 是 test 能以 call site 中真实发生的方式触发 real bug pattern 的地方。如果唯一可用 seam 太 shallow(bug 需要多个 callers,但 test 只有 single-caller;unit test 无法复制触发 bug 的 chain),那里的 regression test 会给出 false confidence。
如果不存在 correct seam,这本身就是发现。 记录下来。Codebase architecture 阻止你锁住 bug。把它标记给下一阶段。
如果存在 correct seam:
- 把 minimised repro 变成该 seam 上的 failing test。
- 看它 fail。
- 应用 fix。
- 看它 pass。
- 重新针对原始(未 minimised)场景运行 Phase 1 feedback loop。
Phase 6 - Cleanup
声明完成前必须做:
- Original repro 不再复现(重跑 Phase 1 loop)
- Regression test 通过(或记录缺少 seam)
- 所有
[DEBUG-...]instrumentation 已移除(grep prefix) - Throwaway prototypes 已删除(或移动到明确标记的 debug location)
- 正确 hypothesis 已写进 commit / PR message,让下一个 debugger 能学习
Related skills
More from vinvcn/mattpocock-skills-zh-cn and the wider catalog.

domain-modeling
Build and refine your project's domain model through active terminology and decision capture.

edit-article
Restructure and refine articles by reorganizing sections, improving clarity, and tightening prose.

git-guardrails-claude-code
Prevent destructive git operations in Claude Code by blocking dangerous commands before execution.

grill-me
Continuous questioning interview to refine plans or designs.

grill-with-docs
Iterative questioning interviews to refine plans and designs while generating ADRs and glossaries.

grilling
Systematically pressure-test plans, decisions, or ideas through structured questioning.