writing-unit-tests
riekelt/principal-engineer
Behavior-first unit testing: one claim per test, deterministic setup, mocks only at external boundaries.
What is writing-unit-tests?
A guide to writing maintainable unit tests that verify contracts rather than implementations. Use when creating new tests, refactoring existing ones, or fixing flaky or unreadable tests. Emphasizes clear naming, deterministic execution, and strategic mocking.
- Test through public contracts, not implementation details, so refactors don't break passing tests
- Name tests as behavioral claims (subject, scenario, outcome) that serve as documentation
- Arrange-act-assert structure with no branching or logic inside test bodies
- Mock only external boundaries you don't own; use real collaborators for your own code
- Eliminate flakiness by injecting time/randomness and avoiding sleeps or real I/O
- Assert outcomes with values, not absence of exceptions, and test failure paths as first-class subjects
How to install writing-unit-tests
npx skills add https://github.com/riekelt/principal-engineer --skill writing-unit-tests- Familiarity with the principal-engineering skill
- Understanding of your testing framework (Jest, pytest, etc.)
How to use writing-unit-tests
- 1.Identify the public contract (inputs, outputs, side effects) of the unit under test
- 2.Write a test name that states the claim: subject, scenario, and expected outcome
- 3.Arrange test state using builders or role-named fixtures, keeping setup visible
- 4.Act by calling the unit once through its public interface
- 5.Assert the outcome with literal expected values (derived independently, not from production code)
- 6.Review for common mistakes: mirror tests, mega-tests, mocks of your own code, sleeps, or conditional assertions
Use cases
- Writing a new test file for a module with clear behavioral contracts
- Refactoring an existing test suite to eliminate flaky sleeps and improve readability
- Fixing a test that breaks after a refactor by distinguishing contract from implementation
- Debugging a test that mocks internal methods instead of external boundaries
- Creating deterministic tests for time-dependent or randomness-dependent code
- Software engineers writing or maintaining unit tests
- Teams adopting behavior-driven testing practices
- Engineers refactoring legacy test suites
- Anyone debugging flaky or hard-to-read tests
writing-unit-tests FAQ
Mock boundaries you don't own (network, filesystem, clock, third-party services). Use real collaborators for code you own within the unit's reach. For your own wrappers around external resources, mock at the seam where your code last touches the unowned resource.
Inject the clock or time provider and control it in the test. For async work, poll the observable outcome with a deadline rather than sleeping for a fixed duration.
It creates a mirror test that stays green even when the implementation is wrong. Expected values should be literals worked out independently—by hand, from a spec, or from real data—with the derivation documented in a comment.
Either widen the unit to something with real behavior, or accept that this seam needs an integration test instead and document which one.
A test with branching, loops, or conditional logic can itself be wrong, and failures become ambiguous. Keep tests simple and logic-free; move generation and iteration into builders and helpers.
Full instructions (SKILL.md)
Source of truth, from riekelt/principal-engineer.
name: writing-unit-tests description: "Use when writing or refactoring unit tests - a new test file, added cases, a flaky test, an unreadable one. Encodes behavior-first testing: one behavior per test, names that state the claim, deterministic setup, mocks only at boundaries you do not own. Use whenever a test is being written, even a quick one, and whenever a test needs a sleep, a mock of your own code, or a copy of the implementation's math."
Writing unit tests
REQUIRED BACKGROUND: the principal-engineering skill. testing-changes governs which tests a change owes; this skill is the craft of the tests themselves.
Overview
A unit test is a behavioral claim with a name, read by the next engineer during a red build. Core principle: test the contract, not the implementation. The name carries the claim, and the test stays simple enough that it cannot itself be wrong.
Contract over implementation
- Test through the public contract of the unit. A refactor that preserves behavior should not break tests; when it does, the tests were asserting the implementation.
- Do not assert call sequences, internal state, or that method A called method B, unless the interaction IS the contract (a required side effect on a boundary).
- Never derive the expected value from the production arithmetic, neither by reimplementing the formula nor by invoking the shared helper that computes it. Expected values are literals worked out independently (by hand, from a spec, from real data), with the derivation in a comment.
One behavior per test, named as the claim
- One behavior per test; splitting is cheaper than archaeology on a multi-assert failure.
- The name states subject, scenario, and expected outcome:
expired_token_is_rejected_with_401, nottest_auth_3. Test names describe behavior, state transitions, and invariants; never delivery order, ticket keys, or phases. - Arrange, act, assert, visibly and in that order. No branching, loops, or logic in a test: a test with logic needs its own test. Shared setup earns a builder or a role-named fixture; a mystery blob fixture hides which arranged fact the assertion depends on. Generation and iteration live in builders and helpers, not in the test body. Property-based tests are the accepted form for invariants and follow their framework's shape; example-based tests stay logic-free.
Determinism
- No real time, real randomness, real network, or real filesystem inside a unit test: inject the clock, seed or inject the randomness, mock the boundary.
- No sleeps. Waiting for async work is condition-based (poll the observable outcome with a deadline), never duration-based.
- A flaky test is red: fix it or quarantine it visibly with an owner (see the red-test rule in
testing-changes); re-running until green is silencing a detector.
Mocks are assumptions
- Mock the boundaries you do not own (network, clock, filesystem, third-party services); prefer real collaborators for code you do own within the unit's reach. For owned wrappers around unowned resources (your repository class fronting the database), mock at the seam where owned code last touches the unowned resource, and keep the test data role-named and visible either way. Every mock hardcodes an assumption about a contract; a stale mock is how a suite stays green while the real integration is broken.
- When a test is mostly mock wiring, it is testing the mocks. Either widen the unit to something with real behavior or accept that this seam needs an integration test instead (and say which).
- Fixtures are labeled snapshots of reality: minimal, role-named for their part in the scenario, updated deliberately when the contract changes, never regenerated blindly to make red go green.
Assertions and failure paths
- Assert outcomes with values, not absence of exceptions. "It did not throw" claims almost nothing.
- Failure paths are first-class test subjects: the typed failure surfaces, the degraded mode is entered loudly, the guard actually guards (see
handling-failures). - Tests themselves follow the no-silent-swallows contract: no catch-and-ignore in test code, no conditional assertions that skip silently when a precondition is absent. A test that cannot run must fail or be visibly skipped with the reason.
Common mistakes
- The mirror test: reimplementing the production logic to compute the expectation.
- The mock echo chamber: mocking your own class and asserting the mock.
- The mega-test: twelve assertions, one name, no way to know which claim broke.
- Shared mutable fixtures that make test order matter; every test builds or receives its own state.
- The sleep that "fixes" flakiness by making it rarer.
- Green-checking the fixture: editing expected values to match actual output without deriving why the new value is right.
Related skills
More from riekelt/principal-engineer and the wider catalog.

adding-dependencies
Systematic framework for adding, vetting, updating, and removing dependencies with cost-aware decision-making.

grounding-before-coding
Map real code and data before writing—ground every claim in evidence before any change.

guarding-architecture
Encode structural invariants as named, enforced contracts to prevent architectural decay.

handling-failures
Enforce loud, explicit failure handling—no silent swallows, typed errors, or undocumented degradation.

diagramming-processes
Diagram business processes, workflows, and system interactions as maintainable source code.

documenting-contracts
Document HTTP APIs, message contracts, and file formats with exhaustive wire-level detail at the right abstraction level.