PluginBench
Skill
Pass
Audit score 90

writing-unit-tests

riekelt/principal-engineer

Behavior-first unit testing: one claim per test, deterministic setup, mocks only at external boundaries.

What is writing-unit-tests?

A guide to writing maintainable unit tests that verify contracts rather than implementations. Use when creating new tests, refactoring existing ones, or fixing flaky or unreadable tests. Emphasizes clear naming, deterministic execution, and strategic mocking.

  • Test through public contracts, not implementation details, so refactors don't break passing tests
  • Name tests as behavioral claims (subject, scenario, outcome) that serve as documentation
  • Arrange-act-assert structure with no branching or logic inside test bodies
  • Mock only external boundaries you don't own; use real collaborators for your own code
  • Eliminate flakiness by injecting time/randomness and avoiding sleeps or real I/O
  • Assert outcomes with values, not absence of exceptions, and test failure paths as first-class subjects

How to install writing-unit-tests

npx skills add https://github.com/riekelt/principal-engineer --skill writing-unit-tests
Prerequisites
  • Familiarity with the principal-engineering skill
  • Understanding of your testing framework (Jest, pytest, etc.)
Claude Code
Cursor
Windsurf
Cline

How to use writing-unit-tests

  1. 1.Identify the public contract (inputs, outputs, side effects) of the unit under test
  2. 2.Write a test name that states the claim: subject, scenario, and expected outcome
  3. 3.Arrange test state using builders or role-named fixtures, keeping setup visible
  4. 4.Act by calling the unit once through its public interface
  5. 5.Assert the outcome with literal expected values (derived independently, not from production code)
  6. 6.Review for common mistakes: mirror tests, mega-tests, mocks of your own code, sleeps, or conditional assertions

Use cases

Good for
  • Writing a new test file for a module with clear behavioral contracts
  • Refactoring an existing test suite to eliminate flaky sleeps and improve readability
  • Fixing a test that breaks after a refactor by distinguishing contract from implementation
  • Debugging a test that mocks internal methods instead of external boundaries
  • Creating deterministic tests for time-dependent or randomness-dependent code
Who it's for
  • Software engineers writing or maintaining unit tests
  • Teams adopting behavior-driven testing practices
  • Engineers refactoring legacy test suites
  • Anyone debugging flaky or hard-to-read tests

writing-unit-tests FAQ

When should I mock versus use a real collaborator?

Mock boundaries you don't own (network, filesystem, clock, third-party services). Use real collaborators for code you own within the unit's reach. For your own wrappers around external resources, mock at the seam where your code last touches the unowned resource.

How do I test time-dependent or async code without sleeps?

Inject the clock or time provider and control it in the test. For async work, poll the observable outcome with a deadline rather than sleeping for a fixed duration.

What's wrong with deriving expected values from production code?

It creates a mirror test that stays green even when the implementation is wrong. Expected values should be literals worked out independently—by hand, from a spec, or from real data—with the derivation documented in a comment.

How do I fix a test that's mostly mock wiring?

Either widen the unit to something with real behavior, or accept that this seam needs an integration test instead and document which one.

Why is a test with logic a problem?

A test with branching, loops, or conditional logic can itself be wrong, and failures become ambiguous. Keep tests simple and logic-free; move generation and iteration into builders and helpers.

Full instructions (SKILL.md)

Source of truth, from riekelt/principal-engineer.


name: writing-unit-tests description: "Use when writing or refactoring unit tests - a new test file, added cases, a flaky test, an unreadable one. Encodes behavior-first testing: one behavior per test, names that state the claim, deterministic setup, mocks only at boundaries you do not own. Use whenever a test is being written, even a quick one, and whenever a test needs a sleep, a mock of your own code, or a copy of the implementation's math."

Writing unit tests

REQUIRED BACKGROUND: the principal-engineering skill. testing-changes governs which tests a change owes; this skill is the craft of the tests themselves.

Overview

A unit test is a behavioral claim with a name, read by the next engineer during a red build. Core principle: test the contract, not the implementation. The name carries the claim, and the test stays simple enough that it cannot itself be wrong.

Contract over implementation

  • Test through the public contract of the unit. A refactor that preserves behavior should not break tests; when it does, the tests were asserting the implementation.
  • Do not assert call sequences, internal state, or that method A called method B, unless the interaction IS the contract (a required side effect on a boundary).
  • Never derive the expected value from the production arithmetic, neither by reimplementing the formula nor by invoking the shared helper that computes it. Expected values are literals worked out independently (by hand, from a spec, from real data), with the derivation in a comment.

One behavior per test, named as the claim

  • One behavior per test; splitting is cheaper than archaeology on a multi-assert failure.
  • The name states subject, scenario, and expected outcome: expired_token_is_rejected_with_401, not test_auth_3. Test names describe behavior, state transitions, and invariants; never delivery order, ticket keys, or phases.
  • Arrange, act, assert, visibly and in that order. No branching, loops, or logic in a test: a test with logic needs its own test. Shared setup earns a builder or a role-named fixture; a mystery blob fixture hides which arranged fact the assertion depends on. Generation and iteration live in builders and helpers, not in the test body. Property-based tests are the accepted form for invariants and follow their framework's shape; example-based tests stay logic-free.

Determinism

  • No real time, real randomness, real network, or real filesystem inside a unit test: inject the clock, seed or inject the randomness, mock the boundary.
  • No sleeps. Waiting for async work is condition-based (poll the observable outcome with a deadline), never duration-based.
  • A flaky test is red: fix it or quarantine it visibly with an owner (see the red-test rule in testing-changes); re-running until green is silencing a detector.

Mocks are assumptions

  • Mock the boundaries you do not own (network, clock, filesystem, third-party services); prefer real collaborators for code you do own within the unit's reach. For owned wrappers around unowned resources (your repository class fronting the database), mock at the seam where owned code last touches the unowned resource, and keep the test data role-named and visible either way. Every mock hardcodes an assumption about a contract; a stale mock is how a suite stays green while the real integration is broken.
  • When a test is mostly mock wiring, it is testing the mocks. Either widen the unit to something with real behavior or accept that this seam needs an integration test instead (and say which).
  • Fixtures are labeled snapshots of reality: minimal, role-named for their part in the scenario, updated deliberately when the contract changes, never regenerated blindly to make red go green.

Assertions and failure paths

  • Assert outcomes with values, not absence of exceptions. "It did not throw" claims almost nothing.
  • Failure paths are first-class test subjects: the typed failure surfaces, the degraded mode is entered loudly, the guard actually guards (see handling-failures).
  • Tests themselves follow the no-silent-swallows contract: no catch-and-ignore in test code, no conditional assertions that skip silently when a precondition is absent. A test that cannot run must fail or be visibly skipped with the reason.

Common mistakes

  • The mirror test: reimplementing the production logic to compute the expectation.
  • The mock echo chamber: mocking your own class and asserting the mock.
  • The mega-test: twelve assertions, one name, no way to know which claim broke.
  • Shared mutable fixtures that make test order matter; every test builds or receives its own state.
  • The sleep that "fixes" flakiness by making it rarer.
  • Green-checking the fixture: editing expected values to match actual output without deriving why the new value is right.