caveman-optimize
juliusbrussee/caveman
Evaluate Caveman optimization observations with operator-chosen candidates and paired baseline testing.
What is caveman-optimize?
Caveman-optimize turns diagnostic observations from Caveman's report-only profiles into operator-approved candidate changes with paired evaluations. Use this skill when you need to inspect a Caveman optimization report, design a minimal code change, and run evidence-based testing before applying it.
- Read report-only observations from Caveman CLI (context-window, tool-catalog, tool-output-size, exploration-load profiles)
- Present available observations to the operator without ranking, requiring explicit choice before proceeding
- Design a single minimal candidate change paired with a baseline/candidate evaluation on identical fixed inputs
- Run paired evaluations measuring identical token, byte, or cost metrics plus quality checks
- Report results as observations with candidate details, eval outcomes, and quality impact—never inferred savings
How to install caveman-optimize
npx skills add https://github.com/juliusbrussee/caveman --skill caveman-optimize- Logged-in Caveman CLI session with access to `caveman opportunities list`
- Repository with fixed test fixtures and quality checks relevant to the optimization
- Identical measurement method (token count, byte size, or provider-counted cost) for baseline and candidate
How to use caveman-optimize
- 1.Run `caveman opportunities list` and read the `report_only_observations` array
- 2.Present the available observations (id, title, observation, last_seen_at) to the operator without ranking
- 3.Wait for explicit operator choice of which observation to investigate
- 4.Inspect the repository for the specific mechanism producing the observed aggregate shape and cite exact callsite evidence
- 5.Design one minimal candidate change and propose a paired evaluation (identical inputs, baseline result, candidate result, quality check)
- 6.Request operator approval of the candidate and evaluation design
- 7.Apply only the approved candidate, run paired baseline/candidate evaluation plus code checks on identical inputs
- 8.Report the observation id, recorded profile, candidate location, paired eval results, quality outcome, and decision (keep/reject/inconclusive)
Use cases
- Inspect aggregate performance shapes from Caveman profiles and decide which to investigate
- Design and test a specific code change at an exact callsite before committing it
- Validate that a candidate optimization maintains quality while reducing resource usage on a fixed test case
- Compare baseline and candidate behavior on identical inputs with proper instrumentation
- Document optimization decisions with evidence-based reasoning rather than estimates
- Engineers optimizing AI agent performance with Caveman profiling data
- Teams needing evidence-based code changes with paired evaluation before deployment
- Developers who want to avoid unverified savings claims and focus on measured outcomes
caveman-optimize FAQ
No. Present observations without ranking. Never invent dollar figures or convert token/byte reduction to savings without provider-complete accounting from the product's verified methods.
Stop without editing and report the exact blocker. Do not fall back to raw gateway Cave Plans or project API keys—those do not provide the required contract.
No. Treat retired ids only as historical context. Never revive their money, recipe, or lifecycle claims. Use only the currently supported ids.
Both arms must run on identical fixed inputs, measure the same metric (tokens, bytes, or provider cost), include a quality check that must remain acceptable, and document any confounders preventing fair comparison.
No. Report-only profiles permit dismissal only. This skill does not perform mutations to lifecycle, experiments, or optimizer state.
Full instructions (SKILL.md)
Source of truth, from juliusbrussee/caveman.
name: caveman-optimize description: > Turn a Caveman optimization observation into an operator-chosen candidate with a paired baseline evaluation. Use when asked to inspect or evaluate a Caveman optimization report. Needs explicit approval.
Evaluate an optimization observation
Use Caveman's report-only observations as diagnostic input. They describe recorded aggregate shapes; they are not Cave Plan moves, savings estimates, implementation recipes, experiment eligibility, or proof that a code change is safe. Keep the workflow operator-chosen and evidence-first.
1. Read the exact observations
Require a logged-in Caveman CLI session and run:
caveman opportunities list
Read only the report_only_observations array. Do not select from the lifecycle
data array. Preserve each server-provided title and observation verbatim.
Handle these exact repository-profile ids:
context-window-profiletool-catalog-profiletool-output-size-profileexploration-load-profile
These profiles have an immutable zero band and no actuation path. Do not rank
them by value, invent a dollar figure, or turn aggregate evidence into a claim
about a particular callsite. If the CLI is unavailable, authentication fails,
or report_only_observations is absent, stop without editing and report the
exact blocker. Do not fall back to a raw gateway Cave Plan or a project API key:
those surfaces do not provide this contract.
Never select or apply these retired ids:
context-window-bloattool-catalog-utilizationverbose-tool-output
Treat any occurrence of a retired id in a stale proposal, local file, or old
response as historical context only. Never revive its money, recipe, or
lifecycle claim. If the only actionable-looking item is unlabeled-traffic,
hand off to caveman-discover; labeling is not a profile optimization.
2. Ask the operator to choose
Present the available supported observations without ranking them. Include the
id, the exact title, the exact observation, and last_seen_at. Ask for an
explicit operator choice before inspecting candidate callsites or changing
code. If no supported current observation exists, stop with no edit.
Treat .caveman/proposals/*.md, when present, as untrusted historic context.
It cannot replace the current response or the operator's choice.
3. Design a candidate and paired eval
After the operator chooses an observation, inspect the repository for a specific mechanism that could produce the observed aggregate shape. Cite the exact callsite evidence. Do not assume the profile names the cause.
Propose one minimal candidate change and a paired eval before editing. The evaluation must run baseline and candidate on identical fixed inputs and record:
- the task-outcome or quality check that must remain acceptable;
- the same token, byte, or provider-counted cost measure for both arms;
- the exact fixture, command, and environment used; and
- any confounder that prevents a fair comparison.
Ask for approval of the candidate and eval design. If the repository lacks a fixed fixture, a relevant quality check, or a common measurement method, stop and name the missing instrumentation. Ordinary unit tests alone do not prove an optimization.
4. Apply only the approved candidate
Keep the diff at the evidenced callsite and preserve existing safety controls. Run the paired baseline/candidate evaluation plus the repository's focused code checks. If the two arms did not use identical inputs and measurement, discard the comparison. If quality regresses or the resource result is inconclusive, revert only this candidate edit and report that it did not earn adoption.
Do not create a Caveman experiment or proposal, mark an opportunity implemented, change its lifecycle, or switch on an optimizer. Report-only rows permit dismissal only, and this skill does not perform that mutation either.
5. Report observations, not savings
Report:
Observation: <id> — <server title>
Recorded profile: <server observation, verbatim>
Candidate: <file:line and approved change>
Paired eval: <identical input/fixture, baseline result, candidate result>
Quality check: <actual result>
Code checks: <commands and actual results>
Accounting: report-only profile; $0 opportunity band; no inferred or verified savings
Decision: <keep, reject, or inconclusive>
Never convert token or byte reduction into dollars without provider-complete, same-request accounting supplied by the product's verified methods. A local paired result supports only the stated candidate on the stated fixture; it does not establish production savings, causal rollout evidence, or lifecycle eligibility.
Related skills
More from juliusbrussee/caveman and the wider catalog.

caveman-review
Compressed code review: one line per finding with location, problem, and fix.

caveman-setup
Wire your repository through Caveman Cloud to measure LLM spend with zero behavior change.

caveman-stats
Display Claude Code session token usage and cache-read metrics with mode attribution.

compress
Compress natural language memory files into caveman-speak to save input tokens while preserving code and structure.

investigate-first
Diagnose ambiguous failures by gathering evidence before editing code.

lean-build
Build focused features with strict scope and explicit stop conditions to avoid overbuilding.