caveman-discover
juliusbrussee/caveman
Discover and label LLM workflows in your codebase for Caveman Cloud spend tracking by workflow.
What is caveman-discover?
Identifies every LLM workflow (jobs, handlers, scheduled tasks, CLI commands) in a repository and proposes workflow labels for Caveman Cloud cost attribution. Use this to break down LLM spending by job instead of treating all traffic as one bucket.
- Inventories HTTP handlers, scheduled jobs, CLI commands, and agents that call LLMs
- Proposes workflow names following slug grammar (lowercase, 1–96 chars)
- Wires labels via the lightest available mechanism (SDK options, headers, env vars, or flags)
- Verifies labeled paths work and don't break existing functionality
- Reports which workflows were labeled, where, and what remains unlabeled
How to install caveman-discover
npx skills add https://github.com/juliusbrussee/caveman --skill caveman-discover- Repository must have at least one LLM entry point (HTTP handler, scheduled job, CLI command, or agent)
- Caveman gateway already configured (see caveman-setup skill if not)
- User approval required before any code changes are made
How to use caveman-discover
- 1.Run the skill to inventory all LLM workflows in the repository
- 2.Review the proposed workflow table with names, locations, and labeling mechanisms
- 3.Approve or request changes to the workflow names and labeling strategy
- 4.The skill wires each label using the lightest mechanism available at that callsite (SDK options, headers, env vars)
- 5.Run an existing test or dev script to verify one labeled path still works
- 6.Check the Caveman Cloud dashboard at /activity?tab=workflows to see labeled spend appear as workflows run
Use cases
- Break down LLM API costs by business function (support replies, digest jobs, eval suites) instead of one lump sum
- Audit which workflows are consuming tokens in a multi-job codebase
- Prepare a repository for Caveman Cloud cost tracking before deploying
- Verify that all entry points to LLM calls are properly instrumented for observability
- Engineering teams using Caveman Cloud for LLM cost management
- DevOps and platform engineers optimizing spend visibility across multiple workflows
- Development teams wanting to track which jobs or features consume the most LLM tokens
caveman-discover FAQ
A workflow is one job a human would name: 'answer a support ticket', 'build the nightly digest', 'run the eval suite'. Multiple LLM calls inside the same request handler are one workflow; the same shared helper used by three jobs is three workflows (label at the callers, not the helper).
Yes, but renaming splits the spend history. Choose names that describe the job, not the technology, and that will age well.
Don't label it. Labels only travel on gateway traffic. The skill will list such callsites under 'not wired' in the report.
Run an existing test, dev script, or curl command that exercises one labeled path. The gateway rejects invalid labels with a 400 error; if you see that, fix the slug and try again.
Immediately for on-demand workflows; for scheduled jobs, spend appears when the schedule next fires. The skill will note this in the report.
Full instructions (SKILL.md)
Source of truth, from juliusbrussee/caveman.
name: caveman-discover description: > Find and label every LLM workflow in the repository so Caveman Cloud groups spend by workflow instead of one bucket. Use for "discover workflows" or breaking LLM spend down by workflow.
You are labeling this repository's LLM workflows for Caveman Cloud. A
workflow is a job the code performs — "answer a support ticket", "build the
nightly digest", "run the eval suite" — not a technology. Every gateway
request can carry a workflow label; unlabeled traffic all lands in one
unlabeled-workflow bucket. Your job: find the workflows, name them well,
wire the labels, and verify nothing broke.
This changes code, so it goes through the user's normal review: propose the table first, apply after the user agrees. Re-running on an already-labeled repo must change nothing (idempotent).
This skill is operator-invoked. An unlabeled-traffic Cave Plan observation is
review-only and does not create an advisory file, proposal, or Draft PR. Do not
infer that telemetry selected a callsite or authorized an edit. Independently
inventory the repository, present the labeling table, and wait for the user's
approval before changing code.
Step 1 — Inventory the workflows
Walk the repo from its entry points, not from its imports:
- HTTP/RPC handlers that call an LLM (directly or through layers)
- Scheduled jobs: cron definitions, queue consumers, workers, GitHub Actions that invoke LLM code
- CLI commands and scripts (
scripts/,bin/, package.json scripts) - Eval / test harnesses that burn real tokens
- Distinct agents or chains inside a framework (each LangGraph graph, each crew, each agent definition is usually its own workflow)
One workflow = one job a human would name. Ten callsites inside the same
request handler are one workflow; one shared llm.ts helper used by three
jobs is three workflows (label at the callers, never the shared helper).
Step 2 — Name them
Slug grammar (the gateway enforces this): lowercase [a-z0-9_-], 1–96 chars.
Name the job, not the tech:
- Good:
support-reply,nightly-digest,pr-review,eval-suite,onboarding-email - Bad:
openai-calls(tech),main(says nothing),SupportReply(invalid),johns-test-3(won't age)
Names are forever-ish — renaming later splits the spend history. When a job's
purpose isn't clear from the code, derive the slug from the file name and mark
it review in the table rather than inventing a purpose.
Step 3 — Propose, then apply
Present this table and ask to proceed:
| workflow | job | where | how it gets labeled |
|---|---|---|---|
| support-reply | answers inbound tickets | src/bot/reply.ts:41 | defaultHeaders on the reply client |
| nightly-digest | 02:00 summary job | jobs/digest.ts:12 | header on the digest client |
| eval-suite (review) | scripts/eval.ts:8 — purpose inferred from filename | scripts/eval.ts:8 | env override at invocation |
Then wire each label with the lightest mechanism available at that callsite:
- @caveman-ai/sdk / caveman_cloud SDK: per-trace
workflowoption, ordefaultWorkflowon the client a single-job service constructs. - Raw provider SDKs (OpenAI/Anthropic/LangChain/LiteLLM/Vercel): add
"x-cave-workflow": "<slug>"to the samedefaultHeaders/default_headers/extra_headersblock that already carriesx-cave-api-key. Shared client used by several jobs → pass the header per call (every SDK above accepts per-request header overrides), or give each job its own thin client. - Wrapped coding agents (
caveman wrap):--workflow <slug>flag orCAVE_WORKFLOW=<slug>env at the invocation site (cron line, CI step). - Raw HTTP: add the
x-cave-workflowheader to the request.
Label the callers, keep the diff minimal, match the repo's style. If a callsite is not routed through the Caveman gateway at all, don't label it — list it under "not wired" in the report (labels only travel on gateway traffic; wiring is the caveman-setup skill's job).
Step 4 — Verify
Run whatever the repo already uses to exercise one labeled path (a test, a
dev script, one curl). Then confirm: the request still succeeds (the gateway
rejects an invalid label with 400 cave_invalid_request_header — fix the slug
if so). Labeled spend appears on the dashboard at /activity?tab=workflows as
each workflow next runs; jobs on a schedule show up when the schedule fires,
and that's worth saying in the report rather than pretending they're live.
Step 5 — Report
## Workflows labeled
| workflow | job | where |
|---|---|---|
| support-reply | answers inbound tickets | src/bot/reply.ts:41 |
| nightly-digest | 02:00 summary job | jobs/digest.ts:12 |
Verified: <the labeled path you actually exercised, and what you observed>
Lands at: <DASHBOARD>/activity?tab=workflows — each row appears as that workflow
next runs. Anything still unlabeled shows as `unlabeled-workflow`.
Not wired (no gateway routing, so no label): <list or "none">
Marked review: <slugs whose purpose was inferred from filenames, or "none">
If you found no LLM entry points at all: say exactly that, and point at the
setup skill (<docs origin>/docs/agent-setup.md) instead of manufacturing a
table.
Related skills
More from juliusbrussee/caveman and the wider catalog.

caveman-evidence-review
Read-only review of Caveman Cloud evidence: costs, Cave Score, workflows, traces, and LLM spend optimization.

caveman-explore
Fast read-only repository explorer for code localization and orientation.

caveman-help
Quick-reference card for caveman modes, skills, and commands—terse communication style guide.

caveman-learn
Review token-cost analysis and apply consent-gated optimizations to reduce agent spending.

caveman-manage
Inspect and recommend on Caveman Cloud experiment lifecycle changes with safety gates.

caveman-optimize
Evaluate Caveman optimization observations with operator-chosen candidates and paired baseline testing.