PluginBench
Skill
Review
Audit score 70

skill-comply

affaan-m/ecc

Measure whether agents actually follow skills, rules, and definitions across prompt strictness levels.

What is skill-comply?

skill-comply auto-generates behavioral specs from markdown files, runs agents under three prompt conditions (supportive to competing), and reports compliance rates with full tool-call timelines. Use it to verify that skills and rules are genuinely followed, not just documented.

  • Auto-generates expected behavioral sequences from .md files (skills, rules, agent definitions)
  • Creates three prompt variants with decreasing strictness to test prompt independence
  • Runs agents and captures tool-call traces via stream-json
  • Classifies tool calls against specs using LLM-based matching (not regex)
  • Checks temporal ordering of steps deterministically
  • Generates self-contained reports with specs, prompts, compliance scores, and tool timelines

How to install skill-comply

npx skills add null --skill skill-comply
Prerequisites
  • Python environment with uv
  • Access to Claude via claude -p command
  • Target markdown file (skill, rule, or agent definition)
Claude Code
Cursor
Windsurf
Cline

How to use skill-comply

  1. 1.Run `/skill-comply <path>` or ask the agent 'is this rule actually being followed?'
  2. 2.For a full compliance run: `uv run python -m scripts.run ~/.claude/rules/common/testing.md`
  3. 3.For a dry run (no cost): `uv run python -m scripts.run --dry-run ~/.claude/skills/search-first/SKILL.md`
  4. 4.Optionally specify models: `uv run python -m scripts.run --gen-model haiku --model sonnet <path>`
  5. 5.Review the generated report: specs, scenario prompts, compliance scores, and tool-call timelines

Use cases

Good for
  • Verify a new rule (e.g., testing.md) is actually followed by agents, not just documented
  • Test whether a skill like search-first remains effective when prompts don't explicitly support it
  • Check agent invocation compliance after modifying agent definitions
  • Periodically audit rule adherence as part of quality maintenance
  • Identify which behavioral steps have low compliance for hook promotion recommendations
Who it's for
  • Agent developers and maintainers
  • Team leads verifying rule enforcement
  • QA engineers testing agent behavior consistency
  • Anyone adding or updating skills, rules, or agent definitions

skill-comply FAQ

What's the difference between a full run and a dry run?

A dry run generates the spec and scenario prompts without executing the agent, so it costs nothing. A full run executes the agent under all three prompt conditions and measures compliance.

Does this measure whether agents follow a rule even if the prompt doesn't mention it?

Yes. That's the core concept: prompt independence. It tests compliance across three strictness levels—supportive, neutral, and competing—to see if the rule holds regardless of prompt framing.

What targets can I measure compliance for?

Skills (SKILL.md files), rules (rules/common/*.md), and agent definitions (agents/*.md). Internal workflow verification for agents is not yet supported.

What does the report include?

The auto-generated spec, the three scenario prompts used, compliance scores per scenario, and detailed tool-call timelines with LLM classification labels. It's self-contained and shareable.

Can I use different models for generation and execution?

Yes. Use `--gen-model` to specify the model for spec/scenario generation and `--model` for agent execution.

Full instructions (SKILL.md)

Source of truth, from affaan-m/ecc.


name: skill-comply description: Visualize whether skills, rules, and agent definitions are actually followed — auto-generates scenarios at 3 prompt strictness levels, runs agents, classifies behavioral sequences, and reports compliance rates with full tool call timelines metadata: origin: ECC tools: Read, Bash

skill-comply: Automated Compliance Measurement

Measures whether coding agents actually follow skills, rules, or agent definitions by:

  1. Auto-generating expected behavioral sequences (specs) from any .md file
  2. Auto-generating scenarios with decreasing prompt strictness (supportive → neutral → competing)
  3. Running claude -p and capturing tool call traces via stream-json
  4. Classifying tool calls against spec steps using LLM (not regex)
  5. Checking temporal ordering deterministically
  6. Generating self-contained reports with spec, prompts, and timelines

Supported Targets

  • Skills (skills/*/SKILL.md): Workflow skills like search-first, TDD guides
  • Rules (rules/common/*.md): Mandatory rules like testing.md, security.md, git-workflow.md
  • Agent definitions (agents/*.md): Whether an agent gets invoked when expected (internal workflow verification not yet supported)

When to Activate

  • User runs /skill-comply <path>
  • User asks "is this rule actually being followed?"
  • After adding new rules/skills, to verify agent compliance
  • Periodically as part of quality maintenance

Usage

# Full run
uv run python -m scripts.run ~/.claude/rules/common/testing.md

# Dry run (no cost, spec + scenarios only)
uv run python -m scripts.run --dry-run ~/.claude/skills/search-first/SKILL.md

# Custom models
uv run python -m scripts.run --gen-model haiku --model sonnet <path>

Key Concept: Prompt Independence

Measures whether a skill/rule is followed even when the prompt doesn't explicitly support it.

Report Contents

Reports are self-contained and include:

  1. Expected behavioral sequence (auto-generated spec)
  2. Scenario prompts (what was asked at each strictness level)
  3. Compliance scores per scenario
  4. Tool call timelines with LLM classification labels

Advanced (optional)

For users familiar with hooks, reports also include hook promotion recommendations for steps with low compliance. This is informational — the main value is the compliance visibility itself.