PluginBench
Skill
Pass
Audit score 90

session-guard

wshobson/agents

Monitor session health and prevent context corruption through behavioral self-enforcement.

What is session-guard?

Session Guard detects and prevents silent degradation in long agent sessions by monitoring tool call counts, behavioral drift, and context compaction events. Use it when sessions exceed 40 tool calls, conventions drift, or after context compaction to maintain rule adherence and output quality without external infrastructure.

  • Monitor session health via tool call count thresholds (Yellow zone: 40-60 calls, Red zone: 60+)
  • Detect behavioral drift by tracking contradictions, style/naming convention changes, and unexpected file content
  • Anchor critical rules through context compaction by re-reading project files and reciting active rules
  • Checkpoint progress and create handoff documents before splitting sessions
  • Prevent silent context loss by distinguishing what survives compaction (recent messages, skill body, git state) from what gets dropped (early decisions, verbal rules, tool outputs)

How to install session-guard

npx skills add https://github.com/wshobson/agents --skill session-guard
Claude Code
Cursor
Windsurf
Cline

How to use session-guard

  1. 1.Monitor your tool call count: enter Yellow zone at 40 calls, Red zone at 60+ calls
  2. 2.In Yellow zone: checkpoint progress in one paragraph, recite 3-5 critical active rules aloud, assess if task is nearly complete or needs splitting
  3. 3.In Red zone: stop making tool calls, re-read project rules files (don't trust memory), write state to handoff document, create fresh session with handoff
  4. 4.After any context compaction event: immediately re-read the project's rules file, recite critical rules in your response, verify next action matches those rules before executing
  5. 5.For compaction-safe workflows: keep all critical instructions in files (CLAUDE.md, CONTEXT.md) rather than relying on conversation context

Use cases

Good for
  • Long multi-step coding tasks approaching or exceeding 40 tool calls to prevent rule drift and quality degradation
  • Projects with strict naming conventions or architectural patterns that need reinforcement as sessions grow
  • Recovery after context compaction events to re-anchor critical rules before continuing work
  • Unbounded task scopes that need deliberate splitting into separate sessions before corruption occurs
  • Complex refactoring or migration tasks where early decisions must be verified before proceeding
Who it's for
  • Coding agents (Claude Code, Cursor) running extended sessions
  • Teams managing projects with strict architectural or naming conventions
  • Developers working on complex multi-step tasks requiring consistent rule adherence
  • Anyone using agents in harnesses where context compaction occurs

session-guard FAQ

How do I know if context compaction has occurred?

Compaction is often silent, but signs include sudden loss of earlier context, unexpected behavior changes, or explicit /compact commands. When uncertain, re-read your project rules files immediately—don't assume memory is reliable after 40+ tool calls.

What's the difference between Yellow and Red zones?

Yellow zone (40-60 calls) is preventive: checkpoint, recite rules, and prepare to split if needed. Red zone (60+ calls or drift detected) is critical: stop work, verify rules, and create a handoff for a fresh session.

Why can't I just use hooks or system prompts to prevent this?

Hooks and post-compaction injections don't work because the compaction summary creates narrative momentum that agents follow instead. This skill uses behavioral self-enforcement—disciplined monitoring and re-reading—which works in every harness.

What survives context compaction and what gets lost?

Survives: most recent user messages, the current skill body (capped at 5K tokens), git status, and file contents read after compaction. Lost: early decisions, rules stated only verbally, and tool output context. Keep critical rules in files, not conversation.

Should I re-read all project files to be safe?

No—that wastes tool calls. Be targeted: re-read only the rules file (CLAUDE.md, CONTEXT.md) and specific source files when you detect drift or after compaction. Trust recent reads; verify only when uncertain.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: session-guard description: >- Use when working on complex multi-step tasks, when a session is getting long (40+ tool calls), when the agent starts ignoring rules it followed earlier, when conventions drift, when output quality seems to degrade, or after any context compaction event. Prevents long-session corruption AND context compaction amnesia through behavioral self-enforcement.

Session Guard

Overview

Long sessions corrupt silently. Context compaction drops instructions unannounced. Hooks don't fix it (confirmed by multiple developers on GitHub issues #19471, #9796, #64171). This skill prevents both through behavioral self-enforcement: monitor health, anchor critical rules through compaction, split before damage occurs. No packages, no databases - pure behavioral enforcement that works in every harness.

When to Use

  • Session exceeds 40 tool calls
  • Agent contradicts earlier decisions
  • Style/naming conventions start drifting
  • After any context compaction event
  • Task scope growing unbounded

Health Signals

SignalThresholdAction
Tool call count>40YELLOW: checkpoint + recite critical rules
Tool call count>60RED: split or compact with anchor
Agent contradicts earlier decisionAnyVERIFY: re-read source of truth
Style/naming driftAnyRECITE: state the active rules aloud
File read returns unexpected contentAnyRE-READ: don't trust cached state
Task scope growing unboundedContinuousSPLIT: one task per session

Protocol

Green Zone (0-40 tool calls)

Normal operation. No intervention needed.

Yellow Zone (40-60 tool calls)

  1. CHECKPOINT - summarize progress in one paragraph
  2. RECITE - state the 3-5 most critical active rules aloud: "Active rules: [naming convention], [file structure], [error handling pattern], [testing requirement]"
  3. ASSESS - almost done? Push through. Not done? Prepare split.
  4. REDUCE - no exploratory reads. Only targeted operations.

Red Zone (60+ tool calls OR drift signal)

  1. STOP - do not make more tool calls
  2. VERIFY - re-read project rules (don't trust memory)
  3. CHECKPOINT - write state to handoff document
  4. SPLIT - create handoff, suggest fresh session

Context Anchoring (anti-compaction)

When compaction has occurred (sudden loss of earlier context, or after /compact):

  1. RE-READ the project's rules file immediately
  2. RECITE the 3-5 critical rules aloud in your response
  3. VERIFY your planned next action matches those rules before executing
  4. If uncertain about ANY prior decision, RE-READ the source file - don't guess

What survives compaction:

  • Most recent user messages (high priority)
  • Currently-invoked skill body (capped at 5K tokens, oldest dropped first)
  • Git status and project structure
  • File contents read AFTER compaction

What gets LOST in compaction:

  • Decisions made early in conversation
  • Architectural rules stated only verbally (not in files)
  • Context from tool outputs (file reads, command outputs)

Compaction-Safe Pattern

Keep critical instructions in FILES (CLAUDE.md, CONTEXT.md), NOT in conversation. If a rule matters, it must live in a file the agent can re-read - not in something agreed on earlier.

Common Mistakes

  • Trusting that you remember the rules after 50+ tool calls (you don't - re-read)
  • Re-reading EVERYTHING to be safe (wastes tool calls - be targeted)
  • Feeling fine therefore assuming context is fine (compaction is SILENT)
  • Splitting AFTER noticing problems (split BEFORE - prevention, not recovery)

Why This Matters

Context compaction is the #1 unsolved platform problem in 2026. Hooks don't fix it (confirmed: the agent ignores post-compaction injections because the compaction summary creates narrative momentum). This skill is the lightweight behavioral countermeasure: no infrastructure, no packages - disciplined self-monitoring that works in every harness.