context-engineering-collection
muratcankoylan/agent-skills-for-context-engineering
Comprehensive guidance for building production AI agent systems through context engineering, multi-agent coordination, and reliable operating loops.
What is context-engineering-collection?
A structured collection of skills for designing and optimizing AI agent systems, covering context management, multi-agent architectures, memory systems, and operational patterns. Use when building new agents, optimizing performance, debugging failures, or designing autonomous systems that require effective context handling and reliable execution loops.
- Provides architectural patterns for multi-agent coordination including supervisor, peer-to-peer, and hierarchical structures
- Teaches context optimization techniques including compression, observation masking, and prefix caching to manage token limits
- Covers memory system design from scratchpads to knowledge graphs and filesystem-based context loading
- Explains long-horizon prompting with pseudo-formal task briefs for autonomous agents
- Describes tool design principles emphasizing consolidation, error context, and clear namespacing
- Includes harness engineering patterns for reliable autonomous agents with metrics, logs, and rollback rules
How to install context-engineering-collection
npx skills add https://github.com/muratcankoylan/agent-skills-for-context-engineering --skill context-engineering-collectionHow to use context-engineering-collection
- 1.Review the foundational context engineering section to understand context degradation patterns and attention mechanisms
- 2.Select the architectural pattern that matches your system: supervisor, peer-to-peer, or hierarchical
- 3.Design your memory system based on your context requirements: scratchpad, vector RAG, knowledge graph, or filesystem-based
- 4.Apply context optimization techniques when approaching token limits, prioritizing tokens-per-task over tokens-per-request
- 5.Implement harness engineering patterns with locked metrics, durable logs, and human approval boundaries
- 6.Evaluate your agent system using deterministic checks and multi-dimensional rubrics before deploying
Use cases
- Building multi-agent systems that coordinate across specialized sub-agents with isolated contexts
- Optimizing long-running autonomous agents that need persistent memory and effective context management
- Debugging agent failures caused by context degradation, lost-in-middle phenomena, or information overload
- Designing evaluation frameworks and harnesses for production agent systems with deterministic checks
- Implementing filesystem-based memory systems for agents that need to manage effectively unlimited context
- AI engineers building production agent systems
- Developers optimizing multi-agent architectures
- Teams designing autonomous research or evaluation harnesses
- Engineers implementing memory and persistence layers for agents
- Researchers working on agent evaluation and harness design
context-engineering-collection FAQ
Context engineering encompasses the complete state available to the model at inference time—system instructions, tool definitions, retrieved documents, message history, and tool outputs—while prompt engineering typically focuses on crafting the text input. Effective context engineering means curating all available information for maximum signal-to-noise ratio.
Use multi-agent architectures when you need to isolate context for different specialized tasks, coordinate complex work decomposition, or run parallel portfolios. The primary benefit is context isolation rather than simulating organizational roles.
Language models show U-shaped attention curves where information at the beginning and end of context receives more attention than information in the middle. Address this through context compression, strategic partitioning across sub-agents, and placing critical information at context boundaries.
Use filesystem-based memory patterns for just-in-time context loading, implement structured summarization with explicit sections for files and decisions, apply context compression targeting tokens-per-task rather than tokens-per-request, and use prefix caching to reuse KV blocks across requests.
Latent Briefing compacts an orchestrator's long trajectory in a worker model's KV cache using task-guided attention, allowing workers to receive relevant latent state without full-text replay. Use it in orchestrator-worker systems where supervisors accumulate long trajectories but workers see only narrow text slices.
Full instructions (SKILL.md)
Source of truth, from muratcankoylan/agent-skills-for-context-engineering.
name: context-engineering-collection description: "A comprehensive collection of Agent Skills for context engineering, harness engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, evaluating, or debugging agent systems that require effective context management and reliable operating loops."
Agent Skills for Context Engineering
This collection provides structured guidance for building production-grade AI agent systems through effective context engineering.
When to Activate
Activate these skills when:
- Building new agent systems from scratch
- Optimizing existing agent performance
- Debugging context-related failures
- Designing multi-agent architectures
- Creating or evaluating tools for agents
- Implementing memory and persistence layers
- Designing autonomous research or evaluation harnesses
Skill Map
Foundational Context Engineering
Understanding Context Fundamentals Context is not just prompt text—it is the complete state available to the language model at inference time, including system instructions, tool definitions, retrieved documents, message history, and tool outputs. Effective context engineering means understanding what information truly matters for the task at hand and curating that information for maximum signal-to-noise ratio.
Recognizing Context Degradation Language models exhibit predictable degradation patterns as context grows: the "lost-in-middle" phenomenon where information in the center of context receives less attention; U-shaped attention curves that prioritize beginning and end; context poisoning when errors compound; and context distraction when irrelevant information overwhelms relevant content.
Architectural Patterns
Multi-Agent Coordination Production multi-agent systems converge on three dominant patterns: supervisor/orchestrator architectures with centralized control, peer-to-peer swarm architectures for flexible handoffs, and hierarchical structures for complex task decomposition. The critical insight is that sub-agents exist primarily to isolate context rather than to simulate organizational roles.
Long-Horizon Prompting Long-running autonomous agents and parallel orchestrations succeed or fail on the launch prompt. Pseudo-formal task briefs specify success predicates, non-counting outcomes, persistence rules with audit-gated return conditions, effort floors, diversity policies for parallel portfolios, and contamination guards, applying the discipline of formal verification linguistically to problems with no machine-checkable success condition.
Memory System Design Memory architectures range from simple scratchpads to sophisticated temporal knowledge graphs. Vector RAG provides semantic retrieval but loses relationship information. Knowledge graphs preserve structure but require more engineering investment. The file-system-as-memory pattern enables just-in-time context loading without stuffing context windows.
Filesystem-Based Context
The filesystem provides a single interface for storing, retrieving, and updating effectively unlimited context. Key patterns include scratch pads for tool output offloading, plan persistence for long-horizon tasks, sub-agent communication via shared files, and dynamic skill loading. Agents use ls, glob, grep, and read_file for targeted context discovery, often outperforming semantic search for structural queries.
Hosted Agent Infrastructure Background coding agents run in remote sandboxed environments rather than on local machines. Key patterns include pre-built environment images refreshed on regular cadence, warm sandbox pools for instant session starts, filesystem snapshots for session persistence, and multiplayer support for collaborative agent sessions. Critical optimizations include allowing file reads before git sync completes (blocking only writes), predictive sandbox warming when users start typing, and self-spawning agents for parallel task execution.
Tool Design Principles Tools are contracts between deterministic systems and non-deterministic agents. Effective tool design follows the consolidation principle (prefer single comprehensive tools over multiple narrow ones), returns contextual information in errors, supports response format options for token efficiency, and uses clear namespacing.
Operational Excellence
Context Compression When agent sessions exhaust memory, compression becomes mandatory. The correct optimization target is tokens-per-task, not tokens-per-request. Structured summarization with explicit sections for files, decisions, and next steps preserves more useful information than aggressive compression. Artifact trail integrity remains the weakest dimension across all compression methods.
Context Optimization Techniques include compaction (summarizing context near limits), observation masking (replacing verbose tool outputs with references), prefix caching (reusing KV blocks across requests), and strategic context partitioning (splitting work across sub-agents with isolated contexts).
Latent Briefing (KV Memory Sharing) Orchestrator-worker systems can compound tokens when supervisors accumulate long trajectories but workers see only narrow text slices. Latent Briefing compacts the orchestrator trajectory in the worker model's KV cache using task-guided attention (Attention Matching-style compaction) so workers receive relevant latent state without full-text replay when the stack exposes worker KV state and the models are compatible.
Evaluation Frameworks Production agent evaluation requires deterministic checks and multi-dimensional rubrics covering factual accuracy, completeness, tool efficiency, and process quality. Use model judges only after structure, evidence, and rubric math are valid; route judge design, pairwise comparison, and bias mitigation to Advanced Evaluation.
Harness Engineering Reliable autonomous agents need explicit operating loops around the model: locked metrics, editable surfaces, durable logs, novelty checks, rollback rules, and human approval boundaries. Harnesses prevent agents from weakening the evaluator, losing state across compaction, or turning ambiguous goals into unreviewable changes.
Self-Improvement Loops When the harness itself becomes the optimization target, a different discipline applies: recursive self-improvement, meta-harness search, failure-driven bounded self-edits, evolutionary scaffold search, and context mechanism evolution. The controlling constraints are empirical two-split acceptance gates, filesystem experience archives with raw traces, runtime-enforced constraints outside every editable surface, and diversity preservation to prevent collapse.
Development Methodology
Project Development Effective LLM project development begins with task-model fit analysis: validating through manual prototyping that a task is well-suited for LLM processing before building automation. Production pipelines follow staged, idempotent architectures (acquire, prepare, process, parse, render) with file system state management for debugging and caching. Structured output design with explicit format specifications enables reliable parsing. Start with minimal architecture and add complexity only when proven necessary.
Cognitive Architecture
BDI Mental States Belief-desire-intention modeling provides a formal way to translate structured external context into agent mental states. Use it for rational agency, explainability, and systems that need auditable links between beliefs, goals, and chosen actions.
Core Concepts
The collection is organized around four core themes. First, context fundamentals establish what context is, how attention mechanisms work, and why context quality matters more than quantity. Second, architectural patterns cover the structures and coordination mechanisms that enable effective agent systems. Third, operational excellence addresses optimization, evaluation, and harness reliability. Fourth, development methodology and cognitive architecture cover project execution and formal mental-state modeling.
Practical Guidance
Each skill can be used independently or in combination. Start with fundamentals to establish context management mental models. Branch into architectural patterns based on your system requirements. Reference operational skills when optimizing production systems.
The skills are platform-agnostic and work with Claude Code, Cursor, or any agent framework that supports custom instructions or skill-like constructs.
Integration
This collection integrates with itself—skills reference each other and build on shared concepts. The fundamentals skill provides context for all other skills. Architectural skills (multi-agent, memory, tools) can be combined for complex systems. Operational skills (optimization, evaluation) apply to any system built using the foundational and architectural skills.
References
Internal skills in this collection:
- context-fundamentals
- context-degradation
- context-compression
- multi-agent-patterns
- long-horizon-prompting
- memory-systems
- tool-design
- filesystem-context
- hosted-agents
- context-optimization
- latent-briefing
- evaluation
- advanced-evaluation
- harness-engineering
- self-improvement-loops
- project-development
- bdi-mental-states
External resources on context engineering:
- Research on attention mechanisms and context window limitations
- Production experience from leading AI labs on agent system design
- Framework documentation for LangGraph, AutoGen, and CrewAI
Skill Metadata
Created: 2025-12-20 Last Updated: 2026-07-11 Author: Agent Skills for Context Engineering Contributors Version: 2.5.0
Related skills
More from muratcankoylan/agent-skills-for-context-engineering and the wider catalog.

serenity-skill
Supply-chain bottleneck analysis for technology and advanced-manufacturing investment research.

printing-press
Generate ship-ready Go CLIs from OpenAPI, HAR, or Postman specs via structured research-generate-build-test loop.

printing-press-amend
Turn dogfood friction into PRs for published CLIs in the printing-press library.

printing-press-catalog
Browse and install pre-built Go CLIs for popular APIs from the catalog

printing-press-import
Import a published CLI from the public library into your local workspace, ready for polish or re-publish.

printing-press-output-review
Internal agentic review of printed CLI output for plausibility bugs that rule-based checks miss.