PluginBench
Skill
Pass
Audit score 90

prompt-engineering-patterns

wshobson/agents

Master advanced prompt engineering techniques to maximize LLM performance and reliability.

What is prompt-engineering-patterns?

This skill teaches advanced prompt engineering patterns for production LLM applications, including few-shot learning, chain-of-thought reasoning, structured outputs, and prompt optimization. Use it when you need to design, optimize, or debug prompts for consistent and reliable LLM performance.

  • Few-shot learning with semantic similarity and dynamic example selection
  • Chain-of-thought and self-consistency reasoning patterns
  • Structured outputs using JSON mode and Pydantic schema enforcement
  • Iterative prompt optimization with A/B testing and performance metrics
  • Reusable prompt templates with variable interpolation and conditional sections
  • System prompt design for specialized AI assistants

How to install prompt-engineering-patterns

npx skills add https://github.com/wshobson/agents --skill prompt-engineering-patterns
Claude Code
Cursor
Windsurf
Cline

How to use prompt-engineering-patterns

  1. 1.Define your task and identify which prompt pattern applies (few-shot, chain-of-thought, structured output, etc.)
  2. 2.Create a prompt template using ChatPromptTemplate with variable placeholders
  3. 3.If using structured outputs, define a Pydantic schema for the expected response format
  4. 4.Initialize your LLM with structured output enforcement if needed
  5. 5.Test your prompt on diverse, representative inputs and measure performance metrics
  6. 6.Iterate on the prompt based on accuracy, consistency, and token usage results
  7. 7.Version control your final prompt and document the reasoning behind its structure

Use cases

Good for
  • Designing SQL query generation prompts with structured output validation
  • Building multi-turn conversation templates with role-based composition
  • Optimizing customer service chatbot prompts through A/B testing and consistency measurement
  • Creating few-shot learning systems for domain-specific tasks with dynamic example retrieval
  • Implementing chain-of-thought reasoning for complex problem-solving tasks
Who it's for
  • LLM application developers
  • Prompt engineers optimizing production systems
  • AI product teams building specialized assistants
  • Data scientists fine-tuning model behavior
  • Teams building reliable, consistent AI workflows

prompt-engineering-patterns FAQ

When should I use few-shot vs. zero-shot prompting?

Use zero-shot for simple tasks where the model has strong prior knowledge. Use few-shot when you need consistent formatting, domain-specific behavior, or handling of edge cases. Few-shot is more reliable but uses more tokens.

How many examples should I include in few-shot prompting?

Start with 3-5 examples and increase only if performance plateaus. Balance example count against context window constraints. Quality of examples matters more than quantity.

What's the difference between chain-of-thought and tree-of-thought?

Chain-of-thought elicits linear step-by-step reasoning. Tree-of-thought explores multiple reasoning paths and selects the best one. Use tree-of-thought for complex problems requiring exploration.

How do I handle malformed JSON outputs?

Use Pydantic schema enforcement with structured output mode to guarantee valid JSON. Add explicit error handling and retry logic for edge cases. Include format examples in your prompt.

Should I version control my prompts?

Yes, treat prompts as code. Version them alongside your application code, document changes, and track performance metrics for each version to understand impact of modifications.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: prompt-engineering-patterns description: >- This skill should be used when the user asks to "optimize a prompt", "improve prompt performance", "design a prompt template", "write better prompts", "debug prompt issues", "use chain-of-thought", "structured prompting", "few-shot prompting", or wants to apply advanced prompt engineering patterns for production LLM applications.

Prompt Engineering Patterns

Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability.

When to Use This Skill

  • Designing complex prompts for production LLM applications
  • Optimizing prompt performance and consistency
  • Implementing structured reasoning patterns (chain-of-thought, tree-of-thought)
  • Building few-shot learning systems with dynamic example selection
  • Creating reusable prompt templates with variable interpolation
  • Debugging and refining prompts that produce inconsistent outputs
  • Implementing system prompts for specialized AI assistants
  • Using structured outputs (JSON mode) for reliable parsing

Core Capabilities

1. Few-Shot Learning

  • Example selection strategies (semantic similarity, diversity sampling)
  • Balancing example count with context window constraints
  • Constructing effective demonstrations with input-output pairs
  • Dynamic example retrieval from knowledge bases
  • Handling edge cases through strategic example selection

2. Chain-of-Thought Prompting

  • Step-by-step reasoning elicitation
  • Zero-shot CoT with "Let's think step by step"
  • Few-shot CoT with reasoning traces
  • Self-consistency techniques (sampling multiple reasoning paths)
  • Verification and validation steps

3. Structured Outputs

  • JSON mode for reliable parsing
  • Pydantic schema enforcement
  • Type-safe response handling
  • Error handling for malformed outputs

4. Prompt Optimization

  • Iterative refinement workflows
  • A/B testing prompt variations
  • Measuring prompt performance metrics (accuracy, consistency, latency)
  • Reducing token usage while maintaining quality
  • Handling edge cases and failure modes

5. Template Systems

  • Variable interpolation and formatting
  • Conditional prompt sections
  • Multi-turn conversation templates
  • Role-based prompt composition
  • Modular prompt components

6. System Prompt Design

  • Setting model behavior and constraints
  • Defining output formats and structure
  • Establishing role and expertise
  • Safety guidelines and content policies
  • Context setting and background information

Quick Start

from langchain_anthropic import ChatAnthropic
from langchain_core.prompts import ChatPromptTemplate
from pydantic import BaseModel, Field

# Define structured output schema
class SQLQuery(BaseModel):
    query: str = Field(description="The SQL query")
    explanation: str = Field(description="Brief explanation of what the query does")
    tables_used: list[str] = Field(description="List of tables referenced")

# Initialize model with structured output
llm = ChatAnthropic(model="claude-sonnet-5")
structured_llm = llm.with_structured_output(SQLQuery)

# Create prompt template
prompt = ChatPromptTemplate.from_messages([
    ("system", """You are an expert SQL developer. Generate efficient, secure SQL queries.
    Always use parameterized queries to prevent SQL injection.
    Explain your reasoning briefly."""),
    ("user", "Convert this to SQL: {query}")
])

# Create chain
chain = prompt | structured_llm

# Use
result = await chain.ainvoke({
    "query": "Find all users who registered in the last 30 days"
})
print(result.query)
print(result.explanation)

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

  1. Be Specific: Vague prompts produce inconsistent results
  2. Show, Don't Tell: Examples are more effective than descriptions
  3. Use Structured Outputs: Enforce schemas with Pydantic for reliability
  4. Test Extensively: Evaluate on diverse, representative inputs
  5. Iterate Rapidly: Small changes can have large impacts
  6. Monitor Performance: Track metrics in production
  7. Version Control: Treat prompts as code with proper versioning
  8. Document Intent: Explain why prompts are structured as they are

Common Pitfalls

  • Over-engineering: Starting with complex prompts before trying simple ones
  • Example pollution: Using examples that don't match the target task
  • Context overflow: Exceeding token limits with excessive examples
  • Ambiguous instructions: Leaving room for multiple interpretations
  • Ignoring edge cases: Not testing on unusual or boundary inputs
  • No error handling: Assuming outputs will always be well-formed
  • Hardcoded values: Not parameterizing prompts for reuse

Success Metrics

Track these KPIs for your prompts:

  • Accuracy: Correctness of outputs
  • Consistency: Reproducibility across similar inputs
  • Latency: Response time (P50, P95, P99)
  • Token Usage: Average tokens per request
  • Success Rate: Percentage of valid, parseable outputs
  • User Satisfaction: Ratings and feedback