airflow-dag-patterns
wshobson/agents
Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment.
What is airflow-dag-patterns?
This skill provides patterns and best practices for creating production-ready Apache Airflow DAGs. Use it when designing data pipelines, orchestrating workflows, scheduling batch jobs, or setting up Airflow in production environments.
- Design idempotent, atomic, and observable DAGs following core principles
- Implement task dependencies with linear, fan-out, fan-in, and complex patterns
- Configure default arguments for retries, exponential backoff, and error handling
- Use TaskFlow API for cleaner code and automatic XCom management
- Set up sensors with reschedule mode to free up workers
- Test DAGs locally with unit and integration tests
How to install airflow-dag-patterns
npx skills add https://github.com/wshobson/agents --skill airflow-dag-patterns- Apache Airflow installed and configured
- Python 3.7 or later
- Basic understanding of DAG concepts and task scheduling
How to use airflow-dag-patterns
- 1.Review DAG design principles: idempotent, atomic, incremental, and observable
- 2.Define default_args with appropriate retries, timeouts, and error handling
- 3.Create DAG with schedule, start_date, and catchup settings
- 4.Define tasks using PythonOperator, EmptyOperator, or custom operators
- 5.Set up task dependencies using >> operator (linear, fan-out, fan-in patterns)
- 6.Add logging and monitoring at each task step
- 7.Write unit and integration tests for your DAG
- 8.Deploy to production Airflow environment
Use cases
- Creating data pipeline orchestration with Apache Airflow
- Designing DAG structures and task dependencies for ETL workflows
- Implementing custom operators and sensors for specialized tasks
- Testing Airflow DAGs locally before production deployment
- Setting up production Airflow environments with best practices
- Data engineers building data pipelines
- Platform engineers setting up Airflow infrastructure
- Backend developers orchestrating batch jobs
- DevOps engineers managing workflow automation
airflow-dag-patterns FAQ
Use TaskFlow API for cleaner code and automatic XCom handling. It's the modern approach and reduces boilerplate compared to traditional operators.
Idempotent means running a task twice produces the same result. This is critical for safe retries and ensures data consistency even if tasks are re-executed.
No. Avoid depends_on_past=True as it creates bottlenecks and can cause cascading failures. Design tasks to be independent when possible.
Set task timeouts in your DAG configuration. This ensures tasks that hang or get stuck are automatically terminated rather than consuming resources indefinitely.
retry_delay sets a fixed wait time between retries, while retry_exponential_backoff increases the delay exponentially with each retry attempt, up to max_retry_delay.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: airflow-dag-patterns description: Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
Apache Airflow DAG Patterns
Production-ready patterns for Apache Airflow including DAG design, operators, sensors, testing, and deployment strategies.
When to Use This Skill
- Creating data pipeline orchestration with Airflow
- Designing DAG structures and dependencies
- Implementing custom operators and sensors
- Testing Airflow DAGs locally
- Setting up Airflow in production
- Debugging failed DAG runs
Core Concepts
1. DAG Design Principles
| Principle | Description |
|---|---|
| Idempotent | Running twice produces same result |
| Atomic | Tasks succeed or fail completely |
| Incremental | Process only new/changed data |
| Observable | Logs, metrics, alerts at every step |
2. Task Dependencies
# Linear
task1 >> task2 >> task3
# Fan-out
task1 >> [task2, task3, task4]
# Fan-in
[task1, task2, task3] >> task4
# Complex
task1 >> task2 >> task4
task1 >> task3 >> task4
Quick Start
# dags/example_dag.py
from datetime import datetime, timedelta
from airflow import DAG
from airflow.operators.python import PythonOperator
from airflow.operators.empty import EmptyOperator
default_args = {
'owner': 'data-team',
'depends_on_past': False,
'email_on_failure': True,
'email_on_retry': False,
'retries': 3,
'retry_delay': timedelta(minutes=5),
'retry_exponential_backoff': True,
'max_retry_delay': timedelta(hours=1),
}
with DAG(
dag_id='example_etl',
default_args=default_args,
description='Example ETL pipeline',
schedule='0 6 * * *', # Daily at 6 AM
start_date=datetime(2024, 1, 1),
catchup=False,
tags=['etl', 'example'],
max_active_runs=1,
) as dag:
start = EmptyOperator(task_id='start')
def extract_data(**context):
execution_date = context['ds']
# Extract logic here
return {'records': 1000}
extract = PythonOperator(
task_id='extract',
python_callable=extract_data,
)
end = EmptyOperator(task_id='end')
start >> extract >> end
Detailed patterns and worked examples
Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
Best Practices
Do's
- Use TaskFlow API - Cleaner code, automatic XCom
- Set timeouts - Prevent zombie tasks
- Use
mode='reschedule'- For sensors, free up workers - Test DAGs - Unit tests and integration tests
- Idempotent tasks - Safe to retry
Don'ts
- Don't use
depends_on_past=True- Creates bottlenecks - Don't hardcode dates - Use
{{ ds }}macros - Don't use global state - Tasks should be stateless
- Don't skip catchup blindly - Understand implications
- Don't put heavy logic in DAG file - Import from modules
Related skills
More from wshobson/agents and the wider catalog.

angular-migration
Migrate AngularJS apps to Angular using hybrid mode, incremental rewriting, and dependency injection updates.

anti-reversing-techniques
Understand and analyze anti-reversing, obfuscation, and protection techniques in authorized security contexts.

api-design-principles
Master REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs.

architecture-decision-records
Write and maintain Architecture Decision Records (ADRs) to document significant technical decisions with context, rationale, and consequences.

architecture-patterns
Implement Clean Architecture, Hexagonal Architecture, and Domain-Driven Design for maintainable backend systems.

async-python-patterns
Master asyncio, concurrent programming, and async/await patterns for high-performance Python applications.