PluginBench
Skill
Review
Audit score 70

data-quality-frameworks

wshobson/agents

Implement data quality validation with Great Expectations, dbt tests, and data contracts for reliable pipelines.

What is data-quality-frameworks?

This skill provides production patterns for implementing data quality checks using Great Expectations, dbt tests, and data contracts. Use it when building data quality pipelines, setting up validation rules, establishing data contracts between teams, or monitoring data quality metrics in CI/CD workflows.

  • Validate data completeness, uniqueness, validity, accuracy, consistency, and timeliness
  • Set up Great Expectations datasources and expectation suites with built-in validators
  • Build comprehensive dbt test suites following a testing pyramid approach
  • Establish and version data contracts between teams
  • Generate data quality reports with pass/fail status and detailed failure analysis
  • Integrate data validation into CI/CD pipelines with automated alerting

How to install data-quality-frameworks

npx skills add https://github.com/wshobson/agents --skill data-quality-frameworks
Prerequisites
  • Python environment with pip
  • Great Expectations library (pip install great_expectations)
  • dbt project (for dbt test patterns)
  • Access to data sources to validate
Claude Code
Cursor
Windsurf
Cline

How to use data-quality-frameworks

  1. 1.Install Great Expectations: pip install great_expectations
  2. 2.Initialize a Great Expectations project: great_expectations init
  3. 3.Create a datasource pointing to your data: great_expectations datasource new
  4. 4.Define an expectation suite with validation rules for your tables
  5. 5.Add specific expectations (e.g., ExpectColumnValuesToNotBeNull, ExpectColumnValuesToBeUnique)
  6. 6.Create a checkpoint that runs your expectation suites
  7. 7.Integrate the checkpoint into your pipeline to validate data before downstream processing
  8. 8.Review generated reports and configure alerts for failed validations

Use cases

Good for
  • Validating source data before transformations in ETL pipelines
  • Monitoring critical columns for null values, duplicates, and out-of-range data
  • Cross-table validation to ensure consistency across related datasets
  • Automating data contract enforcement when data moves between teams
  • Generating quality reports that fail pipelines when validation thresholds are breached
Who it's for
  • Data engineers building production pipelines
  • Analytics engineers implementing dbt test suites
  • Data quality teams establishing validation standards
  • Platform teams enforcing data contracts
  • DevOps engineers integrating data validation into CI/CD

data-quality-frameworks FAQ

What data quality dimensions should I test?

Focus on the six key dimensions: Completeness (no nulls), Uniqueness (no duplicates), Validity (values in expected range), Accuracy (matches reality), Consistency (no contradictions), and Timeliness (data is recent). Prioritize critical columns rather than testing everything.

How do I structure data quality tests?

Use a testing pyramid: schema tests at the base (structure validation), unit tests in the middle (single column checks), and integration tests at the top (cross-table relationships). Add tests incrementally as you discover data issues.

Can I use this with dbt?

Yes, this skill covers both Great Expectations and dbt test patterns. You can implement validation rules in dbt YAML or use Great Expectations for more complex validations, then integrate both into your pipeline.

How do I fail a pipeline on data quality issues?

Run your Great Expectations checkpoint and check if all results passed. Raise an exception or exit with error code if any validation fails, which will stop downstream processing in your CI/CD pipeline.

Should I hardcode validation thresholds?

No, avoid hardcoding thresholds. Use dynamic baselines that adapt to your data patterns, and document clear expectations for each validation rule so teams understand what constitutes a failure.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: data-quality-frameworks description: Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.

Data Quality Frameworks

Production patterns for implementing data quality with Great Expectations, dbt tests, and data contracts to ensure reliable data pipelines.

When to Use This Skill

  • Implementing data quality checks in pipelines
  • Setting up Great Expectations validation
  • Building comprehensive dbt test suites
  • Establishing data contracts between teams
  • Monitoring data quality metrics
  • Automating data validation in CI/CD

Core Concepts

1. Data Quality Dimensions

DimensionDescriptionExample Check
CompletenessNo missing valuesexpect_column_values_to_not_be_null
UniquenessNo duplicatesexpect_column_values_to_be_unique
ValidityValues in expected rangeexpect_column_values_to_be_in_set
AccuracyData matches realityCross-reference validation
ConsistencyNo contradictionsexpect_column_pair_values_A_to_be_greater_than_B
TimelinessData is recentexpect_column_max_to_be_between

2. Testing Pyramid for Data

          /\
         /  \     Integration Tests (cross-table)
        /────\
       /      \   Unit Tests (single column)
      /────────\
     /          \ Schema Tests (structure)
    /────────────\

Quick Start

Great Expectations Setup

# Install
pip install great_expectations

# Initialize project
great_expectations init

# Create datasource
great_expectations datasource new
# great_expectations/checkpoints/daily_validation.yml
import great_expectations as gx

# Create context
context = gx.get_context()

# Create expectation suite
suite = context.add_expectation_suite("orders_suite")

# Add expectations
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)

# Validate
results = context.run_checkpoint(checkpoint_name="daily_orders")

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Summary: {total_passed}/{total_tables} tables passed")

    report.append("")

    for table, result in results.items():
        status = "✅" if result.passed else "❌"
        report.append(f"### {status} {table}")
        report.append(f"- Expectations: {result.total_expectations}")
        report.append(f"- Failed: {result.failed_expectations}")

        if not result.passed:
            report.append("- Failed checks:")
            for detail in result.details:
                if not detail["success"]:
                    report.append(f"  - {detail['expectation']}: {detail['observed_value']}")
        report.append("")

    return "\n".join(report)

Usage

context = gx.get_context() pipeline = DataQualityPipeline(context)

tables_to_validate = { "orders": "orders_suite", "customers": "customers_suite", "products": "products_suite", }

results = pipeline.run_all(tables_to_validate) report = pipeline.generate_report(results)

Fail pipeline if any table failed

if not all(r.passed for r in results.values()): print(report) raise ValueError("Data quality checks failed!")


## Best Practices

### Do's

- **Test early** - Validate source data before transformations
- **Test incrementally** - Add tests as you find issues
- **Document expectations** - Clear descriptions for each test
- **Alert on failures** - Integrate with monitoring
- **Version contracts** - Track schema changes

### Don'ts

- **Don't test everything** - Focus on critical columns
- **Don't ignore warnings** - They often precede failures
- **Don't skip freshness** - Stale data is bad data
- **Don't hardcode thresholds** - Use dynamic baselines
- **Don't test in isolation** - Test relationships too