PluginBench
Skill
Official
Fail
Audit score 45

chdb-datastore

clickhouse/agent-skills

Drop-in pandas replacement with ClickHouse speed for filtering, grouping, joining, and cross-source data operations.

What is chdb-datastore?

chdb DataStore is a lazy, ClickHouse-backed pandas API that compiles operations to optimized SQL. Use it when you have tabular data (DataFrame, parquet, CSV, Arrow, JSON) and need faster filtering, aggregation, joins, or cross-source data integration from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, or Delta Lake.

  • Replace pandas with one-line import change; all 209 DataFrame methods work unchanged
  • Connect to 16+ data sources (files, databases, cloud storage) with unified API
  • Filter, group, aggregate, sort, and compute columns using familiar pandas syntax
  • Join data across different sources (e.g., MySQL table with S3 parquet file)
  • Inspect generated SQL with .to_sql() for debugging and optimization
  • Write results to files or databases with insert_into() and execute()

How to install chdb-datastore

npx skills add https://github.com/clickhouse/agent-skills --skill chdb-datastore
Prerequisites
  • Python 3.9 or later
  • macOS or Linux
  • pip install chdb
Claude Code
Cursor
Windsurf
Cline

How to use chdb-datastore

  1. 1.Install chdb: pip install chdb
  2. 2.Import DataStore: from datastore import DataStore or import chdb.datastore as pd
  3. 3.Connect to your data source using DataStore.from_file(), from_mysql(), from_s3(), or .uri()
  4. 4.Use standard pandas methods (filter, groupby, join, sort, assign) on the DataStore object
  5. 5.Call .to_sql() to inspect generated SQL, or .head() to preview results
  6. 6.Execute operations by printing, iterating, or calling .execute() to write results

Use cases

Good for
  • Speed up slow pandas workflows by switching import without rewriting code
  • Join customer data from MySQL with order history from parquet files in one query
  • Aggregate large datasets from S3 without loading into memory
  • Filter and group tabular data from multiple databases in a single operation
  • Preview and debug cross-source joins by inspecting the generated SQL
Who it's for
  • Data analysts working with pandas who need better performance
  • Python developers joining data from multiple sources (databases, files, cloud)
  • Engineers optimizing slow pandas-based data pipelines
  • Anyone analyzing tabular data (CSV, parquet, JSON, Arrow) who wants lazy evaluation

chdb-datastore FAQ

Do I need to rewrite my pandas code?

No. If you change import pandas as pd to import chdb.datastore as pd, your existing code works unchanged. DataStore implements 209 pandas DataFrame methods.

What's the difference between chdb-datastore and chdb-sql?

Use chdb-datastore for pandas-style operations (filter, groupby, join). Use chdb-sql for raw SQL queries and ClickHouse server administration.

Can I join data from different sources?

Yes. Create separate DataStore objects from each source (MySQL, S3, files, etc.) and use .join() to combine them across sources.

How do I debug unexpected results?

Call .to_sql() on your DataStore to see the generated SQL, then inspect the query for correctness.

What if my pandas code is still slow?

Verify chdb is installed (pip install chdb), check that you're using DataStore not pandas, and ensure your data source connection includes the port (e.g., host='db:3306').

Full instructions (SKILL.md)

Source of truth, from clickhouse/agent-skills.


name: chdb-datastore description: >- Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore — same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports chdb.datastore or from datastore import DataStore. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work. license: Apache-2.0 compatibility: Requires Python 3.9+, macOS or Linux. pip install chdb. metadata: author: chdb-io version: "4.1" homepage: https://clickhouse.com/docs/chdb

chdb DataStore — It's Just Faster Pandas

The Key Insight

# Change this:
import pandas as pd
# To this:
import chdb.datastore as pd
# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"
   → DataStore.from_file() / from_mysql() / from_s3() etc.
   → See references/connectors.md

2. "I need to join data from different sources"
   → Create DataStores from each source, use .join()
   → See examples/examples.md #3-5

3. "My pandas code is too slow"
   → import chdb.datastore as pd — change one line, keep the rest

4. "I need raw SQL queries"
   → Use the chdb-sql skill instead

Connect to Any Data Source — One Pattern

from datastore import DataStore

# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)
ds = DataStore.from_file("sales.parquet")

# Database
ds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")

# Cloud storage
ds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)

# URI shorthand — auto-detects source type
ds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes → connectors.md

After Connecting — Full Pandas API

result = ds[ds["age"] > 25]                                          # filter
result = ds[["name", "city"]]                                        # select columns
result = ds.sort_values("revenue", ascending=False)                  # sort
result = ds.groupby("dept")["salary"].mean()                         # groupby
result = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed column
ds["name"].str.upper()                                               # string accessor
ds["date"].dt.year                                                   # datetime accessor
result = ds1.join(ds2, on="id")                                      # join
result = ds.head(10)                                                 # preview
print(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API → api-reference.md

Cross-Source Join — The Killer Feature

from datastore import DataStore

customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")
orders = DataStore.from_file("orders.parquet")

result = (orders
    .join(customers, left_on="customer_id", right_on="id")
    .groupby("country")
    .agg({"amount": "sum", "rating": "mean"})
    .sort_values("sum", ascending=False))
print(result)

More join examples → examples.md

Writing Data

source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
target = DataStore("file", path="summary.parquet", format="Parquet")

target.insert_into("category", "total", "count").select_from(
    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")
).execute()

Troubleshooting

ProblemFix
ImportError: No module named 'chdb'pip install chdb
ImportError: cannot import 'DataStore'Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeoutInclude port in host: host="db:3306" not host="db"
Join returns empty resultCheck key types match (both int or both string); use .to_sql() to inspect
Unexpected resultsCall ds.to_sql() to see the generated SQL and debug
Environment checkRun python scripts/verify_install.py (from skill directory)

References

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.