chdb-datastore
clickhouse/agent-skills
Drop-in pandas replacement with ClickHouse speed for filtering, grouping, joining, and cross-source data operations.
What is chdb-datastore?
chdb DataStore is a lazy, ClickHouse-backed pandas API that compiles operations to optimized SQL. Use it when you have tabular data (DataFrame, parquet, CSV, Arrow, JSON) and need faster filtering, aggregation, joins, or cross-source data integration from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, or Delta Lake.
- Replace pandas with one-line import change; all 209 DataFrame methods work unchanged
- Connect to 16+ data sources (files, databases, cloud storage) with unified API
- Filter, group, aggregate, sort, and compute columns using familiar pandas syntax
- Join data across different sources (e.g., MySQL table with S3 parquet file)
- Inspect generated SQL with .to_sql() for debugging and optimization
- Write results to files or databases with insert_into() and execute()
How to install chdb-datastore
npx skills add https://github.com/clickhouse/agent-skills --skill chdb-datastore- Python 3.9 or later
- macOS or Linux
- pip install chdb
How to use chdb-datastore
- 1.Install chdb: pip install chdb
- 2.Import DataStore: from datastore import DataStore or import chdb.datastore as pd
- 3.Connect to your data source using DataStore.from_file(), from_mysql(), from_s3(), or .uri()
- 4.Use standard pandas methods (filter, groupby, join, sort, assign) on the DataStore object
- 5.Call .to_sql() to inspect generated SQL, or .head() to preview results
- 6.Execute operations by printing, iterating, or calling .execute() to write results
Use cases
- Speed up slow pandas workflows by switching import without rewriting code
- Join customer data from MySQL with order history from parquet files in one query
- Aggregate large datasets from S3 without loading into memory
- Filter and group tabular data from multiple databases in a single operation
- Preview and debug cross-source joins by inspecting the generated SQL
- Data analysts working with pandas who need better performance
- Python developers joining data from multiple sources (databases, files, cloud)
- Engineers optimizing slow pandas-based data pipelines
- Anyone analyzing tabular data (CSV, parquet, JSON, Arrow) who wants lazy evaluation
chdb-datastore FAQ
No. If you change import pandas as pd to import chdb.datastore as pd, your existing code works unchanged. DataStore implements 209 pandas DataFrame methods.
Use chdb-datastore for pandas-style operations (filter, groupby, join). Use chdb-sql for raw SQL queries and ClickHouse server administration.
Yes. Create separate DataStore objects from each source (MySQL, S3, files, etc.) and use .join() to combine them across sources.
Call .to_sql() on your DataStore to see the generated SQL, then inspect the query for correctness.
Verify chdb is installed (pip install chdb), check that you're using DataStore not pandas, and ensure your data source connection includes the port (e.g., host='db:3306').
Full instructions (SKILL.md)
Source of truth, from clickhouse/agent-skills.
name: chdb-datastore
description: >-
Use when the user has tabular data (pandas DataFrame, parquet, csv,
Arrow, json) and wants to filter, group, aggregate, join, or speed
up slow pandas. Provides chDB DataStore — same pandas API,
ClickHouse engine underneath. Also handles reading from S3, MySQL,
PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as
DataFrames and joining across sources.
TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas",
"speed up pandas", or cross-source DataFrame joins; user imports
chdb.datastore or from datastore import DataStore.
SKIP this skill for raw SQL syntax (use chdb-sql instead),
ClickHouse server administration, or non-Python DataStore API work.
license: Apache-2.0
compatibility: Requires Python 3.9+, macOS or Linux. pip install chdb.
metadata:
author: chdb-io
version: "4.1"
homepage: https://clickhouse.com/docs/chdb
chdb DataStore — It's Just Faster Pandas
The Key Insight
# Change this:
import pandas as pd
# To this:
import chdb.datastore as pd
# Everything else stays the same.
DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).
pip install chdb
Decision Tree: Pick the Right Approach
1. "I have a file/database and want to analyze it with pandas"
→ DataStore.from_file() / from_mysql() / from_s3() etc.
→ See references/connectors.md
2. "I need to join data from different sources"
→ Create DataStores from each source, use .join()
→ See examples/examples.md #3-5
3. "My pandas code is too slow"
→ import chdb.datastore as pd — change one line, keep the rest
4. "I need raw SQL queries"
→ Use the chdb-sql skill instead
Connect to Any Data Source — One Pattern
from datastore import DataStore
# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)
ds = DataStore.from_file("sales.parquet")
# Database
ds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
# Cloud storage
ds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)
# URI shorthand — auto-detects source type
ds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")
All 16+ sources and URI schemes → connectors.md
After Connecting — Full Pandas API
result = ds[ds["age"] > 25] # filter
result = ds[["name", "city"]] # select columns
result = ds.sort_values("revenue", ascending=False) # sort
result = ds.groupby("dept")["salary"].mean() # groupby
result = ds.assign(margin=lambda x: x["profit"] / x["revenue"]) # computed column
ds["name"].str.upper() # string accessor
ds["date"].dt.year # datetime accessor
result = ds1.join(ds2, on="id") # join
result = ds.head(10) # preview
print(ds.to_sql()) # see generated SQL
209 DataFrame methods supported. Full API → api-reference.md
Cross-Source Join — The Killer Feature
from datastore import DataStore
customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")
orders = DataStore.from_file("orders.parquet")
result = (orders
.join(customers, left_on="customer_id", right_on="id")
.groupby("country")
.agg({"amount": "sum", "rating": "mean"})
.sort_values("sum", ascending=False))
print(result)
More join examples → examples.md
Writing Data
source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
target = DataStore("file", path="summary.parquet", format="Parquet")
target.insert_into("category", "total", "count").select_from(
source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")
).execute()
Troubleshooting
| Problem | Fix |
|---|---|
ImportError: No module named 'chdb' | pip install chdb |
ImportError: cannot import 'DataStore' | Use from datastore import DataStore or from chdb.datastore import DataStore |
| Database connection timeout | Include port in host: host="db:3306" not host="db" |
| Join returns empty result | Check key types match (both int or both string); use .to_sql() to inspect |
| Unexpected results | Call ds.to_sql() to see the generated SQL and debug |
| Environment check | Run python scripts/verify_install.py (from skill directory) |
References
- API Reference — Full DataStore method signatures
- Connectors — All 16+ data source connection methods
- Examples — 10+ runnable examples with expected output
- Verify Install — Environment verification script
- Official Docs
Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the
chdb-sqlskill. For contributing to chdb source code, see CLAUDE.md in the project root.
Related skills
More from clickhouse/agent-skills and the wider catalog.

chdb-sql
Run ClickHouse SQL on local files, URLs, S3, and remote databases without a server.

clickhouse-architecture-advisor
Workload-aware ClickHouse architecture decisioning with provenance-labeled recommendations.

clickhouse-best-practices
31 ClickHouse best-practice rules for schema design, query optimization, and data ingestion.

clickhouse-js-node-coding
Write idiomatic Node.js code against ClickHouse using the @clickhouse/client package.

clickhouse-js-node-rowbinary
Generate TypeScript/JavaScript code to read and write ClickHouse RowBinary streams in Node.js.

clickhouse-js-node-troubleshooting
Troubleshoot and resolve common issues with the ClickHouse Node.js client (@clickhouse/client).