jDataMunch MCP MCP Server
io.github.jgravelle/jdatamunch-mcp
Query CSV, Excel, Parquet, and JSONL files server-side without loading rows into context—99%+ token savings on large datasets.
What is the jDataMunch MCP MCP server?
jDataMunch MCP is an MCP server that enables AI agents to query and analyze tabular data (CSV, Excel, Parquet, JSONL) without pasting file contents into the context window. It profiles datasets once, then answers questions about structure, filters rows, aggregates data, and joins datasets server-side, reducing token costs by orders of magnitude on large files.
jDataMunch lets coding agents and analysts explore and query spreadsheets and data files efficiently. Instead of loading entire files into the model's context (which costs millions of tokens for large datasets), it indexes data locally and executes queries server-side, returning only results. A million-row CSV that would cost 111 million tokens to paste costs ~3,849 tokens to describe with jDataMunch.
How to install jDataMunch MCP
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
index_local— Index a local CSV, Excel, Parquet, or JSONL file and store its profiledescribe_dataset— Return column names, inferred types, cardinality, null rates, and sample values without reading rows into contextdescribe_column— Get distribution, statistics, and deep-dive analysis on a single columnget_rows— Retrieve filtered rows from a dataset without loading the entire fileaggregate— Perform server-side grouping and aggregation (count, sum, mean, etc.) without returning raw rowsrun_sql— Execute SQL queries against indexed datasetsplan_query— Preview the cost (token and latency) of a query before running itsample_rows— Get random samples from a datasetget_distribution— Analyze the distribution of values in a columnget_correlations— Calculate correlations between numeric columnssuggest_joins— Identify potential join keys across multiple datasetssuggest_keys— Recommend primary and foreign key candidatesjoin_datasets— Perform cross-dataset joins server-sideget_dataset_health— Assess overall data quality and identify issuesdata_health_radar— Visualize data quality metrics across the datasetget_data_hotspots— Find anomalies in null rates, cardinality, and outlier spreadget_schema_drift— Detect schema inconsistencies and changesfind_unused_columns— Identify columns that are rarely or never populatedcheck_column_drop_safe— Verify whether dropping a column is safe before making the changeget_schema_impact— Assess the impact of schema changes
Use cases
- Understand the structure and quality of a large CSV or Excel file in seconds without pasting it into the model
- Filter and retrieve specific rows from a million-row dataset with minimal token cost
- Aggregate and group data server-side (e.g., count by category) without loading raw rows
- Detect data-quality issues like missing values, outliers, and schema drift
- Join and correlate data across multiple files to find relationships and patterns
jDataMunch MCP MCP server FAQ
jDataMunch is an MCP server that profiles and indexes CSV, Excel, Parquet, and JSONL files locally, then answers questions about them by executing queries server-side. Instead of pasting entire files into the model's context, it returns only the results you need, saving 99%+ of tokens on large datasets.
Yes, jDataMunch is free for non-commercial and personal use. Commercial use requires a paid license (starting at $39 one-time for a single developer).
Run `claude mcp add jdatamunch -- uvx jdatamunch-mcp` (or equivalent for your client). For Excel or Parquet support, add extras: `uvx --from "jdatamunch-mcp[excel,parquet]" jdatamunch-mcp`. Full setup instructions are in QUICKSTART.md.
No. jDataMunch is local-first: your data is profiled and indexed on your machine only. No data, column names, or file paths are sent anywhere. The only optional network behavior is an anonymous savings counter (opt out with JDATAMUNCH_SHARE_SAVINGS=0).
CSV, TSV, and JSON Lines are built in. Excel (.xlsx, .xls) and Parquet require optional extras installed via `pip install "jdatamunch-mcp[excel,parquet]"`
On a 255 MB, 1-million-row CSV, describing the dataset costs ~3,849 tokens instead of 111 million tokens (a 25,000× reduction). Filtering to matching rows saves 99%+. Savings scale with file size; small spreadsheets see minimal benefit.
README (reference)
Source of truth, from the repository.
jDataMunch MCP: Tabular Data Retrieval for AI Agents
jDataMunch is an MCP server for coding agents and analysts that answers questions about CSV, Excel, Parquet, and JSONL files without pasting the rows into the context window.
Index a dataset once, then retrieve column profiles, filtered rows, server-side aggregations, and cross-dataset joins — so a million-row file costs thousands of tokens instead of millions.
Install · Quickstart · Benchmarks · Commercial licensing
Free for personal use. Commercial use requires a paid license — terms below.
Why jDataMunch?
The problem. The default way an agent explores a spreadsheet is to paste it into the prompt. A 255 MB CSV with a million rows costs roughly 111 million tokens that way, and the model still has to reason through a million rows to answer "what columns are in here?"
The mechanism. jDataMunch profiles the file once — columns, types, cardinality, null rates, distributions — and stores that locally. Queries then run against the data, not against a copy of it in the prompt: filters, aggregations, and joins execute server-side and return only results.
The outcome. Orientation questions are answered from the profile. Row-level questions return matching rows. The raw file never enters the context window.
Evidence
Measured on a real public dataset, not estimated. Full harness and per-query results in benchmarks/.
Corpus: LAPD crime records — 1,004,894 rows, 28 columns, 255 MB Baseline: 111,028,360 tokens to paste the raw file
describe_dataset: ~3,849 tokens — a 25,333× reduction Methodology & harness · Full results
| Task | Without jDataMunch | With jDataMunch | Reduction |
|---|---|---|---|
| Understand a dataset's shape | Paste 111M tokens | describe_dataset → ~3,849 tokens | ~25,000× |
| Schema + one column deep-dive | Paste 111M tokens | describe_dataset + describe_column → ~4,400 tokens | ~25,000× |
| Filter to matching rows | Load all 1M rows | get_rows with filters → matching rows only | ~99%+ |
| Count by category | Return all rows, aggregate in the model | aggregate(group_by=[...]) → 21 rows | ~99.9% |
What these numbers are and are not. The reduction is measured against pasting the complete file, which is what a naive agent does and what the token bill reflects. It is not measured against a competent human analyst who would never paste a 255 MB CSV. The multiple scales with file size: a 200-row spreadsheet has far less to save, and the honest figure there is closer to "no meaningful difference."
Typical latencies from the same run: describe_column on a single column, 22–33 ms and ~600 tokens.
Install
Requirements: Python 3.10+, any MCP-compatible client.
There is no install step. jdatamunch-mcp is a stdio MCP server with no CLI subcommands, so nothing needs to land on your PATH — point your client at uvx and it fetches and runs the server on demand.
Claude Code setup:
claude mcp add jdatamunch -- uvx jdatamunch-mcp
Nothing else. Don't have uv yet?
Reading Excel or Parquet? Those pull optional extras, which uvx takes on the --from argument:
claude mcp add jdatamunch -- uvx --from "jdatamunch-mcp[excel,parquet]" jdatamunch-mcp
<details>
<summary><b>Prefer a persistent install?</b></summary>
| Command | Use it when |
|---|---|
uv tool install jdatamunch-mcp | You want it resolved once instead of per-launch |
pipx install jdatamunch-mcp | You already standardise on pipx |
pip install jdatamunch-mcp | Inside a virtualenv you manage yourself. ⚠ Refused on PEP 668 distros (Ubuntu 24.04+, Debian 12+) — use one of the two above. |
Extras take the usual bracket form here: uv tool install "jdatamunch-mcp[excel,parquet]". Registering the server still works the same way; substitute jdatamunch-mcp for uvx jdatamunch-mcp in the claude mcp add line above.
Restart Claude Code, then type /mcp — jdatamunch should be listed. That listing is the verification step; running the server directly just waits on stdin.
Full per-client setup, including Claude Desktop, Cursor, and Windsurf: QUICKSTART.md.
Quickstart
Assumes: jDataMunch installed and registered with your client, and a CSV to hand.
Everything happens inside your agent — there is no separate indexing command. Ask it to index:
Using jdatamunch, index ./data/sales.csv
It calls index_local, which returns the dataset name, row and column counts, and detected types. Then:
Using jdatamunch, describe the sales dataset and tell me which columns have missing values.
The agent calls describe_dataset, which returns column names, inferred types, cardinality, null rates, and sample values — without reading a single row into context. _meta.tokens_saved reports what that cost against loading the file.
Next step: describe_column for a distribution on one column, or aggregate to group and count server-side.
What you can do
- Orient in a dataset you have never seen.
describe_dataset,describe_column,sample_rows,get_distribution,get_correlations. - Query without loading rows.
get_rowswith filters,aggregatewithgroup_by,run_sql, andplan_queryto preview cost before running. - Work across datasets.
suggest_joins,suggest_keys,join_datasets. - Find data-quality problems.
get_dataset_health,data_health_radar,get_data_hotspots(null rate, cardinality anomalies, outlier spread),get_schema_drift,find_unused_columns. - Preflight schema changes.
check_column_drop_safeandget_schema_impactbefore you drop or rename. - Search semantically.
search_dataandfind_similar_columnswhen you know what you mean but not what it is called. - Index from GitHub.
index_repopulls CSV, Excel, Parquet, and JSONL straight from a repository, incrementally by HEAD SHA, private repos included.
39 tools in total. Full reference: USER-MANUAL.md.
How it works
Everything runs locally. The dataset is profiled on your machine and the index is stored on your machine; no hosted service is involved in indexing or querying.
data.csv ──► profiler ──► column stats + local index
│
MCP client ◄── query ─┘ (filters, aggregates, joins
execute server-side)
Aggregations and filters run against the stored data rather than being simulated in the model, which is why the row count barely affects the token cost of an answer. Sampling-based statistics report their error bounds (roughly 2% standard error) rather than presenting an estimate as exact.
Supported formats
| Format | Extensions | Install extra |
|---|---|---|
| CSV / TSV | .csv, .tsv | built in |
| JSON Lines | .jsonl | built in |
| Excel | .xlsx, .xls | pip install "jdatamunch-mcp[excel]" |
| Parquet | .parquet | pip install "jdatamunch-mcp[parquet]" |
Security and privacy
Local-first. Your data is profiled and indexed on your machine and is not uploaded.
The base package's only default network behavior is an anonymous savings counter — a random ID plus aggregate token counts. No data, no column names, no file paths, no PII. Opt out completely:
JDATAMUNCH_SHARE_SAVINGS=0
index_repo reaches GitHub only when you invoke it, using a token you supply. Embedding providers are called only when you configure one. There is no scheduler and no background reporting.
Full detail, including what each optional extra pulls in: SECURITY.md.
Limitations
- Savings scale with file size. On a small spreadsheet the difference is negligible; the benchmark figures come from a 255 MB file.
- Sampled statistics are sampled. Distribution and correlation figures on very large files carry a stated error bound rather than being exact.
- Excel and Parquet need optional extras, which pull additional dependencies.
- A default
describe_columnwill not be labelledoffloadable. jDataMunch does not assert index freshness it cannot prove, so the cheap freshness reading answersunknownand the annotation fails closed. That is deliberate — see the annotation section. - jDataMunch does not read code or prose. Code symbols belong to jcodemunch-mcp; documentation sections to jdocmunch-mcp.
Offloadable-work annotation
JMUNCH_OFFLOADABLE=1 (suite-wide) or JDATAMUNCH_OFFLOADABLE=1 (this server only) makes describe_column carry an advisory _meta.offloadable block marking whether the answer is simple and self-contained enough to hand to a cheaper model.
It is a label and nothing else. jDataMunch never calls another model, never routes the request, and never touches your API keys. Off by default; you decide what happens next.
The verdict is tri-state and reason-coded: not_evaluated ("we did not assess it") is not not_offloadable ("this is not simple work"). It fails closed — any unknown bearing on the answer disqualifies, because a false offloadable sends real work to a model that will confabulate over the gap. verify_with names the call that would adjudicate a cheaper model's answer.
Identical field contract across all three jMunch servers, with a pinned contract digest that fails the build in any one of them that drifts.
Documentation
| Doc | What it covers |
|---|---|
| QUICKSTART.md | Zero-to-indexed in three steps |
| USER-MANUAL.md | Full guide for analysts, ops, and non-developers |
| SECURITY.md | Data handling, network behavior, vulnerability reporting |
| benchmarks/METHODOLOGY.md | How the benchmark is run and what it measures |
| CONTRIBUTING.md | Development setup and the CLA requirement |
| CHANGELOG.md | Release history |
Licensing and commercial use
Released under the jDataMunch-MCP Dual-Use License (full terms). Free for non-commercial use. Commercial use requires a paid license, one-time, sold by jMunch LLC.
jDataMunch only: Builder, $39 (1 developer) · Studio, $149 (up to 5) · Platform, $499 (org-wide internal deployment)
Full jMunch suite (code + docs + data): Trio Builder, $99 · Trio Studio, $449 · Trio Platform, $2,499
Individual developers and non-commercial projects need no license. Organizations deploying jDataMunch across internal teams do.
Support and project status
Actively maintained. Issues and bug reports: GitHub Issues. Commercial licensing questions go through jcodemunch.com.
Part of the jMunch suite alongside jcodemunch-mcp (code symbols) and jdocmunch-mcp (documentation sections). All three implement jMRI, the open retrieval interface spec — same response envelope, same token accounting.
Related MCP servers

jDocmunch MCP
Section-level documentation search for agents—retrieve exact sections from .md, .rst, .adoc, .ipynb, .html, .yaml, .json, and OpenAPI specs without loading whole files.

Transparent MCP proxy that reduces token costs by 88-99% through handle-ified queryable responses.

jCodemunch MCP
Token-efficient code exploration via tree-sitter AST parsing. 70+ languages, 86-99% token savings.
Ask a real human (not an LLM) for a judgment, $0.10 USDC via x402. The result returns to the agent.
View repository →
Copilot Studio MCP
Build, test, ship and maintain Microsoft Copilot Studio agents from the editor
Record a flow through your app, check it after every change, and attack a copy of your database.
