parallel-data-enrichment
parallel-web/parallel-agent-skills
Bulk enrich company, people, and product data with web-sourced fields like CEO names, funding, and contact info.
What is parallel-data-enrichment?
Adds missing data fields to CSV files or inline datasets by querying the web. Use this when you need to augment a list of companies, people, or products with structured information like executive names, funding rounds, or contact details. Supports multi-turn workflows by carrying context from prior research tasks.
- Enriches CSV files or inline JSON data with web-sourced fields
- Suggests relevant enrichment columns based on your intent
- Runs enrichment asynchronously and returns results as JSON
- Supports context chaining from previous research interactions
- Handles bulk datasets with progress tracking and error reporting
How to install parallel-data-enrichment
npx skills add https://github.com/parallel-web/parallel-agent-skills --skill parallel-data-enrichment- parallel-cli installed and authenticated
- Internet access for web-sourced enrichment
- Input data as CSV file or inline JSON array
How to use parallel-data-enrichment
- 1.Prepare your input data (CSV file or JSON array of objects)
- 2.Run `parallel-cli enrich suggest` if unsure which fields to add, or specify fields directly with `--intent`
- 3.Execute `parallel-cli enrich run` with `--no-wait` to start the async job and save the returned `taskgroup_id`
- 4.Poll for results using `parallel-cli enrich poll $TASKGROUP_ID` until completion
- 5.If CSV output is needed, convert the returned JSON locally to CSV preserving all original columns and error information
Use cases
- Add CEO names and founding years to a list of 500 companies
- Enrich a prospect list with recent funding information and contact details
- Augment product data with manufacturer, pricing, and availability fields
- Follow up on a prior research task by reusing its context for related enrichment
- Convert enriched JSON results to CSV for spreadsheet workflows
- Sales and business development teams building prospect lists
- Researchers compiling datasets from multiple sources
- Data analysts preparing inputs for downstream analysis
- Anyone needing to bulk-add missing fields to structured data
parallel-data-enrichment FAQ
Duration depends on the number of rows and fields requested. The skill runs asynchronously server-side, so you can check progress via the monitoring URL while other work continues.
The output JSON includes both successful rows (with `output`) and failed rows (with `error`). You can inspect failures and decide whether to rerun, manually fix, or exclude them.
Yes. If you have the `interaction_id` from a prior task, pass it via `--previous-interaction-id` to reuse that context. This is unavailable for Zero Data Retention accounts.
You can use `--intent` with a natural description (e.g., 'CEO name and recent funding'), and the skill will suggest appropriate columns. Alternatively, pass explicit column names via `--enriched-columns`.
Results are returned as JSON (array of rows with `input` and `output`/`error` fields). You can convert to CSV locally if needed, preserving original columns and error information.
Full instructions (SKILL.md)
Source of truth, from parallel-web/parallel-agent-skills.
name: parallel-data-enrichment description: "Bulk data enrichment. Adds web-sourced fields (CEO names, funding, contact info) to lists of companies, people, or products. Use for enriching CSV files or inline data. Supports multi-turn: pass --previous-interaction-id from a prior research task to carry context forward." user-invocable: true argument-hint: <file or entities> with <fields to add> compatibility: Requires parallel-cli and internet access. allowed-tools: Bash(parallel-cli:*) metadata: author: parallel
Data Enrichment
Enrich: $ARGUMENTS
Before starting
Inform the user that enrichment may take several minutes depending on the number of rows and fields requested.
Optional: Suggest output columns
If the user gave a vague intent ("enrich these companies with useful info") and you're not sure what columns to add, ask the API for a suggestion before kicking off the run:
parallel-cli enrich suggest "Find CEO and recent funding info" --json
The response is an envelope: {title, processor, enriched_columns, warnings}. Extract just the enriched_columns array (not the whole envelope) and pass it as the value of --enriched-columns on enrich run, in place of --intent. These flags are alternative ways to specify what to enrich. If suggest returned a processor, pass it explicitly via --processor on the run call. Skip this section if the user already specified the fields they want.
enrich suggestrequiresparallel-cli≥ 0.3.0. If only that command is missing, skip the optional suggestion step and use--intentin step 1. Suggest an installation-specific upgrade from Setup. Do not classify authentication, API or invalid-input failures as an older CLI. An intent-based run itself requests a suggestion; explicit columns default tocore-fast, while intent can select another processor unless--processoroverrides it.
Step 1: Start the enrichment
Use ONE of these command patterns (substitute user's actual data):
For inline data:
parallel-cli enrich run --data '[{"company": "Google"}, {"company": "Microsoft"}]' --intent "CEO name and founding year" --target "output.csv" --no-wait --json
For CSV file:
parallel-cli enrich run --source-type csv --source "input.csv" --target "output.csv" --source-columns '[{"name": "company", "description": "Company name"}]' --intent "CEO name and founding year" --no-wait --json
If this is a follow-up to a previous research task and you have its interaction_id, add context chaining:
parallel-cli enrich run --data '...' --intent "..." --target "output.csv" --no-wait --json --previous-interaction-id "$INTERACTION_ID"
This reuses the prior Task's context. Context chaining is unavailable for Zero Data Retention (ZDR) accounts, so omit the flag there and include the needed context explicitly. Enrichment does not return a new interaction_id; retain the prior Task ID for later follow-ups. A taskgroup_id or Search/Extract session_id is not a Task interaction ID.
IMPORTANT: Always include --no-wait so the command returns immediately instead of blocking.
Save the --json output's taskgroup_id, url and num_runs immediately. There is no interaction_id field. If creation is interrupted or its response is lost, inspect whether the group was created before submitting another run. Immediately tell the user:
- Enrichment has been kicked off
- The monitoring URL where they can track progress
The group runs server-side; polling can resume later using its saved ID.
Step 2: Poll for results
Pick a persistent, run-specific output path (e.g., enrichment-acme-tgrp-<id>.json). Polling overwrites its output file, so inspect any existing file and use a new path unless replacement is intended. The output is JSON regardless of extension: an array of rows with input and either output or error. Async polling does not include basis or per-row interaction IDs; do not invent citations or context IDs.
parallel-cli enrich poll "$TASKGROUP_ID" --timeout 60 --output "enrichment-<descriptive-name>-<group-id>.json"
Important:
- Keep polls bounded;
--timeout 60allows progress updates between waits. - The
--targetfrom step 1 is unused in--no-waitmode. Only--outputhere determines where results are saved, and the file is always JSON. - A completed group can include failed rows. Count rows containing
outputseparately from rows containingerrorand compare their total withnum_runs; an empty or incomplete file is not successful enrichment of the entire input.
If polling times out or is interrupted
Timeout exit 5 or interruption ends the local wait. Check group state before saying it is still running:
parallel-cli enrich status "$TASKGROUP_ID" --json
Inspect is_active, status_counts and num_runs. Resume the same poll for an active group, or retrieve results for an inactive group and report failures or unresolved rows. Do not recreate the group on a timeout or automatically rerun failed rows. A local file-write failure can be retried with the same group ID and a writable output path.
If the user requested CSV
Convert the saved JSON locally into a separate CSV. Preserve every original input column and row, including duplicate and failed rows; keep enrichment fields separate from conflicting input names and include an error column for failures. Do not assume streamed rows match original input order or guess a join when row identity is ambiguous. Validate the row count and leave the input CSV untouched. This is local conversion, not a CSV produced by async polling; report both JSON and CSV paths.
Response format
After step 1: Share the monitoring URL (for tracking progress).
After step 2:
- Report successful, failed and total row counts, with any missing results called out.
- Preview a few successful rows and a representative failure if present, without claiming all rows succeeded.
- Tell the user the full path to the output file
After completion, link the saved output rather than repeating the monitoring URL.
Setup
If parallel-cli is not found, install and authenticate:
/parallel:parallel-cli-setup
If a documented option or command is missing, identify the install method and upgrade through that method: standalone parallel-cli update; pipx pipx upgrade parallel-web-tools; uv uv tool upgrade parallel-web-tools; Homebrew brew upgrade parallel-web/tap/parallel-cli; npm npm update -g parallel-web-cli. Recheck version and help in the agent's terminal before retrying.
For authentication or API errors, inspect the returned message. A 403 can indicate permissions, account policy or billing; it does not prove low balance. Check parallel-cli auth --json and its authenticated boolean when relevant without exposing credentials. Only a billing-specific error warrants a balance check, and adding funds needs explicit confirmation. Reuse saved group IDs; do not automatically retry an ambiguous creation.
Related skills
More from parallel-web/parallel-agent-skills and the wider catalog.

parallel-deep-research
Exhaustive multi-source research with citations when you need comprehensive, in-depth investigation.

parallel-findall
Discover structured lists of entities (companies, people, products) matching natural-language descriptions.

parallel-memory
Recall past Parallel Task, Monitor, and FindAll runs to inform current work; manage memory with retrieve, evict, and clear operations.

parallel-monitor
Continuously monitor web pages and track changes on a recurring schedule with server-side event detection.

parallel-web-extract
CLI-backed URL extraction with JSON output when web_fetch MCP is unavailable.

parallel-web-search
CLI-backed web search with JSON output and keyword filtering.