PluginBench
Skill
Pass
Audit score 90

data-enrichment

hubspot/agent-cli-skills

Match external records to CRM contacts/companies by email or domain, then upsert enriched data in one pass.

What is data-enrichment?

This skill reshapes CSV/JSONL data and uses `hubspot objects upsert` to create or update CRM contacts (matched by email) or companies (matched by domain) with external enrichment in a single operation. Use it when you need to sync spreadsheet data into HubSpot without race conditions or manual search-then-create loops.

  • Upsert contacts by email or companies by domain as the natural key
  • Reshape CSV/JSONL input with jq and pipe directly to the CRM
  • Preview changes with --dry-run before executing
  • Split results by success/failure and retry only failed records
  • Handle rate-limit 429s and validation 400s with targeted fixes
  • Match and read without write-back using OR-search filters (up to 5 groups per call)

How to install data-enrichment

npx skills add https://github.com/hubspot/agent-cli-skills --skill data-enrichment
Prerequisites
  • HubSpot CLI installed and authenticated
  • Read `bulk-operations/SKILL.md` first for JSONL piping, dry-run, and rate-limit patterns
  • CSV-to-JSONL converter (e.g., csvkit, jq) available locally
  • Knowledge of target property names (verify with `hubspot properties list --type contacts` or `--type companies`)
Claude Code
Cursor
Windsurf
Cline

How to use data-enrichment

  1. 1.Convert your CSV to JSONL using csvjson or similar tool
  2. 2.Use jq to reshape the JSONL, mapping external fields to CRM properties and lowercasing the natural key (email or domain)
  3. 3.Run the pipeline with `--dry-run` and inspect the output to confirm field mappings
  4. 4.Remove `--dry-run` and pipe the output to a file for record-keeping
  5. 5.Split results with jq to isolate successes and failures
  6. 6.Inspect error statuses and fix the input data or property names as needed
  7. 7.Retry failed records in smaller chunks if you hit rate limits (429)

Use cases

Good for
  • Sync an external employee directory CSV into CRM contacts with job titles and company names
  • Enrich company records from a domain-keyed dataset without overwriting existing fields
  • Bulk-match email addresses against CRM to find existing contacts before importing
  • Update contact properties from a spreadsheet, creating new records for unmatched emails
  • Audit which records succeeded vs. failed in a large data import and retry failures
Who it's for
  • Data analysts syncing external sources into HubSpot
  • CRM administrators performing bulk contact or company enrichment
  • Sales ops teams importing prospect lists with deduplication
  • Anyone automating spreadsheet-to-CRM workflows

data-enrichment FAQ

What's the difference between upsert and search-then-create?

Upsert matches by a natural key (email or domain) and creates or updates in one CLI call per record, with no race window. Search-then-create requires branching logic and risks duplicate creation if the record is added between search and create.

Why must I lowercase the natural key?

HubSpot CRM matching is case-sensitive and exact. Lowercasing ensures consistent matching across different input sources.

What if I hit a 429 rate-limit error?

Split your input into smaller chunks (e.g., 100 records at a time) and rerun each chunk. The skill will process them sequentially without losing data.

Can I match on a property that isn't email or domain?

Upsert requires the natural key to be a CRM property. For other matching scenarios, use OR-search filters (up to 5 groups per call) to read matches, then update separately.

Is upsert destructive if a field is already populated?

Upsert itself is non-destructive, but your reshape can overwrite existing data. Always run --dry-run first and spot-check the output before executing.

Full instructions (SKILL.md)

Source of truth, from hubspot/agent-cli-skills.


name: data-enrichment description: Match external CSV/JSONL records to CRM contacts (by email) or companies (by domain) and write enriched data back in one pass using hubspot objects upsert. triggers:

  • "spreadsheet to CRM"
  • "match contacts by email"
  • "match companies by domain"
  • "enrich CRM from CSV"
  • "CRM write-back"
  • "create or update by email"

Prereq: read bulk-operations/SKILL.md first — JSONL piping, dry-run/digest, history, and rate-limit hygiene live there. This skill is the upsert-by-natural-key workflow on top.

The core move: upsert, not search-then-create

hubspot objects upsert --type X --id-property <natural-key> reads JSONL on stdin and creates-or-updates each row in one CLI call per record, keyed by a property (email for contacts, domain for companies). No race window, no branching. Do not loop search → empty? → create.

Per line in: {"id":"jane@example.com","properties":{"firstname":"Jane","jobtitle":"VP"}} Per line out: {"id":"123","ok":true,"data":{...,"new":true|false}} or {"ok":false,"error":{...}}. Order matches input.

CSV/JSONL → upsert stream

Reshape with jq, preview with --dry-run, then execute. Always lowercase the natural key — CRM match is exact. Confirm available property names with hubspot properties list --type contacts; never hard-code a list. See bulk-operations/resources/json-patterns.md for reshape idioms.

# CSV → JSONL (any tool); example using csvkit
csvjson external.csv | jq -c '.[]' > external.jsonl

# Preview
cat external.jsonl \
| jq -c '{id:(.email|ascii_downcase), properties:{firstname:.first, lastname:.last, jobtitle:.title, company:.company}}' \
| hubspot objects upsert --type contacts --id-property email --dry-run | head

# Execute (same pipeline, drop --dry-run, capture results)
cat external.jsonl \
| jq -c '{id:(.email|ascii_downcase), properties:{firstname:.first, lastname:.last, jobtitle:.title, company:.company}}' \
| hubspot objects upsert --type contacts --id-property email \
| tee /tmp/upsert.results.jsonl

Companies: swap --type companies --id-property domain and reshape with .domain|ascii_downcase as id.

Handle per-record OK / error output

Split with jq, inspect failure modes, retry just the failures after fixing the inputs:

jq -c 'select(.ok==true)'  /tmp/upsert.results.jsonl > /tmp/upsert.ok.jsonl
jq -c 'select(.ok==false)' /tmp/upsert.results.jsonl > /tmp/upsert.failed.jsonl
jq -r '.error.status' /tmp/upsert.failed.jsonl | sort | uniq -c   # status → count
jq -r '.data.new'    /tmp/upsert.ok.jsonl     | sort | uniq -c   # created vs updated

429s: split the input and rerun smaller chunks (see bulk-operations rate-limit notes). 400s usually mean a bad property name or invalid enum value — fix the reshape, rerun the failed inputs.

Destructive-op safety

upsert itself is non-destructive, but write-back can clobber populated fields. Always --dry-run first and spot-check. For bulk delete or overwrite of existing data, follow the dry-run → digest → confirm flow in bulk-operations/SKILL.md. Recovery: hubspot history --since 1h.

Match without upsert: OR-search → update

When you only want to read matches (no write-back), or the natural key isn't a CRM property, use repeated --filter flags — each flag is one OR group.

Verified cap: 5 OR groups per call. 6+ returns 400 too many filterGroups (count: N, max allowed: 5). Chunk 5 at a time:

# emails.txt: one lowercased email per line
xargs -n5 < emails.txt | while read -r e1 e2 e3 e4 e5; do
  args=()
  for e in "$e1" "$e2" "$e3" "$e4" "$e5"; do [ -n "$e" ] && args+=(--filter "email=$e"); done
  hubspot objects search --type contacts "${args[@]}" --properties email,firstname,company
done > /tmp/matches.jsonl

jq -c '{id, properties:{lifecyclestage:"marketingqualifiedlead"}}' /tmp/matches.jsonl \
| hubspot objects update --type contacts --dry-run

For larger keyed enrichments, prefer upsert — one pipeline, no chunking math.