crm-data-quality
hubspot/agent-cli-skills
Find incomplete records, normalize values in bulk, and dedupe contacts with HubSpot merge operations.
What is crm-data-quality?
This skill identifies data quality issues in HubSpot CRM—missing fields, inconsistent values, and duplicates—then fixes them at scale using bulk operations. Use it to audit and clean contact, company, or deal records before analysis or sync.
- Search for incomplete records with missing or empty fields using NOT_HAS_PROPERTY filters
- Normalize field values in bulk by searching, reshaping with jq, and piping into update
- Deduplicate records by merging secondary contacts into primary ones (irreversible)
- Audit custom properties and group them by type or category
- Perform dry-run previews and digest/confirm gating for operations affecting >100 records
How to install crm-data-quality
npx skills add https://github.com/hubspot/agent-cli-skills --skill crm-data-quality- HubSpot CLI installed and authenticated
- Familiarity with bulk-operations skill (JSONL piping, dry-run/digest/confirm workflow)
- jq installed for reshaping JSON records
- Understanding of HubSpot property types and filter syntax
How to use crm-data-quality
- 1.List available properties for your object type using `hubspot properties list --type <type>`
- 2.Search for incomplete records using `hubspot objects search` with `!fieldname` filters
- 3.Preview results with `--dry-run` before executing any update or merge
- 4.For normalization: pipe search results through jq to reshape values, then into `hubspot objects update`
- 5.For deduplication: collect full record set with pagination loop, group by duplicate key (e.g., email), pipe merge pairs into `hubspot objects merge --dry-run`
- 6.For >100 operations: use `--digest` to preview impact, then `--confirm` to execute
- 7.Check audit trail with `hubspot history --since 1h` after merges to verify correctness
Use cases
- Find all contacts missing email or phone and bulk-update them with normalized values
- Identify duplicate contacts by email address, merge them, and audit the merge trail
- Collapse variant company name spellings ("Acme", "ACME Corp", "Acme Corporation") into one canonical value
- Audit all enumeration properties in your contacts object to find unused or misconfigured fields
- Create a data-quality flag property and mark records with missing critical fields
- CRM administrators managing data hygiene
- Sales operations teams deduplicating contact lists
- Data analysts preparing CRM data for reporting or integration
- Teams integrating HubSpot with external systems that require clean data
crm-data-quality FAQ
The `objects search` API caps at 100 rows per call. You must use the pagination loop from bulk-operations/SKILL.md to collect the full dataset into a file before grouping and merging.
No, merge is irreversible. However, you can restore the secondary record from HubSpot's recycle bin if you merged in the wrong direction. Always use `--dry-run` first.
Use `hubspot properties list --type <type> | jq -c 'select(.type=="enumeration")'` to list enum properties. Option values are not exposed by the CLI; read them from a real record or the HubSpot UI.
`!email` finds records with missing or empty email (NOT_HAS_PROPERTY). Bare `email` finds records that have a non-empty email value.
Only if you want to add a new field (e.g., a data-quality flag). For existing properties, search and reshape directly with jq, then update.
Full instructions (SKILL.md)
Source of truth, from hubspot/agent-cli-skills.
name: crm-data-quality
description: Find incomplete records, normalize field values in bulk, dedupe with hubspot objects merge, and audit custom properties. Builds on bulk-operations for JSONL piping and dry-run/digest/confirm.
triggers:
- "clean up contacts"
- "data quality"
- "deduplicate"
- "missing fields"
- "normalize data"
- "find incomplete records"
- "merge duplicates"
- "audit properties"
Read bulk-operations/SKILL.md first — JSONL piping, batch read, pagination, and dry-run/digest/confirm gating apply to every command below.
Property discovery
Don't guess property names. List them:
hubspot properties list --type contacts --format table
hubspot properties list --type contacts | jq -c 'select(.type=="enumeration") | {name, label}'
Same for --type companies, deals, or any custom type (hubspot objects types).
1. Find incomplete records
!name = NOT_HAS_PROPERTY (missing or empty). Bare name = HAS_PROPERTY. Within one --filter, chain with AND; multiple --filter flags are OR'd.
hubspot objects search --type contacts --filter "!email" --properties firstname,lastname,company
hubspot objects search --type contacts --filter "!phone AND !mobilephone" --properties email
hubspot objects search --type contacts --filter "!hubspot_owner_id" --properties email,lifecyclestage
For >100 results, use the pagination loop from bulk-operations.
2. Normalize field values
Search → reshape with jq → pipe into update. Always --dry-run first; bulk-operations covers digest/confirm escalation for >100 rows. Reshape patterns: bulk-operations/resources/json-patterns.md.
# Collapse spellings into one canonical value
hubspot objects search --type contacts --filter "company~acme" \
| jq -c '{id, properties:{company:"Acme Corporation"}}' \
| hubspot objects update --type contacts --dry-run
# Lowercase emails (read, reshape, write)
hubspot objects search --type contacts --filter "email" --properties email \
| jq -c '{id, properties:{email: (.properties.email | ascii_downcase)}}' \
| hubspot objects update --type contacts --dry-run
3. Dedupe with hubspot objects merge
Secondary is folded into primary and deleted. Irreversible. Dry-run/digest/confirm gating applies.
# Single pair
hubspot objects merge --type contacts --primary 149 --secondary 425 --dry-run
hubspot objects merge --type contacts --primary 149 --secondary 425 # execute (≤100 pairs)
Bulk: pipe JSONL {"primary":"...","secondary":"..."} on stdin (omit --primary/--secondary).
Pagination required. objects search caps at 100 rows per call and jq -s slurps a single stream into memory — running the snippet below against a raw search will silently miss every duplicate that crosses a page boundary. Collect the full set first with the pagination loop from bulk-operations/SKILL.md (write to /tmp/contacts.jsonl), then dedupe from the file:
# /tmp/contacts.jsonl produced by the pagination loop (bulk-operations/SKILL.md)
jq -s -c '
group_by(.properties.email)[]
| select(length > 1)
| sort_by(.id | tonumber)
| .[0].id as $p | .[1:][] | {primary: $p, secondary: .id}
' /tmp/contacts.jsonl \
| hubspot objects merge --type contacts --dry-run | tee /tmp/merge-preview.jsonl
For >100 pairs, lift digest and impact.records_affected from the BulkData line and re-pipe the same producer with --digest/--confirm (see bulk-operations).
4. Audit properties
hubspot properties list (and get, batch-read) emits {name, label, type, fieldType, groupName} per row. Enum option values are not currently exposed by the CLI — read them off a real record (hubspot objects search ... --properties <enum>) or the HubSpot UI.
# Count properties per group (HubSpot groups standard fields; custom groups stand out)
hubspot properties list --type contacts | jq -rs 'group_by(.groupName) | map({group: .[0].groupName, count: length}) | .[]'
# All enumeration properties
hubspot properties list --type contacts | jq -c 'select(.type=="enumeration") | {name, label, fieldType}'
# Create a DQ flag property, then set it via the normalize pattern in section 2
hubspot properties create --type contacts --name dq_missing_phone --label "DQ: Missing Phone" --prop-type string --field-type text
Recovery
Merge is irreversible. After any merge, hubspot history --since 1h captures the audit trail. If wrong direction, restore the secondary from the UI's recycle bin.
Related skills
More from hubspot/agent-cli-skills and the wider catalog.

crm-lookup
Find CRM records by ID, email, domain, or name, and traverse associations for complete account context.

custom-object-management
Discover, create, update, and delete custom CRM object schemas in HubSpot.

customer-retention
Identify inactive/at-risk customers and create follow-up tasks at scale via CRM filters.

data-enrichment
Match external records to CRM contacts/companies by email or domain, then upsert enriched data in one pass.

deal-management
Run the full deal lifecycle from CLI — discover pipelines, qualify MQLs, advance/reassign in bulk, hunt stalled deals, and close.

quote-to-cash
Build product catalogs, assemble quotes with line items, and track invoices and subscriptions through revenue.