PluginBench
Skill
Review
Audit score 70

crm-data-quality

hubspot/agent-cli-skills

Find incomplete records, normalize values in bulk, and dedupe contacts with HubSpot merge operations.

What is crm-data-quality?

This skill identifies data quality issues in HubSpot CRM—missing fields, inconsistent values, and duplicates—then fixes them at scale using bulk operations. Use it to audit and clean contact, company, or deal records before analysis or sync.

  • Search for incomplete records with missing or empty fields using NOT_HAS_PROPERTY filters
  • Normalize field values in bulk by searching, reshaping with jq, and piping into update
  • Deduplicate records by merging secondary contacts into primary ones (irreversible)
  • Audit custom properties and group them by type or category
  • Perform dry-run previews and digest/confirm gating for operations affecting >100 records

How to install crm-data-quality

npx skills add https://github.com/hubspot/agent-cli-skills --skill crm-data-quality
Prerequisites
  • HubSpot CLI installed and authenticated
  • Familiarity with bulk-operations skill (JSONL piping, dry-run/digest/confirm workflow)
  • jq installed for reshaping JSON records
  • Understanding of HubSpot property types and filter syntax
Claude Code
Cursor
Windsurf
Cline

How to use crm-data-quality

  1. 1.List available properties for your object type using `hubspot properties list --type <type>`
  2. 2.Search for incomplete records using `hubspot objects search` with `!fieldname` filters
  3. 3.Preview results with `--dry-run` before executing any update or merge
  4. 4.For normalization: pipe search results through jq to reshape values, then into `hubspot objects update`
  5. 5.For deduplication: collect full record set with pagination loop, group by duplicate key (e.g., email), pipe merge pairs into `hubspot objects merge --dry-run`
  6. 6.For >100 operations: use `--digest` to preview impact, then `--confirm` to execute
  7. 7.Check audit trail with `hubspot history --since 1h` after merges to verify correctness

Use cases

Good for
  • Find all contacts missing email or phone and bulk-update them with normalized values
  • Identify duplicate contacts by email address, merge them, and audit the merge trail
  • Collapse variant company name spellings ("Acme", "ACME Corp", "Acme Corporation") into one canonical value
  • Audit all enumeration properties in your contacts object to find unused or misconfigured fields
  • Create a data-quality flag property and mark records with missing critical fields
Who it's for
  • CRM administrators managing data hygiene
  • Sales operations teams deduplicating contact lists
  • Data analysts preparing CRM data for reporting or integration
  • Teams integrating HubSpot with external systems that require clean data

crm-data-quality FAQ

Why does my deduplication miss duplicates across page boundaries?

The `objects search` API caps at 100 rows per call. You must use the pagination loop from bulk-operations/SKILL.md to collect the full dataset into a file before grouping and merging.

Can I undo a merge?

No, merge is irreversible. However, you can restore the secondary record from HubSpot's recycle bin if you merged in the wrong direction. Always use `--dry-run` first.

How do I find all enumeration properties and their option values?

Use `hubspot properties list --type <type> | jq -c 'select(.type=="enumeration")'` to list enum properties. Option values are not exposed by the CLI; read them from a real record or the HubSpot UI.

What's the difference between `--filter "!email"` and `--filter "email"`?

`!email` finds records with missing or empty email (NOT_HAS_PROPERTY). Bare `email` finds records that have a non-empty email value.

Do I need to create a custom property before normalizing data?

Only if you want to add a new field (e.g., a data-quality flag). For existing properties, search and reshape directly with jq, then update.

Full instructions (SKILL.md)

Source of truth, from hubspot/agent-cli-skills.


name: crm-data-quality description: Find incomplete records, normalize field values in bulk, dedupe with hubspot objects merge, and audit custom properties. Builds on bulk-operations for JSONL piping and dry-run/digest/confirm. triggers:

  • "clean up contacts"
  • "data quality"
  • "deduplicate"
  • "missing fields"
  • "normalize data"
  • "find incomplete records"
  • "merge duplicates"
  • "audit properties"

Read bulk-operations/SKILL.md first — JSONL piping, batch read, pagination, and dry-run/digest/confirm gating apply to every command below.

Property discovery

Don't guess property names. List them:

hubspot properties list --type contacts --format table
hubspot properties list --type contacts | jq -c 'select(.type=="enumeration") | {name, label}'

Same for --type companies, deals, or any custom type (hubspot objects types).

1. Find incomplete records

!name = NOT_HAS_PROPERTY (missing or empty). Bare name = HAS_PROPERTY. Within one --filter, chain with AND; multiple --filter flags are OR'd.

hubspot objects search --type contacts --filter "!email" --properties firstname,lastname,company
hubspot objects search --type contacts --filter "!phone AND !mobilephone" --properties email
hubspot objects search --type contacts --filter "!hubspot_owner_id" --properties email,lifecyclestage

For >100 results, use the pagination loop from bulk-operations.

2. Normalize field values

Search → reshape with jq → pipe into update. Always --dry-run first; bulk-operations covers digest/confirm escalation for >100 rows. Reshape patterns: bulk-operations/resources/json-patterns.md.

# Collapse spellings into one canonical value
hubspot objects search --type contacts --filter "company~acme" \
| jq -c '{id, properties:{company:"Acme Corporation"}}' \
| hubspot objects update --type contacts --dry-run

# Lowercase emails (read, reshape, write)
hubspot objects search --type contacts --filter "email" --properties email \
| jq -c '{id, properties:{email: (.properties.email | ascii_downcase)}}' \
| hubspot objects update --type contacts --dry-run

3. Dedupe with hubspot objects merge

Secondary is folded into primary and deleted. Irreversible. Dry-run/digest/confirm gating applies.

# Single pair
hubspot objects merge --type contacts --primary 149 --secondary 425 --dry-run
hubspot objects merge --type contacts --primary 149 --secondary 425   # execute (≤100 pairs)

Bulk: pipe JSONL {"primary":"...","secondary":"..."} on stdin (omit --primary/--secondary).

Pagination required. objects search caps at 100 rows per call and jq -s slurps a single stream into memory — running the snippet below against a raw search will silently miss every duplicate that crosses a page boundary. Collect the full set first with the pagination loop from bulk-operations/SKILL.md (write to /tmp/contacts.jsonl), then dedupe from the file:

# /tmp/contacts.jsonl produced by the pagination loop (bulk-operations/SKILL.md)
jq -s -c '
    group_by(.properties.email)[]
    | select(length > 1)
    | sort_by(.id | tonumber)
    | .[0].id as $p | .[1:][] | {primary: $p, secondary: .id}
  ' /tmp/contacts.jsonl \
| hubspot objects merge --type contacts --dry-run | tee /tmp/merge-preview.jsonl

For >100 pairs, lift digest and impact.records_affected from the BulkData line and re-pipe the same producer with --digest/--confirm (see bulk-operations).

4. Audit properties

hubspot properties list (and get, batch-read) emits {name, label, type, fieldType, groupName} per row. Enum option values are not currently exposed by the CLI — read them off a real record (hubspot objects search ... --properties <enum>) or the HubSpot UI.

# Count properties per group (HubSpot groups standard fields; custom groups stand out)
hubspot properties list --type contacts | jq -rs 'group_by(.groupName) | map({group: .[0].groupName, count: length}) | .[]'

# All enumeration properties
hubspot properties list --type contacts | jq -c 'select(.type=="enumeration") | {name, label, fieldType}'

# Create a DQ flag property, then set it via the normalize pattern in section 2
hubspot properties create --type contacts --name dq_missing_phone --label "DQ: Missing Phone" --prop-type string --field-type text

Recovery

Merge is irreversible. After any merge, hubspot history --since 1h captures the audit trail. If wrong direction, restore the secondary from the UI's recycle bin.