PluginBench
Skill
Pass
Audit score 90

data360-prepare

forcedotcom/sf-skills

Manage Salesforce Data Cloud data streams, DLOs, transforms, and Document AI ingestion.

What is data360-prepare?

This skill handles the Prepare phase of Salesforce Data Cloud: creating and managing data streams, Data Lake Objects (DLOs), transforms, and Document AI configurations. Use it when setting up ingestion pipelines, configuring how data enters Data Cloud, or preparing unstructured document sources—but delegate connection setup to data360-connect, identity resolution to data360-harmonize, and queries to data360-query.

  • Create, inspect, and run data streams from CRM, database connectors, and Ingestion API sources
  • Manage Data Lake Objects (DLOs) and configure their category (Profile, Engagement, or Other)
  • Set up unstructured document ingestion from SharePoint and similar sources
  • Configure Document AI and transform pipelines for data preparation
  • Diagnose ingestion readiness and health before handoff to harmonization

How to install data360-prepare

npx skills add https://github.com/forcedotcom/sf-skills --skill data360-prepare
Prerequisites
  • Node.js ≥18.0.0
  • Python 3 ≥3.10.0
  • Salesforce CLI (sf) ≥2.0.0
  • pip ≥21.0
  • Access to a Salesforce org with Data Cloud provisioned
  • Source connection already created (via data360-connect skill)
Claude Code
Cursor
Windsurf
Cline

How to use data360-prepare

  1. 1.Run the readiness classifier to verify your org is ready for prepare work: `node ../data360-orchestrate/scripts/diagnose-org.mjs -o <org> --phase prepare --json`
  2. 2.List existing data streams and DLOs to understand current ingestion assets: `sf data360 data-stream list -o <org>` and `sf data360 dlo list -o <org>`
  3. 3.Determine the stream category (Profile, Engagement, or Other) based on your dataset type
  4. 4.Create or inspect a data stream using `sf data360 data-stream create-from-object` or `sf data360 data-stream get`
  5. 5.Run a stream refresh with `sf data360 data-stream run -o <org> --name <stream>` when needed
  6. 6.For unstructured sources, create a DLO with the appropriate directory and file configuration
  7. 7.Verify stream and DLO health before handing off to data360-harmonize for DMO mapping and identity resolution

Use cases

Good for
  • Setting up a new data stream from Salesforce CRM or an external database connector
  • Configuring an unstructured DLO to ingest documents from SharePoint
  • Re-scanning or refreshing an existing data stream after source updates
  • Preparing an Ingestion API-backed stream to receive external system records
  • Validating stream and DLO health before mapping to DMOs and identity resolution
Who it's for
  • Data engineers building Data Cloud ingestion pipelines
  • Salesforce administrators setting up data connectors and streams
  • Teams integrating external data sources into Data Cloud
  • Organizations preparing unstructured document ingestion

data360-prepare FAQ

When should I use data360-prepare vs. data360-connect?

Use data360-connect to create and test source connections. Use data360-prepare once the connection is working and you're ready to create data streams, DLOs, and configure ingestion into Data Cloud.

What's the difference between data-stream run and connection run-existing?

data-stream run refreshes a specific stream and is the preferred method for unstructured document rescans. connection run-existing runs at the connection level and is less reliable for unstructured sources; use it only for specific connector workflows.

How do I ingest data from an external system using the Ingestion API?

Create the connector in data360-connect, upload the schema with `sf data360 connection schema-upsert`, create the stream (often in the UI), then use the local Ingestion API example in `examples/ingestion-api/` to send records with JWT authentication and a Data Cloud token.

What happens if I delete a data stream?

Deleting a data stream can also delete the associated DLO unless you specify otherwise in the delete mode. Verify your deletion intent before proceeding.

When should I use the UI instead of the CLI for stream creation?

Use the UI for initial unstructured document setup (SharePoint, etc.) when you need the richer end-to-end pipeline with additional metadata fields. Some external database connectors also require UI or browser automation for first-time stream creation.

Full instructions (SKILL.md)

Source of truth, from forcedotcom/sf-skills.


name: data360-prepare description: "Salesforce Data Cloud Prepare phase. Use this skill when the user creates or manages Data Cloud data streams, DLOs, transforms, or Document AI configurations. TRIGGER when: user creates or manages Data Cloud data streams, DLOs, transforms, or Document AI configurations, or asks about ingestion into Data Cloud. DO NOT TRIGGER when: the task is connection setup only (use data360-connect), DMOs and identity resolution (use data360-harmonize), or query/search work (use data360-query)." metadata: cliTools: - tool: ["node"] semver: ">=18.0.0" - tool: ["pip"] semver: ">=21.0" - tool: ["python3"] semver: ">=3.10.0" - tool: ["sf"] semver: ">=2.0.0" relatedSkills: - "data360-connect" - "data360-harmonize" - "data360-orchestrate" - "data360-query" version: "1.0" domains: ["Data 360"]

data360-prepare: Data Cloud Prepare Phase

Use this skill when the user needs ingestion and lake preparation work: data streams, Data Lake Objects (DLOs), transforms, Document AI, unstructured ingestion, or the handoff from connector setup into a live stream.

When This Skill Owns the Task

Use data360-prepare when the work involves:

  • sf data360 data-stream *
  • sf data360 dlo *
  • sf data360 transform *
  • sf data360 docai *
  • choosing how data should enter Data Cloud
  • rerunning or rescanning ingestion after a source update
  • preparing Ingestion API-backed streams after connector setup is complete

Delegate elsewhere when the user is:

  • still creating/testing source connections → data360-connect
  • mapping to DMOs or designing IR/data graphs → data360-harmonize
  • querying ingested data → data360-query

Required Context to Gather First

Ask for or infer:

  • target org alias
  • source connection name
  • source object / dataset / document source
  • desired stream type
  • DLO naming expectations
  • whether the user is creating, updating, running, or deleting a stream
  • whether the source is CRM, a database connector, an unstructured file source, or an Ingestion API feed

Core Operating Rules

  • Verify the external plugin runtime before running Data Cloud commands.
  • Run the shared readiness classifier before mutating ingestion assets: node ../data360-orchestrate/scripts/diagnose-org.mjs -o <org> --phase prepare --json.
  • Prefer inspecting existing streams and DLOs before creating new ingestion assets.
  • Suppress linked-plugin warning noise with 2>/dev/null for normal usage.
  • Treat DLO naming and field naming as Data Cloud-specific, not CRM-native.
  • Confirm whether each dataset should be treated as Profile, Engagement, or Other before creating the stream.
  • Distinguish stream-level refresh from connection-level reruns when working with unstructured sources.
  • Use UI setup intentionally when initial stream or unstructured asset creation is platform-gated.
  • Hand off to Harmonize only after ingestion assets are clearly healthy.

Recommended Workflow

1. Classify readiness for prepare work

node ../data360-orchestrate/scripts/diagnose-org.mjs -o <org> --phase prepare --json

2. Inspect existing ingestion assets

sf data360 data-stream list -o <org> 2>/dev/null
sf data360 dlo list -o <org> 2>/dev/null

3. Confirm the stream category before creation

Use these rules when suggesting categories:

CategoryUse forTypical requirement
Profileperson/entity recordsprimary key
Engagementtime-based events or interactionsprimary key + event time field
Otherreference/configuration/supporting datasetsprimary key

When the source is ambiguous, ask the user explicitly whether the dataset should be treated as Profile, Engagement, or Other.

4. Create or inspect streams intentionally

sf data360 data-stream get -o <org> --name <stream> 2>/dev/null
sf data360 data-stream create-from-object -o <org> --object Contact --connection SalesforceDotCom_Home 2>/dev/null
sf data360 data-stream create -o <org> -f stream.json 2>/dev/null
sf data360 data-stream run -o <org> --name <stream> 2>/dev/null

5. Check DLO shape

sf data360 dlo get -o <org> --name Contact_Home__dll 2>/dev/null

6. Choose the right refresh mechanism

Use the smaller refresh scope that matches the user goal:

sf data360 data-stream run -o <org> --name <stream> 2>/dev/null
sf data360 connection run-existing -o <org> --name <connection-id> 2>/dev/null
  • data-stream run is the closest match to a stream-level refresh or re-scan.
  • connection run-existing runs at the connection level and can be useful for some connector workflows, but it is not a reliable replacement for stream refresh on unstructured sources.
  • For unstructured document connectors, prefer data-stream run when the goal is to re-scan newly added or changed files.

7. Handle unstructured sources deliberately

For SharePoint-style document ingestion, a minimal unstructured DLO payload can look like:

{
  "name": "my_udlo",
  "label": "My UDLO",
  "category": "Directory_Table",
  "dataSource": {
    "sourceType": "SF_DRIVE",
    "directoryAndFilesDetails": [
      {
        "dirName": "SPUnstructuredDocument/<CONNECTION_ID>/<SITE_ID>",
        "fileName": "*"
      }
    ],
    "sourceConfig": {
      "reservedPrefix": "$dcf_content$"
    }
  }
}

Use the UI for the first-time unstructured setup when the user needs the richer end-to-end pipeline. The UI path can seed additional document metadata fields and downstream assets that a bare CLI DLO create flow may not provision automatically.

8. Use the local Ingestion API example for send-data workflows

For external systems pushing records into Data Cloud:

  1. create the connector in data360-connect
  2. upload the schema with sf data360 connection schema-upsert
  3. create the stream in the UI when required
  4. send records with the local example in examples/ingestion-api/
cd examples/ingestion-api
cp .env.example .env
python3 send-data.py

Key details:

  • auth is a staged flow: JWT → Salesforce token → Data Cloud token
  • the ingestion endpoint uses the tenant URL, not the Salesforce instance URL
  • 202 means the payload was accepted for processing, not that records are queryable immediately
  • validation failures often surface in the Problem Records DLO family

9. Only then move into harmonization

Once the stream and DLO are healthy, hand off to data360-harmonize.


High-Signal Gotchas

  • CRM-backed stream behavior is not the same as fully custom connector-framework ingestion.
  • sf data360 data-stream run and sf data360 connection run-existing are not interchangeable; prefer stream-level refresh for unstructured rescans.
  • SFDC streams sync on a platform-managed schedule; data-stream run is not the general control path for CRM connector refresh.
  • Some external database connectors can be created via API while stream creation still requires UI flow or org-specific browser automation. Do not promise a pure CLI stream-creation path for every connector type.
  • Initial SharePoint-style unstructured setup can be richer in the UI than in a minimal CLI DLO create flow.
  • Stream deletion can also delete the associated DLO unless the delete mode says otherwise.
  • DLO field naming differs from CRM field naming, including __c → _c transformations.
  • Query DLO record counts with Data Cloud SQL instead of assuming list output is sufficient.
  • CdpDataStreams means the stream module is gated for the current org/user; guide the user to provisioning/permissions review instead of retrying blindly.

Output Format

Prepare task: <stream / dlo / transform / docai>
Source: <connection + object>
Target org: <alias>
Artifacts: <stream names / dlo names / json definitions>
Verification: <passed / partial / blocked>
Next step: <harmonize or retrieve>

References

  • examples/ingestion-api/README.md
  • ../data360-orchestrate/assets/definitions/data-stream.template.json
  • ../data360-orchestrate/references/plugin-setup.md
  • ../data360-orchestrate/references/feature-readiness.md