agentforce-observe
forcedotcom/sf-skills
Analyze production Agentforce agent behavior via session traces, Data Cloud, and Agent Health Monitoring alerts.
What is agentforce-observe?
Observe production Agentforce agent performance by querying session traces and Data Cloud records, reproduce issues using live preview, and manage Agent Health Monitoring (AHM) alerts on agent metrics. Use this when investigating production failures, regressions, performance issues, or setting up monitoring for escalation rates and deflection metrics.
- Query STDM session data and Data Cloud trace records to analyze agent behavior in production
- Reproduce reported issues using sf agent preview with live conversation simulation
- Manage Agent Health Monitoring (AHM) alerts on agent metrics (escalation rate, deflection, etc.)
- Create, list, update, and delete AHM data alerts with notification tracking
- Investigate agent failures, regressions, and performance metrics across sessions
- Resolve agent names and locate .agent files from org or local project
How to install agentforce-observe
npx skills add https://github.com/forcedotcom/sf-skills --skill agentforce-observe- Authenticated Salesforce org (sf org login web if needed)
- Agent API name (MasterLabel or DeveloperName)
- Data Cloud with STDM (Session Trace Data Model) activated, or fallback to test suites
- CLI tools: git >=2.0.0, jq >=1.6.0, python3 >=3.10.0, sf >=2.136.8
How to use agentforce-observe
- 1.Gather required inputs: org alias, agent API name, optional session IDs or days to look back
- 2.Run Phase 0 to discover the correct Data Cloud Data Space (default or custom)
- 3.Execute Phase 1 (Observe) to query STDM sessions and surface issues from the last 7 days
- 4.If issues found, run Phase 2 (Reproduce) using sf agent preview to simulate the problematic conversation
- 5.Edit the .agent file directly to fix issues, then validate and publish
- 6.For AHM alerts: create with specific metric thresholds, list existing alerts, update conditions, or delete as needed
- 7.Check alert notifications and Incidents view to confirm alerts are firing correctly
Use cases
- Investigate why a production agent is failing or has regressed in performance
- Analyze session traces to understand conversation flow and identify bottlenecks
- Set up alerts to monitor escalation rates and deflection metrics over time
- Reproduce a customer-reported issue in preview before making fixes
- Check whether AHM alerts have fired and review notification counts
- Agentforce developers troubleshooting production agents
- DevOps engineers setting up monitoring and alerting for agents
- Support teams investigating customer-reported agent failures
- Product managers tracking agent health metrics and performance trends
agentforce-observe FAQ
agentforce-generate is for creating and debugging .agent files during development; agentforce-observe is for analyzing production behavior via session traces and managing monitoring alerts.
If STDM is not available, the skill falls back to test suites and local preview traces. Enable STDM in Setup > Data Cloud > Data Streams under 'Agentforce Activity'.
Use Phase 2 (Reproduce): run sf agent preview with the agent API name to simulate conversations live and test fixes before deploying.
AHM lets you create data alerts on agent metrics (escalation rate, deflection, etc.) and receive notifications when thresholds are crossed; use this skill to create, list, update, and delete those alerts.
Check the alert status and metric values using the Incidents view; verify the threshold condition matches your metric data and that notifications are enabled for your user.
Full instructions (SKILL.md)
Source of truth, from forcedotcom/sf-skills.
name: agentforce-observe description: "Analyze production Agentforce agent behavior using session traces and Data Cloud, and manage Agent Health Monitoring (AHM) alerts. TRIGGER when: user queries STDM session data or Data Cloud trace records; investigates production agent failures, regressions, or performance issues; asks about session traces, conversation logs, or agent metrics; wants to reproduce a reported production issue in preview; runs findSessions or trace analysis queries; creates, lists, updates, or deletes an AHM data alert on an agent metric (escalation rate, deflection, etc.); asks why an alert is not firing or wants to check whether alerts have fired (notification counts, or the per-alert Incidents view). DO NOT TRIGGER when: user creates, modifies, or debugs .agent files during development (use agentforce-generate); writes or runs test specs (use agentforce-test); uses sf agent preview for local development iteration; deploys or publishes agents." allowed-tools: Bash Read Write Edit Glob Grep metadata: relatedSkills: - "agentforce-generate" - "agentforce-test" version: "0.9" domains: ["Agentforce", "Data 360"] cliTools: - tool: ["git"] semver: ">=2.0.0" - tool: ["jq"] semver: ">=1.6.0" - tool: ["python3"] semver: ">=3.10.0" - tool: ["sf"] semver: ">=2.136.8"
Agentforce Observability
Improve Agentforce agents using session trace data and live preview testing.
Three-phase workflow:
- Observe -- Query STDM sessions from Data Cloud (if available), OR run test suites + preview with local traces as fallback
- Reproduce -- Use
sf agent previewto simulate problematic conversations live - Improve -- Edit the
.agentfile directly, validate, publish, verify
Platform Notes
- Shell examples below use bash syntax. On Windows, use PowerShell equivalents or Git Bash.
- Replace
python3withpythonon Windows. - Replace
/tmp/with$env:TEMP\(PowerShell) or%TEMP%\(cmd). - Replace
jqwithpython -c "import json,sys; ..."if jq is not installed.
Routing
Gather these inputs before starting:
- Org alias (required) -- must be authenticated (else
sf org login web) - Agent API name (required for preview and deploy; ask if not provided)
- Agent file path (optional) -- path to the
.agentfile, typicallyforce-app/main/default/aiAuthoringBundles/<AgentName>/<AgentName>.agent. Auto-detect if not provided. - Session IDs (optional) -- analyze specific sessions; if absent, query last 7 days
- Days to look back (optional, default 7)
- Alert owner user (optional, alerts only) -- user whose alerts to list/manage; defaults to the current user
Determine intent from user input:
- No specific action -> run all three analysis phases: Observe -> surface issues -> ask if user wants to Reproduce and/or Improve
- "analyze" / "sessions" / "what's wrong" -> Phase 1 only, then suggest next steps
- "reproduce" / "test" / "preview" -> Phase 2 (run Phase 1 first if no issues in hand)
- "fix" / "improve" / "update" -> Phase 3 (run Phase 1 first if no issues in hand)
- "create alert" / "set up monitoring" / "alert me when" -> AHM (create)
- "list alerts" / "show my alerts" / "update alert" / "delete alert" -> AHM (list / update via PUT / delete)
- "get / list notifications" / "notifications for a specific alert" / "have my alerts fired" / "why isn't my alert firing" -> AHM: fetch notifications with header
X-UNS-Type-Filter: all, report status + list; for a specific alert, filter by alertId fromtargetPageRef.state.c__alertId(15/18-char-safe), not metricId (+ Incidents view + metric verify)
Resolve agent name
Before any STDM query, resolve the user-provided agent name against the org to get the exact MasterLabel and DeveloperName:
sf data query --json \
--query "SELECT Id, MasterLabel, DeveloperName FROM GenAiPlannerDefinition WHERE MasterLabel LIKE '%<user-provided-name>%' OR DeveloperName LIKE '%<user-provided-name>%'" \
-o <org>
MasterLabel= display name used by STDMfindSessionsand Agent Builder UI (e.g. "Order Service")DeveloperName= API name with version suffix used in metadata (e.g. "OrderService_v9")- The
--api-nameflag forsf agent preview/activate/publishusesDeveloperNamewithout the_vNsuffix (e.g. "OrderService")
Store these values:
AGENT_MASTER_LABEL-- forfindSessions()agent filterAGENT_API_NAME--DeveloperNamewithout_vNsuffix, forsf agentCLI commandsPLANNER_ID-- the Salesforce record ID for this agent
Locate the .agent file
Step 1 -- Search locally:
find <project-root>/force-app/main/default/aiAuthoringBundles -name "*.agent" 2>/dev/null
If the user provided an agent file path, use that directly. Otherwise, search for files matching AGENT_API_NAME.
Step 2 -- If not found locally, retrieve from the org:
sf project retrieve start --json --metadata "AiAuthoringBundle:<AGENT_API_NAME>" -o <org>
Known bug:
sf project retrieve startcreates a double-nested path:force-app/main/default/main/default/aiAuthoringBundles/.... Fix it immediately after retrieve:
if [ -d "force-app/main/default/main/default/aiAuthoringBundles" ]; then
mkdir -p force-app/main/default/aiAuthoringBundles
cp -r force-app/main/default/main/default/aiAuthoringBundles/* \
force-app/main/default/aiAuthoringBundles/
rm -rf force-app/main/default/main
fi
Step 3 -- Validate the retrieved file:
Read the .agent file and verify it has proper Agent Script structure:
system:block withinstructions:config:block withdeveloper_name:start_agentorsubagentblocks withreasoning: instructions:- Each subagent should have distinct
instructions:content (not identical across subagents)
Store the resolved path as AGENT_FILE for Phase 3.
Phase 0: Discover Data Space
Before running any STDM query, determine the correct Data Cloud Data Space API name.
sf api request rest "/services/data/v63.0/ssot/data-spaces" -o <org>
Note: sf api request rest is a beta command -- do not add --json (that flag is unsupported and causes an error).
The response shape is:
{
"dataSpaces": [
{
"id": "0vhKh000000g3DjIAI",
"label": "default",
"name": "default",
"status": "Active",
"description": "Your org's default data space."
}
],
"totalSize": 1
}
The name field is the API name to pass to AgentforceOptimizeService.
Decision logic:
- If the command fails (e.g. 404 or permission error), fall back to
'default'and note it as an assumption. - Filter to only
status: "Active"entries. - If exactly one active Data Space exists, use it automatically and confirm to the user: "Using Data Space:
<name>". - If multiple active Data Spaces exist, show the list (label + name) and ask the user which to use.
Store the selected name value as DATA_SPACE for all subsequent steps.
Prerequisite check: STDM DMOs
After deploying the helper class (step 1.0), run a quick probe to verify the STDM Data Model Objects exist in Data Cloud:
sf apex run -o <org> -f /dev/stdin << 'APEX'
ConnectApi.CdpQueryInput qi = new ConnectApi.CdpQueryInput();
qi.sql = 'SELECT ssot__Id__c FROM "ssot__AiAgentSession__dlm" LIMIT 1';
try {
ConnectApi.CdpQueryOutputV2 out = ConnectApi.CdpQuery.queryAnsiSqlV2(qi, '<DATA_SPACE>');
System.debug('STDM_CHECK:OK rows=' + (out.data != null ? out.data.size() : 0));
} catch (Exception e) {
System.debug('STDM_CHECK:FAIL ' + e.getMessage());
}
APEX
If STDM_CHECK:FAIL: STDM is not activated. Inform the user and switch to Phase 1-ALT:
STDM (Session Trace Data Model) is not available in this org. To enable: Setup -> Data Cloud -> Data Streams and verify "Agentforce Activity" is active. Proceeding with fallback: test suites + local traces.
If STDM_CHECK:OK, proceed to Phase 1 (STDM path).
Phase 1-ALT: Observe Without STDM (Fallback Path)
When STDM is not available, use test suites and sf agent preview --authoring-bundle with local trace analysis.
| Data source | When to use | Pros | Cons |
|---|---|---|---|
| STDM (Phase 1) | Historical production analysis | Real user data, volume | Requires Data Cloud, 15-min lag |
| Test suites + local traces (Phase 1-ALT) | Dev iteration, orgs without STDM | Instant, full LLM prompt, variable state | Preview only, no real user data |
1-ALT.1 Run existing test suite (if available)
sf agent test list --json -o <org>
sf agent test run --json --api-name <TestSuiteName> --wait 10 --result-format json -o <org> | tee /tmp/test_run.json
JOB_ID=$(python3 -c "import json; print(json.load(open('/tmp/test_run.json'))['result']['runId'])")
sf agent test results --json --job-id "$JOB_ID" --result-format json -o <org>
1-ALT.2 Derive test utterances from .agent file (if no test suite)
If no test suite exists, derive utterances: one per non-entry subagent (from description: keywords), one per key action, one guardrail test, one multi-turn test.
1-ALT.3 Preview with --authoring-bundle (local traces)
Run each test utterance through preview to generate local trace files:
sf agent preview start --json --authoring-bundle <BundleName> --simulate-actions -o <org> | tee /tmp/preview_start.json
SESSION_ID=$(python3 -c "import json; print(json.load(open('/tmp/preview_start.json'))['result']['sessionId'])")
sf agent preview send --json --session-id "$SESSION_ID" --authoring-bundle <BundleName> \
--utterance "$UTT" -o <org> | tee /tmp/preview_response.json
sf agent preview end --json --session-id "$SESSION_ID" --authoring-bundle <BundleName> -o <org>
Trace file location: .sfdx/agents/{BundleName}/sessions/{sessionId}/traces/{planId}.json
1-ALT.4 Local trace diagnosis
| Issue type | Trace command |
|---|---|
| Subagent misroute | jq -r '.plan[] | select(.type=="NodeEntryStateStep") | .data.agent_name' "$TRACE" |
| Action not called | jq -r '.plan[] | select(.type=="EnabledToolsStep") | .data.enabled_tools[]' "$TRACE" |
| LOW adherence | jq -r '.plan[] | select(.type=="ReasoningStep") | {category, reason}' "$TRACE" |
| Variable capture fail | jq -r '.plan[] | select(.type=="VariableUpdateStep") | .data.variable_updates[]' "$TRACE" |
| Vague instructions | jq -r '.plan[] | select(.type=="LLMStep") | .data.messages_sent[0].content' "$TRACE" |
DefaultTopic trace quirk: With --authoring-bundle, the root .topic field often shows "DefaultTopic" even when routing works. Always use NodeEntryStateStep.data.agent_name for the real subagent chain.
Entry answering directly (SMALL_TALK pattern): If start_agent trace shows SMALL_TALK grounding and transition tools visible but none invoked, add "You are a router only. Do NOT answer questions directly." to start_agent instructions.
1-ALT.5 Classify and present
Classify issues using the categories in references/issue-classification.md. After presenting findings, automatically proceed to agent config evidence analysis.
Phase 1: Observe -- Query STDM
Full STDM query details, Apex service deployment, and response parsing: see
references/stdm-queries.md
1.0 Deploy helper class (once per org)
Deploy AgentforceOptimizeService Apex class to the org. Check if already deployed first:
sf data query --json --query "SELECT Id, Name FROM ApexClass WHERE Name = 'AgentforceOptimizeService'" -o <org>
If not deployed, copy from skill directory and deploy. See references/stdm-queries.md for full steps.
1.1 Find sessions
Query recent sessions using findSessions(). Parse DEBUG|STDM_RESULT: from the Apex debug log. If findSessions returns empty, switch to Phase 1-ALT.
1.2 Get conversation details
Use getMultipleConversationDetails() for up to 5 sessions (most recent first). Returns turn-by-turn data with messages, steps, topics, and action results.
1.2b Get LLM prompt/response (optional)
When LOW adherence detected, use getLlmStepDetails() to get the actual LLM prompt and response.
1.2c Get aggregated metrics (recommended first step)
Use getAggregatedMetrics() for high-level health dashboard: session rates, top intents, quality distribution, RAG averages.
1.2d Get moment insights (per-session detail)
Use getMomentInsights() for intent summaries, quality scores (1-5), and retriever metrics per session.
1.2e Run observability queries (RAG deep-dive)
Use runObservabilityQuery() for targeted RAG analysis: KnowledgeGap, Hallucination, RetrievalQuality, AnswerRelevancy, Leaderboard.
1.3 Reconstruct conversations
Render turn-by-turn timeline from ConversationData JSON for each session.
1.4 Identify issues
Full issue pattern table and classification categories: see
references/issue-classification.md
Check each session for: action errors, subagent misroutes, missing actions, wrong inputs, variable capture failures, no transitions, slow actions, LOW adherence, abandoned sessions, dead subagents, publish drift, dead hub anti-pattern, entry answering directly, and safety issues.
Voice agents (has modality voice: block): Also check for:
- Response verbosity — flag any agent response over 3 sentences (voice UX anti-pattern; also a silence/nudge-timer trigger)
- Visual formatting in responses — lists, links, markdown that don't render in speech
- Missing confirmation patterns — actions modifying data without repeating back key details
- Missing voice wiring — voice agent lacks a
VoiceCallIdlinked variable (@VoiceCall.Id) or theconnection customer_web_client:block, or someone added a non-existentconnection voice:block - Latency anti-patterns — cross-reference trace step durations against the field-verified patterns in
/agentforce-generatereferences/voice-latency-heuristics.md: synchronous writes on the live-call path, bulky retrieval returned raw to the reasoning LLM, chained external callouts, over-decomposed subagent routing, and slow actions with no ack phrase. Latency fixes are flag-only unless purely instructional (ack phrase, turn-length, spoken-form rule). - TTS garble / missing spoken-form rule — action outputs or responses that surface prices, phone numbers, or IDs without a spoken-form instruction rule.
Priority: P1 = action errors, misroutes, LOW adherence; P2 = missing actions, variable bugs, knowledge gaps; P3 = performance, abandoned sessions, voice UX issues, voice latency anti-patterns.
1.5 Present findings and agent config evidence
Present sessions analyzed, issues grouped by root cause category, and uplift estimate. Then automatically proceed to analyze the .agent file to confirm root causes.
Full structural analysis checks, cross-reference procedures, and publish drift detection: see
references/issue-classification.md
Retrieve the .agent file from the org, run automated checks (subagent count vs action blocks, dead hub detection, orphan actions, cross-subagent variable dependencies), and cross-reference STDM symptoms against the file structure.
Phase 2: Reproduce -- Live Preview
Full preview procedures, trace diagnosis commands, and classification criteria: see
references/reproduce-reference.md
Build one test scenario per confirmed issue from Phase 1. Run each through sf agent preview with --authoring-bundle (generates local traces). Run each scenario 3 times and classify:
| Verdict | Criteria |
|---|---|
[CONFIRMED] | Same failure in 3/3 runs |
[INTERMITTENT] | Failure in 1-2 of 3 runs |
[NOT REPRODUCED] | Passes in 3/3 runs |
Only [CONFIRMED] and [INTERMITTENT] issues proceed to Phase 3.
Key commands:
sf agent preview start --json --authoring-bundle <Name> --simulate-actions -o <org>
sf agent preview send --json --session-id "$SID" --utterance "<text>" --authoring-bundle <Name> -o <org>
sf agent preview end --json --session-id "$SID" --authoring-bundle <Name> -o <org>
Run these from the Salesforce project directory. start requires an action mode with --authoring-bundle (--simulate-actions or --use-live-actions); that flag is rejected by send and end.
Trace location: .sfdx/agents/{Name}/sessions/{sessionId}/traces/{planId}.json
Phase 3: Improve -- Edit .agent File Directly
Full procedures for pre-flight checks, fix mapping, instruction principles, regression prevention, deployment chain, verification, safety re-verification, and test case creation: see
references/improve-reference.md
3.0 Pre-flight
Verify all action targets exist and are registered in the org before editing. If targets are missing, present options: deploy stubs, remove actions, register via UI, or proceed with routing-only fixes.
3.1-3.3 Map issue, edit, and follow instruction principles
Map each confirmed issue to a fix location in the .agent file (description, instructions, actions, bindings, transitions). Use the Edit tool for targeted changes. Follow instruction principles: name actions explicitly, state pre-conditions, scope tightly, keep persona in system: only.
3.4 Regression prevention
Establish baseline before editing. Make minimal edits. Test immediately after each edit. One fix per publish cycle. Check cross-subagent dependencies. Test adjacent subagents.
3.5 Apply fixes
Read the .agent file, edit with the Edit tool, and show the diff. Preserve the
file's existing structural indentation for a surgical edit so you do not create a
mixed-style file. Generate new files with 4 spaces. If normalization is needed,
convert the entire structural indentation as a separate change and validate it;
do not partially convert a tab-indented file.
3.6 Validate, deploy, publish, activate
# Validate (dry run)
sf agent validate authoring-bundle --json --api-name <AGENT_API_NAME> -o <org>
# Publish (compile + deploy + activate)
sf agent publish authoring-bundle --json --api-name <AGENT_API_NAME> -o <org>
If publish fails, use deploy + activate fallback (note: incomplete -- does not propagate reasoning: actions: to live metadata).
3.7 Verify
Run Phase 2 scenarios post-fix. Check trace for correct routing, grounding, tools, and variables. After 24-48 hours, re-run Phase 1 to compare against baseline.
3.7b Safety re-verification (required)
Re-run safety review (Section 15 of /agentforce-generate) on the modified .agent file. Revert any changes that introduce BLOCK findings.
3.8 Update Testing Center test cases
Create regression test cases from confirmed issues in Testing Center YAML format. Deploy with sf agent test create and verify all previously-broken scenarios pass.
Agent Health Monitoring (AHM)
Full AHM alert procedures -- POST schema, enums, SDM/metric/agent discovery, thresholds, metric verification, troubleshooting: see
references/ahm-alerts.md
Proactive complement to the reactive Observe phases: AHM data alerts fire when an agent metric (escalation, deflection, etc.) crosses a threshold. Driven through sf api request rest against tableau/dataAlerts (dataAlertType: "agenthealthmonitoring") -- there is no sf agent alert subcommand yet. Reuse Phase 0 for dataspace. Build the POST body and discover the SDM, its _mtc metric id, and the agent filter value per the reference first, then:
# List/describe -- resolve owner id first (required); single-alert GET is 405 => list + filter client-side
USER_ID=$(sf org display user --target-org <org> --json | python3 -c 'import sys,json;print(json.load(sys.stdin)["result"]["id"])')
sf api request rest "/services/data/v66.0/tableau/dataAlerts?ownerId=$USER_ID" -o <org>
# Create / PUT update-in-place / delete (204). PUT & DELETE are addressed by $ALERT_ID in the path; PUT takes the full POST-style body. Bodies per reference into a 0600 mktemp file.
sf api request rest "/services/data/v66.0/tableau/dataAlerts" -X POST -H "Content-Type: application/json" -b "@$body" -o <org>
sf api request rest "/services/data/v66.0/tableau/dataAlerts/$ALERT_ID" -X PUT -H "Content-Type: application/json" -b "@$body" -o <org>
sf api request rest "/services/data/v66.0/tableau/dataAlerts/$ALERT_ID" -X DELETE -o <org> --include
# Notifications REQUIRE header X-UNS-Type-Filter: all (else custom-types-only => AHM empty = false 0). Report status + list.
sf api request rest "/services/data/v66.0/connect/notifications/status" -H "X-UNS-Type-Filter: all" -o <org>
sf api request rest "/services/data/v66.0/connect/notifications" -H "X-UNS-Type-Filter: all" -o <org>
# AHM notifications: type=templatized_data_alert; no per-alert endpoint. Attribute by alertId in targetPageRef.state.c__alertId (15/18-char-safe), NOT metricId (collides).
Three silent-failure traps to carry into any alert work:
- Thresholds are raw 0-1 ratios, not display percentages -- 5% is
"0.05","1"means 100%. - POST field names/casing differ from GET -- POST uses
utterance(notalertName) and PascalCasetypediscriminators, so copying a GET response back into a POST fails. - Notifications need
X-UNS-Type-Filter: all-- without itconnect/notifications*returns custom types only, so AHM notifications (typetemplatized_data_alert) come back empty; an empty list is a false negative, not "never fired."
Reference Files
| Reference | Contents |
|---|---|
references/stdm-queries.md | STDM query procedures, Apex service deployment, response parsing |
references/ahm-alerts.md | AHM data-alert create/list/update/delete, trigger history, SDM/metric/agent discovery, threshold schema, metric verification |
references/issue-classification.md | Issue pattern table, root cause categories, structural analysis checks |
references/reproduce-reference.md | Phase 2 preview procedures, trace diagnosis, classification criteria |
references/improve-reference.md | Phase 3 editing, deployment chain, verification, safety, test cases |
references/stdm-schema.md | DMO field schemas, data hierarchy, quality notes, agent name resolution |
Related skills
More from forcedotcom/sf-skills and the wider catalog.

agentforce-test
Write, run, and analyze functional and security test suites for Agentforce agents.

analyzing-omnistudio-dependencies
Detect namespaces, map dependencies, and visualize impact across OmniScripts, FlexCards, Integration Procedures, and Data Mappers.

applying-cms-brand
Search, extract, and apply Salesforce CMS brand guidelines to generated content.

applying-slds
Apply SLDS-compliant UI using the correct blueprints, styling hooks, utility classes, and icons. Use when building any UI that needs SLDS, choosing between Lightning Base Components and SLDS Blueprints, applying styling hooks for theming, using utility classes for layout and spacing, or selecting icons. Triggers include \"build a modal\", \"create a form\", \"data table\", \"SLDS styling\", \"style with hooks\", \"add an icon\".

automation-flow-generate
Generate Salesforce Flows programmatically using a 3-step MCP pipeline.

automation-sandbox-post-copy-config-generate
Convert sandbox-refresh SOPs into JSON config for Salesforce post-copy automation.