argent-qa-flows
software-mansion/argent
Create repeatable QA regression E2E tests as Argent flows from test cases and acceptance criteria.
What is argent-qa-flows?
Builds deterministic, automated regression test flows from QA requirements, test cases, or acceptance criteria. Use this when you need a reusable, verifiable test scenario with structural and visual evidence that passes consistently. Supports iOS, Android, Chromium, and Vega (Fire TV); Apple TV and Android TV require argent-tv-interact instead.
- Define a QA contract mapping user actions to executable evidence (structural checks, assertions, snapshots)
- Record deterministic test scenarios with setup, navigation, and verification steps
- Verify state changes, persistence, absence claims, and visual rendering with stable selectors
- Support multiple platforms: iOS (simulator and physical devices), Android, Chromium, and Vega (Fire TV)
- Ensure tests pass twice consecutively with fresh baseline and no manual cleanup required
- Use recorded tv-remote steps for Vega instead of touch directives
How to install argent-qa-flows
npx skills add https://github.com/software-mansion/argent --skill argent-qa-flows- Argent authoring environment (uses argent-create-flow as the engine)
- Test case, ticket, or acceptance criteria with clear user actions and expected outcomes
- For physical iPhones: device udid and stable selectors (may differ from simulator)
- For Vega (Fire TV): understanding of tv-remote navigation instead of touch gestures
How to use argent-qa-flows
- 1.Write a compact QA contract table: app state, user actions, expected outcomes, and executable evidence for each row
- 2.Record the scenario walkthrough using argent-create-flow, capturing structural checks and assertions as state appears
- 3.Add snapshots during polish for visual verification on settled, deterministic screens
- 4.Ensure the first non-echo/script step is launch: with deterministic setup (no manual cleanup needed)
- 5.Run the flow twice consecutively: pass 1 with fresh services, pass 2 immediately after, both unchanged
Use cases
- Automate regression testing for a feature ticket with acceptance criteria (e.g., 'Dark theme selection persists and renders correctly')
- Create reusable test flows for critical user journeys that must pass on every release
- Verify state persistence across app navigation (e.g., settings saved, then re-enter to confirm)
- Test absence claims and negative cases with discriminating evidence (e.g., old UI element hidden after action)
- Build deterministic test suites that work on physical iOS devices by recording with udid and stable selectors
- QA engineers automating regression test suites
- Product teams defining acceptance criteria as executable tests
- Mobile app developers verifying cross-platform behavior (iOS, Android, Chromium, Fire TV)
- Teams needing deterministic, repeatable test scenarios with visual and structural evidence
argent-qa-flows FAQ
Use argent-qa-flows when you have formal acceptance criteria, test cases, or tickets requiring repeatable regression tests with deterministic setup and two consecutive passes. Use argent-test-ui-flow for one-off UI checks or argent-create-flow for replayable paths without acceptance criteria.
Pass the device udid as --device and keep it connected. Selectors authored on a simulator may not match the physical device's flow tree, so read argent-ios-device-interact for the app-scoped contract. Pinch and rotate steps will fail on hardware; use the app's own zoom UI instead.
These platforms are out of scope for argent-qa-flows. Use argent-tv-interact instead and report the limitation. The runner does not reject touch directives there, so they fail at the gesture layer rather than with authoring guidance.
For dynamic content, assert controlled state or stable app chrome; use anchored structure for unavoidable dynamic values. For repeated controls, prefer a stable id. If unavailable, use flow-only within with a stable container and text.in to prove rendered membership.
For state changes, prove both the new state and the old state's absence. For absence, prove the containing screen and record the same stable selector as visible, then the action, then hidden. For persistence, cross the commit boundary: cancel/save, leave, re-enter, then verify the stored state.
Full instructions (SKILL.md)
Source of truth, from software-mansion/argent.
name: argent-qa-flows description: Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria. Use when the user asks to generate or preserve an automated regression scenario, with deterministic setup, stable targets, executable structural or visual evidence, and two consecutive full passes. For one-off UI checks or replayable paths without acceptance criteria, use argent-test-ui-flow or argent-create-flow. Supports iOS, Android, Chromium, and Vega (Fire TV), where recorded tv-remote steps replace touch directives. Apple TV and Android TV are unsupported; use argent-tv-interact there and report the limitation.
Create a QA regression flow
Load argent-create-flow as the authoring engine. Follow its required references for recorder syntax, selectors, polish, platform exceptions, and repair. This skill adds the QA contract and completion gate.
Vega supports every item below: launch: { vega: ... }, await:/assert: selectors, snapshot:, and idle all run there. Only the touch directives are missing, because Vega is remote-driven. Navigate with recorded tool: tv-remote steps and type with tool: keyboard, which leaves item 5 with nothing to govern. A D-pad path is relative to where focus already is, so gate every move with item 4's identity check rather than assuming the cursor landed. Read argent-tv-interact for focus reading and remote navigation.
Apple TV and Android TV are out of scope. The runner does not reject touch directives there, so they fail at the gesture layer instead of with authoring guidance. Use argent-tv-interact and report the limitation.
Physical iPhones run QA flows, with three hardware limits: replay never auto-binds a phone, so pass its udid as device (CLI --device) and keep it connected; pinch/rotate steps fail there like the live tools, so drive the app's own zoom UI instead; the flow tree is the describe tree (same ids and roles), so a selector authored on a simulator can miss there. Read argent-ios-device-interact for the app-scoped contract before recording.
Definition of done
A QA flow is complete only when:
- The first step that is not
echo:orscript:islaunch:. In-flow setup proves a deterministic data baseline. Repeated runs do not accumulate artifacts or require manual cleanup. - The first walkthrough recorded every action and live structural check. Only the three documented polish insertions are unrecorded.
- Every requirement maps to a hard
await:,assert:, or reviewedsnapshot:. Echoes and screenshots are not verdicts. A negative check needs the same stable selector established as visible earlier. - Every screen change has destination identity followed by
idlereadiness. - Targets satisfy the stable-selector and coordinate-fallback rules. QA keeps coordinates only for genuinely unlabeled targets. Vacuous on Vega, which has no coordinate targets.
- The unchanged YAML passes twice with the same runner. Pass 1 starts with fresh mobile Argent services, and pass 2 follows immediately.
1. Define the test contract
Before touching the app, write a compact table. Restate it in the final report. Include:
- App, platform, and named start state.
- Ordered user actions.
- One row for each expected outcome, persistence rule, or absence claim.
- Stable executable evidence for each row.
- Required data and side effects.
Use structural checks for semantic state, snapshots for pixels, and both for mixed requirements. One behavioral scenario becomes one qa-<area>-<behavior> flow.
Do not invent a material value or weaken ambiguity. Choose the strongest UI-verifiable reading and report it. Ask when the choice changes test meaning.
Make repeated runs deterministic:
- Inspect the required baseline without mutation.
- If the account is dirty, record a safe reset or seed flow. Alternatively, include safe normalization in setup.
- After setup navigation, echo the named baseline and hard-check it before the first scenario mutation. Use
assert:or a destinationawait:that fully proves the baseline. - Prefer to restore the baseline at the end.
A flow has two fixture mechanisms:
run:replays a separately recorded reset or seed flow.script:runs requested local setup or cleanup. Record it withflow-add-scriptwhere it belongs in the walkthrough.
Ask before cleanup that creates or deletes meaningful user data outside the request.
Compact example
Ticket: select Dark in Settings. Verify Dark is selected, Light is absent, and the screen renders in dark mode.
| Contract row | Action | Evidence | State effect |
|---|---|---|---|
| Signed-in Home | Launch | await: { visible: { id: home-screen } }, then await: { idle: true } | Existing account |
| Open Settings | Tap settings-tab | await: { visible: { id: settings-screen } }, then await: { idle: true } | None |
| Prove Light selected | Inspect Settings | assert: { visible: { id: theme-light-selected } } | Fails if already Dark |
| Prove Dark selected | Tap theme-dark-option | await: { visible: { id: theme-dark-selected } } | Theme becomes Dark |
| Prove Light absent | Inspect settled screen | assert: { hidden: { id: theme-light-selected } } | None |
| Verify dark rendering | Inspect settled screen | snapshot: settings-dark | None |
| Restore baseline | Tap theme-light-option | await: { visible: { id: theme-light-selected } } | Next run starts clean |
The initial Light check establishes the selector used by the later hidden check. The final restore makes pass 2 independent.
2. Record the scenario
Follow argent-create-flow's start order and live-authoring cycle. Record each structural contract check when its state appears.
A snapshot has no recorder form. Inspect its stable state during the walkthrough, then add the planned snapshot during polish. If direct recovery changes state, re-record the affected behavior. A recovered walkthrough is not proof.
3. Make evidence discriminating
- State change: prove the new state and the old state's absence when both can otherwise match.
- Cancel/persistence: cross the commit boundary. After cancel or save, leave, re-enter, then verify the stored state: unchanged after cancel or updated after save.
- Absence: prove the containing screen and record the same stable selector as
visible, then the action, thenhidden. Do not add an unestablishedhiddencheck only to strengthen a positive baseline. In a collection, viewport absence is not global absence. Use fixed seeded position, count, empty state, or other collection-wide evidence. - Overlays: use the create-flow obscured-target procedure.
- Repeated controls: prefer an id. Otherwise use flow-only
withinwith a stable container. Usetext.into prove rendered membership inside that container. - Dynamic content: assert controlled state or stable app chrome. Use anchored structure for unavoidable dynamic values and disclose the dependency.
- Visual state: snapshot only a correct, settled, deterministic screen. Use full screen for global changes and
cropOnfor one component.
Never put acceptance evidence inside when:. Use when: only for optional setup that reconverges to the required path.
4. Finish and audit
Complete the create-flow polish and blocking audit. Then:
- Map every contract row to an executed action or hard check.
- Build a navigation table with one row per screen change, naming both the identity gate and the readiness gate. A row missing either is a blocking defect.
- Confirm setup and end state permit an immediate second run.
| Action | Destination | Identity | Readiness |
|---|---|---|---|
Tap settings-tab | Settings | settings-screen visible | idle |
The two are repaired differently. A missing identity check must be recorded live on the restored screen. A missing idle check is added in YAML, because await: { idle: true } has no recorder form and is one of argent-create-flow's three permitted polish insertions. Re-record any missing action or other structural check.
5. Prove two consecutive passes
After the last edit and audit, set the streak to zero:
- Choose one runner for both passes. Use
flow-executelocally orargent flow run <name> --platform <platform>for CI. Switching runners resets the streak. - Seed, review, and freeze snapshot baselines. Baseline updates do not count as passes.
- Before mobile pass 1, recycle Argent services for this flow's device: two warm passes are correlated evidence, because a fixed timing margin can pass twice simply because environment speed did not change. Scope
stop-all-simulator-serverstodevices: [<device>]. Never omit the scope — a bare call is the machine-wide sweep, and step 7 restarts this proof often enough to reap every other agent's devices repeatedly. Use the MCP call forflow-execute, orargent run stop-all-simulator-servers --devices <device>from the standalone runner's install. The reset must not change app or account data. For Chromium, let the runner boot the declared app and omitdevice. Vega owns no recyclable Argent services, so the teardown is a no-op there and both passes are warm. - Run from the flow's launch and setup without baseline-update mode. Count a pass only when
ok: trueand every acceptance check executed. A falsewhen:can skip optional setup only. An errored step does not advance the streak, and the count mixes two kinds — read each reason. One that could not run (an unreadable tree underidle, an unresolvablerun:target) is environment: fix it and rerun. A failedlaunch:also scoreserrored, and it is a verdict about the app — an app that no longer installs or starts is the regression this test exists to catch, so report it instead of rerunning. - Resolve every passing-step warning before completion. Also resolve recorded-wait warnings from
flow-finish-recording. Follow Live waits and checks. For runner warnings,await: { idle: true }raises six different warnings, so read which one it is first. Two say the screen was moving. One says the wait ran out mid-hold and needs a largertimeout:. One says the tree stayed empty. One — settled on the UI tree alone — says the hierarchy did hold still and only the screenshot pairs were missing, so inspect the capture path rather than the app's rendering. One says the step ended with no evidence either way. Inspect the screen, disclose the cause, and verify that surrounding acceptance checks use stable elements rather than stillness. A selector-less gesture — a coordinatetap/long-press/swipe, or apinch/rotatewith noon:— warns in a different shape: a tree-source outage left it unsettled, so it dispatched blind and the green says only that the gesture was sent. Restore the tree source, usually by relaunching the app so the instrumentation loads, and rerun. Accepting that warning needs an app that serves no tree, which cannot satisfy this contract anyway. - Run the same YAML again immediately with the same runner. Do not manually reset app or account data.
- Reset the streak after any failure, edit, re-recording, baseline update, or state-changing manual recovery. Repair through
argent-create-flow, audit again, and restart with fresh services.
Finish only when the streak reaches two. If the intended runner is unavailable, report proof as blocked. If product behavior fails, keep the strong check and report the regression. Never weaken it to obtain green output.
Record the runner and fresh-service setup used.
6. Report
Report:
- Flow name, path, platform, and standalone command.
- Contract rows mapped to actions and checks.
- Navigation table.
- Baseline setup, end-state restoration, and accepted data dependencies.
- Both pass results, runner, fresh-service setup, and resolved warnings.
- Snapshot scope, reviewed baseline status, and mismatch tolerance.
- Coordinate or raw-gesture exceptions.
- Remaining manual judgment or blocker.
Related skills
More from software-mansion/argent and the wider catalog.

argent-react-native-app-workflow
Step-by-step workflows for developing and debugging React Native apps on iOS simulator or Android emulator.

argent-react-native-optimization
Profile React Native apps to find real bottlenecks, then fix mechanical issues systematically.

argent-react-native-profiler
Profile React Native Hermes apps to measure re-render and CPU performance with ranked issue reports.

argent-screen-recording
Record video of iOS simulator or Android device screens with touch visualization and automatic static-frame trimming.

argent-screenshot-diff
Compare app screenshots to detect visual regressions and UI changes pixel-by-pixel.

argent-settings-permissions
Control iOS and Android app permissions directly without navigating Settings—for test setup when the app can't change them itself.