On this page
Test changes against a pinned version
Create repeatable cases for verification, missing account data, tool failures and escalation. Run them against a specific publication, then inspect the evidence behind each failed assertion.
This billing-support example uses synthetic cases, verdicts and costs in Studio. Select any screenshot to open it at full size.
Save the scenario and assertions
The caller asks about the account without immediately confirming identity.
The structural assertion checks node visits; the judged criterion checks account disclosure against the conversation.
Open the agent's Tests tab and create a suite. For “Verify before account disclosure”:
- Give the simulated caller a goal: ask about an invoice without immediately confirming identity.
- Write the caller behavior: ask who is calling; confirm identity only after the purpose is explained.
- Add a node assertion for the relevant verification and lookup nodes. Visiting nodes is a structural check; it is not, alone, proof that they occurred in the correct business order.
- Add a judged criterion: “Did the agent verify identity before sharing account details?” Define a threshold and inspect the judge's reasoning when it fails.
- Mock
lookup_accountwith a fixed synthetic response so the test does not depend on a production account record. - Save the case and validate it against the target draft or publication.
Mock each customer HTTP action the case needs. An unmocked customer HTTP action fails with unmocked_action before sending a request. Built-in actions have simulated defaults.
Use a separate case for each meaningful failure path. A useful starting suite includes caller refusal, wrong person, missing account, lookup timeout, unavailable transfer and a request outside policy. Your assertion choice must match the behavior you want to establish; a judge is probabilistic and can be wrong.
Read test suite creation, suite edits and suite validation for the current case, assertion and mock fields.
Run the exact candidate
The test run records the publication and suite revision that were evaluated. The failed case links to its individual trial.
In Studio, selecting the draft in the run dialog first publishes it as a test version with live: false, then estimates and starts the suite against that version. Through the API, call Publish an agent with live: false yourself, then pass the returned version to test_runs.estimate and test_runs.create. Review the trial count and estimate before starting, and follow the test run until it settles.
The record ties together the agent publication, suite revisions, case results and trials. Save a baseline test-run ID when you want to compare another candidate against it. A changed suite changes what you are measuring: inspect both the agent version and the suite revision before interpreting an improvement.
The estimate, create, results and trial endpoints support the same workflow from code. Build and test recipes shows the publication-first sequence.
Inspect the failing evidence
The failed criterion points to a specific turn and explains the verdict. In this synthetic example, account details appeared before verification.
Open the failing case, choose a trial and select the assertion. Review its evidence turns, node path, action calls and judge rationale. Determine whether the failure is in the agent, the scenario, the assertion or the judge before editing the flow.
For this case, move account disclosure after verification and cover the refusal path explicitly. Publish a new test version and run the same suite again. Add additional trials when you need to understand variability; one passing result does not establish a reliable pass rate.
Decide whether to promote
Use the result as one part of release acceptance:
| Check | What it establishes |
|---|---|
| Draft validation | The definition meets compiler checks |
| Simulation assertions and evidence | The tested conversational behaviors under the configured scenarios |
| Browser voice test | Behavior through a browser voice session |
| Authorized carrier call | Real number routing and phone/media behavior |
| Real integration acceptance | The intended downstream service works with the actual credentials and business contract |
Simulation mocks do not prove live tool delivery. Text-based simulations do not prove audio quality or carrier routing. Promote the tested publication only after the acceptance checks appropriate to your use case pass, then inspect real run evidence.
Keep the test-run ID with your release review. It links the publication, suite revision and trial evidence you used to approve the change.



