On this page
Release and operate a tested agent
Validate a draft, publish a test version, compare it with the baseline and promote it after the required checks pass. Studio keeps the version changes, test results and run records available for review.
The screenshots show synthetic publications, cases, verdicts and costs. Select a screenshot to inspect it at full size.
Keep the candidate separate from live behavior
Edit the draft, validate it and publish a test version with live: false. Existing runs continue with their original publication. New production runs keep using the live version until you change it.
The version list tells you which publication is live, which candidate is a test and which publications were used previously. Review the candidate's source and definition before changing the live version.
Open the change summary and check the actual behavior change. A canvas layout change is different from a prompt, action, condition or provider change. The screenshot's candidate tightens the outcome statement; the review should ask whether it can still claim completion before the external action is confirmed.
Design the acceptance cases
Choose the cases that establish the change is safe. For an order-resolution release, cover authorization refusal, missing account facts, negative eligibility, caller refusal, stale data, uncertain writes and unavailable transfers.
Use structural assertions for the path and actions, and judged criteria for conversational requirements. State the expected behavior, such as “verify identity before sharing invoice details,” in each case.
Mocks give repeatable responses without reaching your business endpoints. An unmocked customer HTTP action fails closed in Cumulus suite simulations. Use real integration acceptance separately to establish credentials, routing and the business transaction contract.
Review results on the exact publication
Check both the agent publication and the suite revision. If the suite changed, the results may no longer measure the same behavior. Read failed, invalid, errored and flaky outcomes separately; they call for different investigation.
For a case error, inspect the test setup and execution failure. For a failed criterion, read the evidence and judge rationale before deciding whether the agent needs a change.
Follow a failure to the deciding evidence
Open a trial and select the assertion. Use the highlighted turn, node path and action record to determine whether the fault is in the agent behavior, test setup, criterion or judge.
In this example, the candidate shared invoice details before verification. Fix the sequence, publish another test version and rerun the same case. Preserve the failed trial as evidence of what the earlier candidate did.
Compare the candidate with its baseline
Use the same case across versions and compare the trial evidence, not only the summary score. The baseline in this illustration waited for verification; the candidate took the lookup path too early. The aligned view helps locate that behavioral difference.
Keep suite revisions, caller behavior, mocks, judge configuration and trial count consistent when the goal is a meaningful comparison. When you change them, record that change rather than calling the result an agent improvement without qualification.
Record a judge disagreement
If the judge is wrong, record the human verdict through the feedback control and retain the supporting explanation in your review process. This feedback is an evaluation of the trial; it does not rewrite the original test result.
Calibrate model judges against cases your team has reviewed. Use deterministic assertions for requirements you can check mechanically.
Set the information your operators will see
Define analysis fields in business terms. An operator should be able to tell whether the case was resolved, what remains to be done and why human review is needed. Distinguish extraction from the actual transaction result: a model-produced resolved flag does not override a failed action receipt.
Choose the per-agent data-storage behavior and review workspace retention. If transcripts or recordings are not retained, an operator may not have the evidence a conversational criterion needs later. Read workspace retention before deciding how long to keep review material.
Promote, then verify the real path
Use publication promotion for the tested version. Preserve the version and suite IDs in your release review. Where a suite is configured as required, review the applicable publication checks and their results.
After promotion, verify an authorized real workflow: carrier routing and audio for phone agents, real tool authentication and business contracts, signed webhook handling, trigger input and output readiness. Then inspect a production run's publication and external-action evidence.
Use run review to confirm which publication executed and what the business service returned.









