Cumulus TalosDocs
Open Talos
On this page

Release and operate a tested agent

Validate a draft, publish a test version, compare it with the baseline and promote it after the required checks pass. Studio keeps the version changes, test results and run records available for review.

The screenshots show synthetic publications, cases, verdicts and costs. Select a screenshot to inspect it at full size.

Keep the candidate separate from live behavior

Published versions showing live, test and previous states with the available actions

Edit the draft, validate it and publish a test version with live: false. Existing runs continue with their original publication. New production runs keep using the live version until you change it.

The version list tells you which publication is live, which candidate is a test and which publications were used previously. Review the candidate's source and definition before changing the live version.

A candidate publication expanded to show behavioral differences from the earlier version

Open the change summary and check the actual behavior change. A canvas layout change is different from a prompt, action, condition or provider change. The screenshot's candidate tightens the outcome statement; the review should ask whether it can still claim completion before the external action is confirmed.

Design the acceptance cases

A saved persona case asking about an account without immediately confirming identity

Node and judged assertions for verification before account disclosure

Choose the cases that establish the change is safe. For an order-resolution release, cover authorization refusal, missing account facts, negative eligibility, caller refusal, stale data, uncertain writes and unavailable transfers.

Use structural assertions for the path and actions, and judged criteria for conversational requirements. State the expected behavior, such as “verify identity before sharing invoice details,” in each case.

Mocks give repeatable responses without reaching your business endpoints. An unmocked customer HTTP action fails closed in Cumulus suite simulations. Use real integration acceptance separately to establish credentials, routing and the business transaction contract.

Review results on the exact publication

A test run summary with its pinned version, suite revision, result and comparison context

Check both the agent publication and the suite revision. If the suite changed, the results may no longer measure the same behavior. Read failed, invalid, errored and flaky outcomes separately; they call for different investigation.

For a case error, inspect the test setup and execution failure. For a failed criterion, read the evidence and judge rationale before deciding whether the agent needs a change.

Follow a failure to the deciding evidence

A failed trial showing the missing verification node and the judge's evidence turn

Open a trial and select the assertion. Use the highlighted turn, node path and action record to determine whether the fault is in the agent behavior, test setup, criterion or judge.

In this example, the candidate shared invoice details before verification. Fix the sequence, publish another test version and rerun the same case. Preserve the failed trial as evidence of what the earlier candidate did.

Compare the candidate with its baseline

Aligned baseline and candidate conversations with the point of path divergence

Use the same case across versions and compare the trial evidence, not only the summary score. The baseline in this illustration waited for verification; the candidate took the lookup path too early. The aligned view helps locate that behavioral difference.

Keep suite revisions, caller behavior, mocks, judge configuration and trial count consistent when the goal is a meaningful comparison. When you change them, record that change rather than calling the result an agent improvement without qualification.

Record a judge disagreement

The human feedback control for a judged assertion

If the judge is wrong, record the human verdict through the feedback control and retain the supporting explanation in your review process. This feedback is an evaluation of the trial; it does not rewrite the original test result.

Calibrate model judges against cases your team has reviewed. Use deterministic assertions for requirements you can check mechanically.

Set the information your operators will see

Configured post-run analysis fields for request type, resolution and the next step

Define analysis fields in business terms. An operator should be able to tell whether the case was resolved, what remains to be done and why human review is needed. Distinguish extraction from the actual transaction result: a model-produced resolved flag does not override a failed action receipt.

The per-agent data-storage setting

Choose the per-agent data-storage behavior and review workspace retention. If transcripts or recordings are not retained, an operator may not have the evidence a conversational criterion needs later. Read workspace retention before deciding how long to keep review material.

Promote, then verify the real path

Use publication promotion for the tested version. Preserve the version and suite IDs in your release review. Where a suite is configured as required, review the applicable publication checks and their results.

After promotion, verify an authorized real workflow: carrier routing and audio for phone agents, real tool authentication and business contracts, signed webhook handling, trigger input and output readiness. Then inspect a production run's publication and external-action evidence.

Use run review to confirm which publication executed and what the business service returned.

Content version 459f89d1Markdown source
Release and operate a tested agent · Cumulus Talos docs