Source: https://docs.cumuluslabs.io/testing
Content version: 2c4b20f9

# Test an agent

Test a pinned publication before making it live. Use text simulations or browser voice sessions before placing real phone calls.

## Prerequisites {#prerequisites}

Publish a [test version](/docs/recipes#publish-test-version), choose your sandbox, and ensure run admission is open.

## Testing methods {#testing}

A test suite holds cases: a simulated caller (a `persona` with a goal, or a `script` of lines) talks to the agent in text, and assertions check what happened. `POST /api/v1/test-runs` runs suites against one published version, live or test, so a draft is tested by publishing it with `"live": false` first. No phone call is placed, and no trunk or number is needed; the environment's run admission must be open (Cumulus staff open it; until then a test run is refused with `tenant_not_configured` or `tenant_admission_closed`).

**Create a suite and run it**

```
POST /api/v1/test-suites
{
  "agent_id": "fairway-rpc-outbound",
  "name": "Right-party contact",
  "spec": { "kind": "simulation", "cases": [{
    "case_id": "rpc-confirms", "name": "The debtor confirms who they are", "language": "en",
    "caller": { "kind": "persona", "goal": "Confirm you are Jane Doe and ask what the call is about",
                "brief": "You are Jane Doe. You answer briefly and politely." },
    "dynamic_variables": { "debtorName": "Jane Doe", "debtorFirstName": "Jane" },
    "assertions": [
      { "type": "node_visited", "id": "identity-checked", "node_ids": ["<a node id from the draft>"] },
      { "type": "agent_says", "id": "no-balance-before-id", "expect": "never", "match": { "patterns": ["balance", "you owe"] },
        "scope": { "until_node": "<the node that confirms identity>" } }
    ]
  }] }
}

POST /api/v1/test-runs
{ "agent_id": "fairway-rpc-outbound", "version": 0, "suite_ids": ["<suite_id>"] }
```

*   Deterministic assertions check the path and the words: `node_visited`, `node_not_visited`, `final_node`, `path`, `agent_says`, `tool_called`, `variable`, `ended`, `transfer`, `disposition`. A `rubric` assertion is judged by a model against your criteria.
*   `mocks` answer the agent's HTTP actions by `action_id`, so a test never reaches your systems.
*   Each case runs `trials` times (default 1, at most 10) and passes by `pass_policy` (`majority` or `all`). `POST /api/v1/test-runs/estimate` gives the trials, cost and minutes first.
*   Poll `GET /api/v1/test-runs/{test_run_id}` until `state` is `completed`; `…/results` lists every case and assertion verdict, and `…/cases/{case_id}/trials/{trial}` shows the transcript a trial was judged on.
*   `POST /api/v1/test-suites/{suite_id}/validate` checks a suite's node, action and variable references against a version or the open draft.

| Operation | MCP tool | Key role | Purpose |
| --- | --- | --- | --- |
| [`GET /api/v1/test-suites`](https://docs.cumuluslabs.io/reference/test_suites.list) | `test_suites_list` | viewer | List test suites (optionally one agent's) with their last run |
| [`GET /api/v1/test-suites/{suite_id}`](https://docs.cumuluslabs.io/reference/test_suites.get) | `test_suites_get` | viewer | Get a test suite (its current revision, or ?revision) with its last run and flaky cases |
| [`POST /api/v1/test-suites`](https://docs.cumuluslabs.io/reference/test_suites.create) | `test_suites_create` | member | Create a test suite on an agent (201; an idempotent replay returns 200) |
| [`PATCH /api/v1/test-suites/{suite_id}`](https://docs.cumuluslabs.io/reference/test_suites.edit) | `test_suites_edit` | member | Edit a test suite (rename, upsert, remove or reorder cases) as a compare-and-set on its revision |
| [`DELETE /api/v1/test-suites/{suite_id}`](https://docs.cumuluslabs.io/reference/test_suites.delete) | `test_suites_delete` | member | Archive a test suite; its runs and results are kept |
| [`POST /api/v1/test-suites/{suite_id}/validate`](https://docs.cumuluslabs.io/reference/test_suites.validate) | `test_suites_validate` | member | Check a suite's node, action and variable references against a published version or the open draft |
| [`POST /api/v1/test-runs/estimate`](https://docs.cumuluslabs.io/reference/test_runs.estimate) | `test_runs_estimate` | viewer | Estimate a test run's trials, cost and duration before starting it |
| [`POST /api/v1/test-runs`](https://docs.cumuluslabs.io/reference/test_runs.create) | `test_runs_create` | member | Run test suites against one publication version (201; an idempotent replay returns 200) |
| [`GET /api/v1/test-runs`](https://docs.cumuluslabs.io/reference/test_runs.list) | `test_runs_list` | viewer | List test runs, newest first |
| [`GET /api/v1/test-runs/{test_run_id}`](https://docs.cumuluslabs.io/reference/test_runs.get) | `test_runs_get` | viewer | Get a test run: state, verdict, progress, summary and cost |
| [`GET /api/v1/test-runs/{test_run_id}/results`](https://docs.cumuluslabs.io/reference/test_runs.results) | `test_runs_results` | viewer | List a test run's case results with each trial's assertion verdicts |
| [`GET /api/v1/test-runs/{test_run_id}/cases/{case_id}/trials/{trial}`](https://docs.cumuluslabs.io/reference/test_runs.trial) | `test_runs_trial` | viewer | Get one trial: its assertion trace and, while retained, the transcript it was judged on |
| [`GET /api/v1/test-runs/{test_run_id}/events`](https://docs.cumuluslabs.io/reference/test_runs.events) | `test_runs_events` | viewer | Read a test run's journal events (JSON page or text/event-stream) |
| [`POST /api/v1/test-runs/{test_run_id}/cancel`](https://docs.cumuluslabs.io/reference/test_runs.cancel) | `test_runs_cancel` | member | Cancel a test run: no new trials start and running trials are cancelled; settled results stay |


## Next step {#next-step}

Read the results, fix failures in a new draft and repeat. Promote the tested publication only when you are ready to go live. See [recipes](/docs/recipes).
