On this page
Connect the Cumulus MCP
The Cumulus MCP server gives an AI assistant the Talos operations as tools: it can build and publish agents, start runs, follow them and read their transcripts, with the same API key and permissions as your code.
Documentation for coding agents
Connect the public documentation server at https://api.cumuluslabs.io/mcp/docs without a key. The authenticated customer server below offers the same docs tools alongside your permitted operations. See Build with an AI agent for Codex and Claude setup.
| Tool | Use |
|---|---|
docs_search | Find sections by question, operation, SDK method, MCP tool or error. |
docs_get | Read a page or section with a canonical URL, content version and pagination. |
docs_get_operation | Read registry-derived operation schemas and workflow guidance. |
Clients supporting MCP resources can browse the corpus; clients supporting prompts can select build-and-test, connection and debugging workflows. Retrieval never accesses customer data.
Quickstart
- Pick the environment. Build and test in your sandbox; repeat the setup in production when it works. Agents, trunks, numbers, secrets and tests are separate in each, and a key works only in the environment it was created in.
- Create a key. In Talos, switch to that environment, open Settings → API keys, name the key after the client (for example “Claude Code, Shivam”) and give it the member role. A member key builds, tests, publishes and places calls; keep admin work (trunks, numbers, secrets, keys, team) in Talos.
- Connect your client with the key: Claude Code or Claude Desktop below.
- Check it. Ask: “Which Cumulus environment am I connected to, and what agents do I have?”
- Give it the job. For example:
Prompt
Set up our Fairway outbound collections agent for SIP calling through our Twilio trunk, and test it.
Our Retell export is ./fairway-retell-export.json. The trunk and our caller IDs are already connected in Talos.
Test on a version that isn't live, show me the results, and ask me before you make anything live or place a call.
When your client connects, the server gives your assistant a setup guide: the order below, where each input comes from, and which steps need you or Cumulus. Each tool's description says what to call it with and what to call next.
| Step | Who | Notes |
|---|---|---|
| Voice capacity and run admission for the environment | Cumulus staff | Once per environment (sandbox and production) |
| Elastic SIP trunk: Termination URI, Credential List, Secure Trunking, your numbers | You, in the Twilio console | See Phone |
| Connect the trunk (the SIP password becomes a sip_auth secret) and import each caller ID | An admin, in Talos (Phone numbers) | Or secrets_put, trunks_create, numbers_import with an admin key |
| Import the Retell export, or create an agent from a template | Your assistant: agents_import_inspect, agents_import | A draft; nothing goes live |
| Set the trunk and caller ID; fix what validation reports | Your assistant: drafts_edit, drafts_validate | Fixes are ready-made edits |
| Publish a test version and run simulated calls against it | Your assistant: publications_publish (live: false), test_suites_create, test_runs_create | No phone call is placed |
| Make it live | Your assistant, after you agree: publications_promote | |
| Place calls | Your dialer or assistant: runs_create | Dials a real number |
Calls are real
runs_createon a phone agent dials the number you give it, also withtest: true(a test run is only left out of billing, webhooks and triggers). The server marks it as destructive, like every tool that changes access, credentials or what live calls do, so a client that confirms destructive tools asks you first. Keep that on.
What Cumulus staff do
Cumulus sets each environment's concurrent call cap and opens it for runs. Until then, connecting a trunk or importing a number answers
tenant_not_configured, a call answerstenant_not_configuredortenant_disabled, and a test runtenant_not_configuredortenant_admission_closed;GET /api/v1/runs/capacity(runs_capacity) shows the limit and whether admission is open. Ask your Cumulus contact to “set voice capacity and open run admission” for the environment, quoting its slug andtenant_idfromenvironments_list. You never need a Cumulus trunk: calls go out on yours.
The server
- URL:
https://api.cumuluslabs.io/mcp - Transport: Streamable HTTP, stateless, JSON responses (
POSTonly). No local process to install. - Authentication: an API key as a bearer token:
Authorization: Bearer cmls_…. There is no OAuth sign-in; create the key in Talos (Get an API key). A request without a key gets401. - Permissions: the key's role decides which tools the client sees. A viewer key lists only read tools; a member key adds runs and authoring; an admin key adds keys, secrets, trunks, numbers and the budget.
- Session cookies are not accepted: an MCP client always sends an API key.
- Untrusted content: results that carry text written by callers or other parties (transcripts, events, test trials, imported prompts, external responses) are marked
untrusted_content: true, and the text your assistant reads starts with a notice to treat it as data, never as instructions.
Treat the key as a password
The assistant acts with the key's full role. Give it its own member key (name it after the client), a viewer key when it only needs to read, and revoke it in Talos when you stop using it. If you do let an assistant connect a trunk with an admin key, use a separate key for that and revoke it afterwards.
Set up your client
Claude Code
With the key in CUMULUS_API_KEY, run:
Terminal
claude mcp add --transport http cumulus https://api.cumuluslabs.io/mcp \
--header "Authorization: Bearer $CUMULUS_API_KEY"
Add --scope user to use it in every project. Check it with claude mcp list, or /mcp inside a session.
Claude Desktop
Claude Desktop's custom connectors sign in with OAuth, and the Cumulus MCP takes an API key, so connect through the open-source mcp-remote bridge (needs Node.js). Add this to claude_desktop_config.json (Settings → Developer → Edit Config) and restart Claude Desktop:
claude_desktop_config.json
{
"mcpServers": {
"cumulus": {
"command": "npx",
"args": [
"-y", "mcp-remote", "https://api.cumuluslabs.io/mcp",
"--header", "Authorization:${CUMULUS_AUTH}"
],
"env": { "CUMULUS_AUTH": "Bearer cmls_..." }
}
}
}
Keep Authorization: and the variable together without a space: some clients split arguments on spaces.
Cursor
Add the server to ~/.cursor/mcp.json (every project) or .cursor/mcp.json (one project). Cursor reads the key from your environment:
~/.cursor/mcp.json
{
"mcpServers": {
"cumulus": {
"url": "https://api.cumuluslabs.io/mcp",
"headers": { "Authorization": "Bearer ${env:CUMULUS_API_KEY}" }
}
}
}
VS Code
Add the server to .vscode/mcp.json. VS Code asks for the key once and stores it securely:
.vscode/mcp.json
{
"inputs": [
{ "type": "promptString", "id": "cumulus-api-key", "description": "Cumulus API key", "password": true }
],
"servers": {
"cumulus": {
"type": "http",
"url": "https://api.cumuluslabs.io/mcp",
"headers": { "Authorization": "Bearer ${input:cumulus-api-key}" }
}
}
}
Any other MCP client that speaks Streamable HTTP and can send a header works the same way.
How tools work
Each tool is one API operation, named after it with dots as underscores (runs.create is runs_create). Its arguments mirror the request: params for the path, query and body. Tools that create something also take an optional idempotency_key; the server generates one when the client sends none.
tools/call
{
"name": "runs_create",
"arguments": {
"body": { "channel": "cloud", "agent_id": "…", "mode": "task", "input": { "ticket": "I was charged twice." }, "test": true },
"idempotency_key": "<a new UUID for each request>"
}
}
The result is { ok, status, result | error, idempotency_key }, with the same error codes as the API. When a command returns outcome_unknown, the assistant retries with the same idempotency_key, so nothing runs twice.
Result
{ "ok": true, "status": 201, "result": { "run_id": "…", "state": "admitted", … }, "idempotency_key": "<the key used>" }
First prompts to try
- “Import the Retell export in ./agent.json as a draft, show me what the import changed or dropped, and fix what validation reports.”
- “List my Talos agents and tell me which ones have unpublished draft changes.”
- “Start a test cloud run of the Ticket triage agent with the ticket "I was charged twice", wait for it to end and summarize its transcript.”
- “Show my last 20 phone runs that reached voicemail, grouped by agent.”
- “How many phone calls can I run at once right now, and how many are active?”
- “Read the analysis of run <run_id> and explain its disposition.”
- “Publish the open draft of agent <agent_id> as its next version, without making it live.”
Ask for test runs (test: true) while you experiment: they run normally but never reach your webhooks or triggers and are not billed. A phone test run still places a real call; simulated calls are test suites.
Tools
| Operation | MCP tool | Key role | Purpose |
|---|---|---|---|
GET /api/v1/agents | agents_list | viewer | List agents with their head publication, open draft and last run; search, filter and order server-side |
GET /api/v1/agents/{agent_id} | agents_get | viewer | Get one agent |
POST /api/v1/agents | agents_create | member | Create an agent from a template (201; an idempotent replay returns 200) |
DELETE /api/v1/agents/{agent_id} | agents_delete | member | Delete (archive) an agent: it leaves the catalog and takes no new drafts, publications or runs; history is kept |
POST /api/v1/agents/{agent_id}/clone | agents_clone | member | Clone an agent's head, chosen version or never-published draft into a new agent's first draft (201; replay 200) |
POST /api/v1/agents/import | agents_import | member | Import a Retell agent export as a draft, never live: a new agent gets it as its first draft; an existing agent's draft is replaced and its head is unchanged. Replacing a draft with unpublished changes needs replace_draft {draft_id, expected_revision} naming it (409 reason draft_has_changes otherwise). Publish the draft with publications.publish (201; re-importing the same export onto its unedited draft returns 200) |
POST /api/v1/agents/import/inspect | agents_import_inspect | member | Compile a Retell export and say what agents.import would do (new agent, new or replaced draft, unpublished changes it would discard) without importing it |
GET /api/v1/agents/{agent_id}/draft | drafts_get | viewer | Read the open draft |
POST /api/v1/agents/{agent_id}/draft | drafts_open | member | Open a draft from the head or a chosen version |
PATCH /api/v1/agents/{agent_id}/draft | drafts_edit | member | Apply domain edits to the draft (revision compare-and-set) |
DELETE /api/v1/agents/{agent_id}/draft | drafts_discard | member | Discard the draft (revision compare-and-set) |
POST /api/v1/agents/{agent_id}/draft/rebase | drafts_rebase | member | Rebase the draft onto the live version after a head change (keeps its definition; 409 head_moved when live moved again) |
POST /api/v1/agents/{agent_id}/draft/validate | drafts_validate | member | Compile the draft and return diagnostics and the would-be digest |
POST /api/v1/agents/{agent_id}/publications | publications_publish | member | Publish the draft as the next immutable version; moves head unless live is false (201; replay 200) |
GET /api/v1/agents/{agent_id}/publications | publications_list | viewer | List publications of an agent |
GET /api/v1/agents/{agent_id}/publications/{version} | publications_get | viewer | Get one publication with its compiled plan |
POST /api/v1/agents/{agent_id}/head | publications_promote | member | Move head to an existing version (manual promote or rollback), rebasing the named open draft in the same transaction |
GET /api/v1/actions | actions_list | viewer | List the action descriptors available to publications |
POST /api/v1/runs | runs_create | member | Admit a cloud or phone run of a publication (201; an idempotent replay returns 200) |
GET /api/v1/runs | runs_list | viewer | List runs |
GET /api/v1/runs/capacity | runs_capacity | viewer | Read admission capacity per channel: configured limit and runs holding it |
GET /api/v1/runs/{run_id} | runs_get | viewer | Get one run |
PATCH /api/v1/runs/{run_id} | runs_update | member | Update a live phone run's dynamic variables or metadata; returns a durable receipt (202 while pending, 200 once settled) |
GET /api/v1/runs/{run_id}/updates/{update_id} | runs_update_get | viewer | Read a live update's receipt |
POST /api/v1/runs/{run_id}/cancel | runs_cancel | member | Cancel an active run (accepted: a live call settles cancelled when its media session ends) |
POST /api/v1/runs/{run_id}/messages | runs_send_message | member | Send a user message to a cloud chat run |
GET /api/v1/runs/{run_id}/events | runs_events | viewer | Read the run's semantic events (JSON page or text/event-stream) |
GET /api/v1/runs/{run_id}/effects | runs_effects | viewer | Read the run's governed effects and receipts |
GET /api/v1/runs/{run_id}/transcript | runs_transcript | viewer | Read the run transcript |
GET /api/v1/runs/{run_id}/result | runs_result | viewer | Read what a cloud task run produced: its output object, validated against the publication's output schema |
GET /api/v1/runs/{run_id}/recording | runs_recording | viewer | Get a short-lived recording URL |
GET /api/v1/runs/{run_id}/analysis | runs_analysis | viewer | Read post-run analysis |
GET /api/v1/runs/{run_id}/usage | runs_usage | viewer | Read metered usage and cost of one run |
POST /api/v1/test-sessions | test_sessions_create | member | Start a browser voice test of one agent version: admits a web run and returns a short-lived LiveKit room token (the same Idempotency-Key rejoins the live run with a fresh token) |
GET /api/v1/usage | usage_summary | viewer | Summarize metered usage by agent, run or day: the billable runs, or (purpose=test) the test runs, metered and not billed |
GET /api/v1/usage/budget | budgets_get | viewer | Read the tenant budget: period, cap and metered spend (configured: false until metered pricing is installed) |
PUT /api/v1/usage/budget | budgets_update | admin | Set the tenant budget cap (compare-and-set on the active policy; refused until metered pricing is installed) |
GET /api/v1/usage/rate-card | usage_rate_card | viewer | Read the customer rates of the tenant's active pricing (configured: false until metered pricing is installed) |
GET /api/v1/voices | voices_list | viewer | List the speech models and platform TTS voices an agent can use |
GET /api/v1/telephony/trunks | trunks_list | viewer | List BYOC SIP trunks |
POST /api/v1/telephony/trunks | trunks_create | admin | Connect a BYOC SIP trunk |
PATCH /api/v1/telephony/trunks/{trunk_id} | trunks_update | admin | Update a BYOC SIP trunk |
DELETE /api/v1/telephony/trunks/{trunk_id} | trunks_delete | admin | Disconnect a BYOC SIP trunk |
GET /api/v1/telephony/numbers | numbers_list | viewer | List BYOC phone numbers |
POST /api/v1/telephony/numbers | numbers_import | admin | Import a phone number from a connected trunk |
PATCH /api/v1/telephony/numbers/{number_id} | numbers_update | admin | Assign agents or rename a phone number |
DELETE /api/v1/telephony/numbers/{number_id} | numbers_release | admin | Release a phone number |
GET /api/v1/secrets | secrets_list | viewer | List secret metadata (never values) |
PUT /api/v1/secrets/{secret_id} | secrets_put | admin | Create a secret or add a new version |
DELETE /api/v1/secrets/{secret_id} | secrets_delete | admin | Delete a secret and all its versions |
GET /api/v1/test-suites | test_suites_list | viewer | List test suites (optionally one agent's) with their last run |
GET /api/v1/test-suites/{suite_id} | test_suites_get | viewer | Get a test suite (its current revision, or ?revision) with its last run and flaky cases |
POST /api/v1/test-suites | test_suites_create | member | Create a test suite on an agent (201; an idempotent replay returns 200) |
PATCH /api/v1/test-suites/{suite_id} | test_suites_edit | member | Edit a test suite (rename, upsert, remove or reorder cases) as a compare-and-set on its revision |
DELETE /api/v1/test-suites/{suite_id} | test_suites_delete | member | Archive a test suite; its runs and results are kept |
POST /api/v1/test-suites/{suite_id}/validate | test_suites_validate | member | Check a suite's node, action and variable references against a published version or the open draft |
POST /api/v1/test-runs/estimate | test_runs_estimate | viewer | Estimate a test run's trials, cost and duration before starting it |
POST /api/v1/test-runs | test_runs_create | member | Run test suites against one publication version (201; an idempotent replay returns 200) |
GET /api/v1/test-runs | test_runs_list | viewer | List test runs, newest first |
GET /api/v1/test-runs/{test_run_id} | test_runs_get | viewer | Get a test run: state, verdict, progress, summary and cost |
GET /api/v1/test-runs/{test_run_id}/results | test_runs_results | viewer | List a test run's case results with each trial's assertion verdicts |
GET /api/v1/test-runs/{test_run_id}/cases/{case_id}/trials/{trial} | test_runs_trial | viewer | Get one trial: its assertion trace and, while retained, the transcript it was judged on |
GET /api/v1/test-runs/{test_run_id}/events | test_runs_events | viewer | Read a test run's journal events (JSON page or text/event-stream) |
POST /api/v1/test-runs/{test_run_id}/cancel | test_runs_cancel | member | Cancel a test run: no new trials start and running trials are cancelled; settled results stay |
GET /api/v1/keys | keys_list | admin | List API keys (never secrets) |
POST /api/v1/keys | keys_create | admin | Create an API key with an explicit role |
DELETE /api/v1/keys/{key_id} | keys_revoke | admin | Revoke an API key; queued effects of its runs are denied |
POST /api/v1/keys/{key_id}/rotate | keys_rotate | admin | Replace an API key's secret |
GET /api/v1/environments | environments_list | viewer | List this organization's production and sandbox environments |
POST /api/v1/environments/sandbox | environments_ensure_sandbox | admin | Create this organization's sandbox environment on first use (same members) |