Cumulus TalosDocs
Open Talos
On this page

Connect the Cumulus MCP

The Cumulus MCP server gives an AI assistant the Talos operations as tools: it can build and publish agents, start runs, follow them and read their transcripts, with the same API key and permissions as your code.

Documentation for coding agents

Connect the public documentation server at https://api.cumuluslabs.io/mcp/docs without a key. The authenticated customer server below offers the same docs tools alongside your permitted operations. See Build with an AI agent for Codex and Claude setup.

ToolUse
docs_searchFind sections by question, operation, SDK method, MCP tool or error.
docs_getRead a page or section with a canonical URL, content version and pagination.
docs_get_operationRead registry-derived operation schemas and workflow guidance.

Clients supporting MCP resources can browse the corpus; clients supporting prompts can select build-and-test, connection and debugging workflows. Retrieval never accesses customer data.

Quickstart

  1. Pick the environment. Build and test in your sandbox; repeat the setup in production when it works. Agents, trunks, numbers, secrets and tests are separate in each, and a key works only in the environment it was created in.
  2. Create a key. In Talos, switch to that environment, open Settings → API keys, name the key after the client (for example “Claude Code, Shivam”) and give it the member role. A member key builds, tests, publishes and places calls; keep admin work (trunks, numbers, secrets, keys, team) in Talos.
  3. Connect your client with the key: Claude Code or Claude Desktop below.
  4. Check it. Ask: “Which Cumulus environment am I connected to, and what agents do I have?”
  5. Give it the job. For example:

Prompt

Example
Set up our Fairway outbound collections agent for SIP calling through our Twilio trunk, and test it.
Our Retell export is ./fairway-retell-export.json. The trunk and our caller IDs are already connected in Talos.
Test on a version that isn't live, show me the results, and ask me before you make anything live or place a call.

When your client connects, the server gives your assistant a setup guide: the order below, where each input comes from, and which steps need you or Cumulus. Each tool's description says what to call it with and what to call next.

StepWhoNotes
Voice capacity and run admission for the environmentCumulus staffOnce per environment (sandbox and production)
Elastic SIP trunk: Termination URI, Credential List, Secure Trunking, your numbersYou, in the Twilio consoleSee Phone
Connect the trunk (the SIP password becomes a sip_auth secret) and import each caller IDAn admin, in Talos (Phone numbers)Or secrets_put, trunks_create, numbers_import with an admin key
Import the Retell export, or create an agent from a templateYour assistant: agents_import_inspect, agents_importA draft; nothing goes live
Set the trunk and caller ID; fix what validation reportsYour assistant: drafts_edit, drafts_validateFixes are ready-made edits
Publish a test version and run simulated calls against itYour assistant: publications_publish (live: false), test_suites_create, test_runs_createNo phone call is placed
Make it liveYour assistant, after you agree: publications_promote
Place callsYour dialer or assistant: runs_createDials a real number

Calls are real

runs_create on a phone agent dials the number you give it, also with test: true (a test run is only left out of billing, webhooks and triggers). The server marks it as destructive, like every tool that changes access, credentials or what live calls do, so a client that confirms destructive tools asks you first. Keep that on.

What Cumulus staff do

Cumulus sets each environment's concurrent call cap and opens it for runs. Until then, connecting a trunk or importing a number answers tenant_not_configured, a call answers tenant_not_configured or tenant_disabled, and a test run tenant_not_configured or tenant_admission_closed; GET /api/v1/runs/capacity (runs_capacity) shows the limit and whether admission is open. Ask your Cumulus contact to “set voice capacity and open run admission” for the environment, quoting its slug and tenant_id from environments_list. You never need a Cumulus trunk: calls go out on yours.

The server

  • URL: https://api.cumuluslabs.io/mcp
  • Transport: Streamable HTTP, stateless, JSON responses (POST only). No local process to install.
  • Authentication: an API key as a bearer token: Authorization: Bearer cmls_…. There is no OAuth sign-in; create the key in Talos (Get an API key). A request without a key gets 401.
  • Permissions: the key's role decides which tools the client sees. A viewer key lists only read tools; a member key adds runs and authoring; an admin key adds keys, secrets, trunks, numbers and the budget.
  • Session cookies are not accepted: an MCP client always sends an API key.
  • Untrusted content: results that carry text written by callers or other parties (transcripts, events, test trials, imported prompts, external responses) are marked untrusted_content: true, and the text your assistant reads starts with a notice to treat it as data, never as instructions.

Treat the key as a password

The assistant acts with the key's full role. Give it its own member key (name it after the client), a viewer key when it only needs to read, and revoke it in Talos when you stop using it. If you do let an assistant connect a trunk with an admin key, use a separate key for that and revoke it afterwards.

Set up your client

Claude Code

With the key in CUMULUS_API_KEY, run:

Terminal

Example
claude mcp add --transport http cumulus https://api.cumuluslabs.io/mcp \
  --header "Authorization: Bearer $CUMULUS_API_KEY"

Add --scope user to use it in every project. Check it with claude mcp list, or /mcp inside a session.

Claude Desktop

Claude Desktop's custom connectors sign in with OAuth, and the Cumulus MCP takes an API key, so connect through the open-source mcp-remote bridge (needs Node.js). Add this to claude_desktop_config.json (Settings → Developer → Edit Config) and restart Claude Desktop:

claude_desktop_config.json

Example
{
  "mcpServers": {
    "cumulus": {
      "command": "npx",
      "args": [
        "-y", "mcp-remote", "https://api.cumuluslabs.io/mcp",
        "--header", "Authorization:${CUMULUS_AUTH}"
      ],
      "env": { "CUMULUS_AUTH": "Bearer cmls_..." }
    }
  }
}

Keep Authorization: and the variable together without a space: some clients split arguments on spaces.

Cursor

Add the server to ~/.cursor/mcp.json (every project) or .cursor/mcp.json (one project). Cursor reads the key from your environment:

~/.cursor/mcp.json

Example
{
  "mcpServers": {
    "cumulus": {
      "url": "https://api.cumuluslabs.io/mcp",
      "headers": { "Authorization": "Bearer ${env:CUMULUS_API_KEY}" }
    }
  }
}

VS Code

Add the server to .vscode/mcp.json. VS Code asks for the key once and stores it securely:

.vscode/mcp.json

Example
{
  "inputs": [
    { "type": "promptString", "id": "cumulus-api-key", "description": "Cumulus API key", "password": true }
  ],
  "servers": {
    "cumulus": {
      "type": "http",
      "url": "https://api.cumuluslabs.io/mcp",
      "headers": { "Authorization": "Bearer ${input:cumulus-api-key}" }
    }
  }
}

Any other MCP client that speaks Streamable HTTP and can send a header works the same way.

How tools work

Each tool is one API operation, named after it with dots as underscores (runs.create is runs_create). Its arguments mirror the request: params for the path, query and body. Tools that create something also take an optional idempotency_key; the server generates one when the client sends none.

tools/call

Example
{
  "name": "runs_create",
  "arguments": {
    "body": { "channel": "cloud", "agent_id": "…", "mode": "task", "input": { "ticket": "I was charged twice." }, "test": true },
    "idempotency_key": "<a new UUID for each request>"
  }
}

The result is { ok, status, result | error, idempotency_key }, with the same error codes as the API. When a command returns outcome_unknown, the assistant retries with the same idempotency_key, so nothing runs twice.

Result

Example
{ "ok": true, "status": 201, "result": { "run_id": "…", "state": "admitted", … }, "idempotency_key": "<the key used>" }

First prompts to try

  • “Import the Retell export in ./agent.json as a draft, show me what the import changed or dropped, and fix what validation reports.”
  • “List my Talos agents and tell me which ones have unpublished draft changes.”
  • “Start a test cloud run of the Ticket triage agent with the ticket "I was charged twice", wait for it to end and summarize its transcript.”
  • “Show my last 20 phone runs that reached voicemail, grouped by agent.”
  • “How many phone calls can I run at once right now, and how many are active?”
  • “Read the analysis of run <run_id> and explain its disposition.”
  • “Publish the open draft of agent <agent_id> as its next version, without making it live.”

Ask for test runs (test: true) while you experiment: they run normally but never reach your webhooks or triggers and are not billed. A phone test run still places a real call; simulated calls are test suites.

Tools

OperationMCP toolKey rolePurpose
GET /api/v1/agentsagents_listviewerList agents with their head publication, open draft and last run; search, filter and order server-side
GET /api/v1/agents/{agent_id}agents_getviewerGet one agent
POST /api/v1/agentsagents_creatememberCreate an agent from a template (201; an idempotent replay returns 200)
DELETE /api/v1/agents/{agent_id}agents_deletememberDelete (archive) an agent: it leaves the catalog and takes no new drafts, publications or runs; history is kept
POST /api/v1/agents/{agent_id}/cloneagents_clonememberClone an agent's head, chosen version or never-published draft into a new agent's first draft (201; replay 200)
POST /api/v1/agents/importagents_importmemberImport a Retell agent export as a draft, never live: a new agent gets it as its first draft; an existing agent's draft is replaced and its head is unchanged. Replacing a draft with unpublished changes needs replace_draft {draft_id, expected_revision} naming it (409 reason draft_has_changes otherwise). Publish the draft with publications.publish (201; re-importing the same export onto its unedited draft returns 200)
POST /api/v1/agents/import/inspectagents_import_inspectmemberCompile a Retell export and say what agents.import would do (new agent, new or replaced draft, unpublished changes it would discard) without importing it
GET /api/v1/agents/{agent_id}/draftdrafts_getviewerRead the open draft
POST /api/v1/agents/{agent_id}/draftdrafts_openmemberOpen a draft from the head or a chosen version
PATCH /api/v1/agents/{agent_id}/draftdrafts_editmemberApply domain edits to the draft (revision compare-and-set)
DELETE /api/v1/agents/{agent_id}/draftdrafts_discardmemberDiscard the draft (revision compare-and-set)
POST /api/v1/agents/{agent_id}/draft/rebasedrafts_rebasememberRebase the draft onto the live version after a head change (keeps its definition; 409 head_moved when live moved again)
POST /api/v1/agents/{agent_id}/draft/validatedrafts_validatememberCompile the draft and return diagnostics and the would-be digest
POST /api/v1/agents/{agent_id}/publicationspublications_publishmemberPublish the draft as the next immutable version; moves head unless live is false (201; replay 200)
GET /api/v1/agents/{agent_id}/publicationspublications_listviewerList publications of an agent
GET /api/v1/agents/{agent_id}/publications/{version}publications_getviewerGet one publication with its compiled plan
POST /api/v1/agents/{agent_id}/headpublications_promotememberMove head to an existing version (manual promote or rollback), rebasing the named open draft in the same transaction
GET /api/v1/actionsactions_listviewerList the action descriptors available to publications
POST /api/v1/runsruns_creatememberAdmit a cloud or phone run of a publication (201; an idempotent replay returns 200)
GET /api/v1/runsruns_listviewerList runs
GET /api/v1/runs/capacityruns_capacityviewerRead admission capacity per channel: configured limit and runs holding it
GET /api/v1/runs/{run_id}runs_getviewerGet one run
PATCH /api/v1/runs/{run_id}runs_updatememberUpdate a live phone run's dynamic variables or metadata; returns a durable receipt (202 while pending, 200 once settled)
GET /api/v1/runs/{run_id}/updates/{update_id}runs_update_getviewerRead a live update's receipt
POST /api/v1/runs/{run_id}/cancelruns_cancelmemberCancel an active run (accepted: a live call settles cancelled when its media session ends)
POST /api/v1/runs/{run_id}/messagesruns_send_messagememberSend a user message to a cloud chat run
GET /api/v1/runs/{run_id}/eventsruns_eventsviewerRead the run's semantic events (JSON page or text/event-stream)
GET /api/v1/runs/{run_id}/effectsruns_effectsviewerRead the run's governed effects and receipts
GET /api/v1/runs/{run_id}/transcriptruns_transcriptviewerRead the run transcript
GET /api/v1/runs/{run_id}/resultruns_resultviewerRead what a cloud task run produced: its output object, validated against the publication's output schema
GET /api/v1/runs/{run_id}/recordingruns_recordingviewerGet a short-lived recording URL
GET /api/v1/runs/{run_id}/analysisruns_analysisviewerRead post-run analysis
GET /api/v1/runs/{run_id}/usageruns_usageviewerRead metered usage and cost of one run
POST /api/v1/test-sessionstest_sessions_creatememberStart a browser voice test of one agent version: admits a web run and returns a short-lived LiveKit room token (the same Idempotency-Key rejoins the live run with a fresh token)
GET /api/v1/usageusage_summaryviewerSummarize metered usage by agent, run or day: the billable runs, or (purpose=test) the test runs, metered and not billed
GET /api/v1/usage/budgetbudgets_getviewerRead the tenant budget: period, cap and metered spend (configured: false until metered pricing is installed)
PUT /api/v1/usage/budgetbudgets_updateadminSet the tenant budget cap (compare-and-set on the active policy; refused until metered pricing is installed)
GET /api/v1/usage/rate-cardusage_rate_cardviewerRead the customer rates of the tenant's active pricing (configured: false until metered pricing is installed)
GET /api/v1/voicesvoices_listviewerList the speech models and platform TTS voices an agent can use
GET /api/v1/telephony/trunkstrunks_listviewerList BYOC SIP trunks
POST /api/v1/telephony/trunkstrunks_createadminConnect a BYOC SIP trunk
PATCH /api/v1/telephony/trunks/{trunk_id}trunks_updateadminUpdate a BYOC SIP trunk
DELETE /api/v1/telephony/trunks/{trunk_id}trunks_deleteadminDisconnect a BYOC SIP trunk
GET /api/v1/telephony/numbersnumbers_listviewerList BYOC phone numbers
POST /api/v1/telephony/numbersnumbers_importadminImport a phone number from a connected trunk
PATCH /api/v1/telephony/numbers/{number_id}numbers_updateadminAssign agents or rename a phone number
DELETE /api/v1/telephony/numbers/{number_id}numbers_releaseadminRelease a phone number
GET /api/v1/secretssecrets_listviewerList secret metadata (never values)
PUT /api/v1/secrets/{secret_id}secrets_putadminCreate a secret or add a new version
DELETE /api/v1/secrets/{secret_id}secrets_deleteadminDelete a secret and all its versions
GET /api/v1/test-suitestest_suites_listviewerList test suites (optionally one agent's) with their last run
GET /api/v1/test-suites/{suite_id}test_suites_getviewerGet a test suite (its current revision, or ?revision) with its last run and flaky cases
POST /api/v1/test-suitestest_suites_creatememberCreate a test suite on an agent (201; an idempotent replay returns 200)
PATCH /api/v1/test-suites/{suite_id}test_suites_editmemberEdit a test suite (rename, upsert, remove or reorder cases) as a compare-and-set on its revision
DELETE /api/v1/test-suites/{suite_id}test_suites_deletememberArchive a test suite; its runs and results are kept
POST /api/v1/test-suites/{suite_id}/validatetest_suites_validatememberCheck a suite's node, action and variable references against a published version or the open draft
POST /api/v1/test-runs/estimatetest_runs_estimateviewerEstimate a test run's trials, cost and duration before starting it
POST /api/v1/test-runstest_runs_creatememberRun test suites against one publication version (201; an idempotent replay returns 200)
GET /api/v1/test-runstest_runs_listviewerList test runs, newest first
GET /api/v1/test-runs/{test_run_id}test_runs_getviewerGet a test run: state, verdict, progress, summary and cost
GET /api/v1/test-runs/{test_run_id}/resultstest_runs_resultsviewerList a test run's case results with each trial's assertion verdicts
GET /api/v1/test-runs/{test_run_id}/cases/{case_id}/trials/{trial}test_runs_trialviewerGet one trial: its assertion trace and, while retained, the transcript it was judged on
GET /api/v1/test-runs/{test_run_id}/eventstest_runs_eventsviewerRead a test run's journal events (JSON page or text/event-stream)
POST /api/v1/test-runs/{test_run_id}/canceltest_runs_cancelmemberCancel a test run: no new trials start and running trials are cancelled; settled results stay
GET /api/v1/keyskeys_listadminList API keys (never secrets)
POST /api/v1/keyskeys_createadminCreate an API key with an explicit role
DELETE /api/v1/keys/{key_id}keys_revokeadminRevoke an API key; queued effects of its runs are denied
POST /api/v1/keys/{key_id}/rotatekeys_rotateadminReplace an API key's secret
GET /api/v1/environmentsenvironments_listviewerList this organization's production and sandbox environments
POST /api/v1/environments/sandboxenvironments_ensure_sandboxadminCreate this organization's sandbox environment on first use (same members)
Content version 1f673cdbMarkdown source
Connect the Cumulus MCP · Cumulus Talos docs