Realtime
WebSocket endpoints. Two are for telephony providers to stream call audio into; two are for your own clients to watch a call as it happens.
| Endpoint | Used by | Auth |
|---|---|---|
/api/voice/ws/plivo | Plivo | None — the provider connects |
/api/voice/ws/exotel | Exotel | None |
/api/voice/ws/browser/:flow_uuid | Your browser test console | ?token= |
/api/voice/ws/live/:call_uuid | Anything watching a live call | ?token= |
/api/voice/ws/agent-assist/:call_uuid | Live coaching overlay | ?token= |
Browser clients cannot set an Authorization header, so these take the session
token as a query parameter.
Watch a live call
wss://your-host/api/voice/ws/live/<call_uuid>?token=<session-token>
Receive-only. Each message is one JSON event:
{ "type": "transcript", "speaker": "customer", "text": "Yes, speaking.", "is_final": true, "ts": "2026-09-20T14:32:07Z" }
{ "type": "status", "status": "connected" }
{ "type": "status", "status": "ended" }
| Field | Notes |
|---|---|
type | transcript or status |
speaker | agent or customer |
is_final | false for interim recognition that will be revised |
status | connected, ended |
Interim results arrive before their final version, so a UI that appends unconditionally will show the same sentence twice. Replace the last non-final turn instead.
| Status | Cause |
|---|---|
401 | Missing or invalid token |
403 | The call belongs to another organization |
Both AI calls and human click-to-calls stream in the same shape, so watching one is no different from watching the other.
Browser test calls
wss://your-host/api/voice/ws/browser/<flow_uuid>?token=<session-token>
Runs a full conversation against a flow's agent from a browser, with no telephony involved. Nothing is persisted, billed, or counted — these calls are marked ephemeral.
On connect:
{ "type": "ready", "call_uuid": "cal-…" }
Then send raw PCM audio as binary frames, and receive synthesised audio back as binary plus the same transcript events as above. Text frames from the client are ignored.
The flow UUID is resolved against your token's organization, so another
tenant's flow is a 404.
Agent assist
wss://your-host/api/voice/ws/agent-assist/<call_uuid>?token=<session-token>
Streams live transcription and suggestion events for a human-handled call, for
an overlay shown to the teammate on the phone. Registered only when live
transcription of human calls is enabled
(LIVE_TRANSCRIBE_HUMAN_CALLS).
Provider media sockets
/api/voice/ws/plivo and /api/voice/ws/exotel are what the provider's
<Stream> verb connects back to. You do not call these — the answer webhook
hands the provider the URL.
Worth knowing about them:
- The origin check is disabled on the upgrade, because a provider does not
send a browser
Originheader. Browser-facing sockets verify the token inside the handler instead. - Plivo's start frame arrives with empty
from/to, which is why the answer webhook puts those on the socket URL. Without them an inbound call has nothing to match a flow against. - Exotel takes the sample rate on the URL (
?sample-rate=16000) and defaults to 8000 without it. It must match the rate the transport is built with, or audio plays at the wrong speed.
Reconnecting
There is no replay. A client that reconnects mid-call sees events from that
moment on; fetch GET /api/voice/calls/{uuid}
afterwards for the complete transcript.
Live-stream events are not persisted separately — the stored transcript is written when the call ends.