Overview
nextneural_converse is an autonomous voice agent. It answers inbound calls and places
outbound campaign calls, holding the conversation itself: speech recognition,
an LLM, and speech synthesis run in a loop over the telephony provider's media
stream.
How the pieces fit
Voice Agent ──► Flow ──► Customers ──► Calls
│ │ │
│ ├─ inbound: answers a DID └─► Webhooks, transcripts
│ └─ outbound: dials a list
│
└─ persona, voice, language Workflow ──► handoff to a human
- Agents define the persona — name, voice, language, tone
- Flows bind an agent to either an inbound number or an outbound contact list
- Campaigns are the run control for an outbound flow
- Customers are the people being called
- Calls are the result, with transcript and recording
- Workflows decide when to hand off to a person
Base paths
| Surface | Base | Credential |
|---|---|---|
| Console API | /api | Session token |
| Developer API | /v1 | API key (public API) |
| Auth | /api/auth | — |
| Telephony webhooks | /api/voice/webhooks | None — called by the provider |
| Media and live streams | /api/voice/ws | Token in query string |
Console endpoints below are written in full, e.g. GET /api/voice/agents.
A call, end to end
Inbound. The provider fetches the answer webhook; the dialled number is matched against active inbound flows. No match means a hangup rather than answering in the wrong company's voice. On a match, the provider is handed XML pointing at the media socket, and the pipeline takes over.
Outbound. The campaign runner walks the contact list, skipping anyone
already reached and anyone outside the calling window, and places calls through
Plivo at calls_per_minute. Each call row is written before dialling, so a
call that fails at the provider still leaves a record.
During the call. The caller's audio is transcribed; the reply is generated and synthesised sentence by sentence so the first is playing while the rest is still being written. Barge-in cancels generation and flushes buffered audio.
Ending. Transcript, duration, and recording URL are written when the media socket closes. The hangup webhook closes out calls that never connected at all — a no-answer or busy produces no media stream.
What you need configured
| For | Requires |
|---|---|
| Any call | LLM_*, an STT provider, a TTS provider |
| Real telephony | A connected Plivo or Exotel integration |
| Outbound campaigns | PUBLIC_BASE_URL — the provider must reach your webhooks |
| Knowledge-base answers | AWS_* / BEDROCK_* for embeddings |
| Human handoff | A workflow with handoff enabled and a destination number |