Voices
Two related things: the catalog of stock voices with audition clips, and cloned voices built from a reference recording.
The catalog
GET /api/voice/voice-samples/catalog
Every speaker and language a clip could be generated for, whether or not one exists yet. This is what a voice picker renders — a voice with no clip can still be chosen, and offered for generation.
Generated clips
GET /api/voice/voice-samples
The audition clips that exist, with the URL each is served from.
{ "samples": [] }
Generate clips
POST /api/voice/voice-samples/generate
{
"speakers": ["anushka", "abhilash"],
"languages": ["en-IN", "hi-IN"],
"force": false
}
| Field | Default | Notes |
|---|---|---|
speakers | all | Unknown names are rejected before the run starts |
languages | ["en-IN"] | A full run is 407 clips, so it must be asked for explicitly |
model | bulbul:v2 | |
force | false | Regenerate clips that already exist |
202 Accepted:
{ "status": "generating" }
Returns immediately — a full run is minutes of synthesis and would time out the connection. Watch progress by listing samples, or:
GET /api/voice/voice-samples/generate/status
{ "running": true }
| Status | Cause |
|---|---|
400 | SARVAM_API_KEY is not configured, or an unknown speaker or language |
409 | {"error": "a generation run is already in progress"} |
Cloned voices
Builds a new speaker from a short reference recording, through the configured
voice models service (NEXTNEURAL_TTS_URL / NEXTNEURAL_TTS_TOKEN).
List cloned voices
GET /api/voice/models/speakers
{ "speakers": [ { "name": "priya-clone", "deployed": true } ] }
Clone a voice
POST /api/voice/models/speakers — multipart/form-data
| Part | Required | Notes |
|---|---|---|
reference_audio | ✓ | 6 to 30 seconds of clean speech |
name | ✓ | Referenced later in voice_config.speaker |
language | — | Defaults to en |
curl -X POST https://your-host/api/voice/models/speakers \
-H "Authorization: Bearer <token>" \
-F "name=priya-clone" \
-F "language=hi" \
-F "[email protected]"
| Status | Cause |
|---|---|
400 | {"detail": "a reference audio clip is required"} |
400 | {"detail": "the reference clip is too large — 6 to 30 seconds of speech is enough"} |
409 | {"detail": "a voice with this name already exists"} |
502 | {"detail": "the voice models service could not be reached"} |
503 | Voice cloning is not configured on this deployment |
Deploy and undeploy
POST /api/voice/models/speakers/{name}/deploy
{ "name": "priya-clone", "deployed": true }
DELETE /api/voice/models/speakers/{name}/deploy — undeploys without deleting.
A voice must be deployed before it can synthesise. Undeploy frees the resource while keeping the clone.
Delete a cloned voice
DELETE /api/voice/models/speakers/{name}
| Status | Cause |
|---|---|
404 | {"detail": "speaker not found"} |
Deleting or undeploying a voice that an agent references leaves that agent
unable to speak. Clear or repoint the agent's voice_config.speaker first.
Using a voice
Reference it from the agent's voice_config:
{ "provider": "sarvam", "speaker": "priya-clone", "language_code": "hi-IN" }
Mid-call language switching
Some providers can retarget the voice mid-call; the rest keep the voice they started with, which is still correct for a single-language call. The switch only fires when the settled language changes, because providers treat a settings change as a boundary and doing it every turn interrupts synthesis for nothing. See language behaviour.