Runs
One practice call: a scenario played against a deployed agent, then scored.
Create a run
POST /api/nn-coach/runs — any member
{
"scenario_uuid": "sc1a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8",
"target_kind": "flow",
"target_uuid": "f1g2h3i4-j5k6-7890-abcd-ef1234567890",
"mode": "text",
"agent_prompt": "You are a renewals specialist for Pragati AI…",
"agent_model": "anthropic/claude-sonnet-4",
"seed": 42,
"wait": false
}
| Field | Required | Notes |
|---|---|---|
scenario_uuid | ✓ |
| target_kind | — | flow (the default), voice_agent, workflow, or human_rep |
| target_uuid | ✓ | The deployed agent under test |
| mode | ✓ | text or voice |
| agent_prompt | — | The prompt the rep is played with — the thing under test |
| agent_model | — | Model playing the rep |
| proxy_model | — | Overrides the service default for this run |
| tts_provider | — | Overrides the service default for this run |
| agent_opens | — | Absent means "use what the scenario says" |
| seed | — | For reproducible runs |
| wait | — | Run inline and return the finished run |
An empty agent_prompt falls back to a generic sales agent — useful as a
baseline to measure a real prompt against.
proxy_model and tts_provider are sent per run rather than stored, because
the coach keeps no per-organization settings table. The caller holds the
preference and states it each time.
| Status | Cause |
|---|---|
400 | {"detail": "scenario_uuid is required"} |
400 | mode must be "text" or "voice" |
400 | voice mode can only call a flow, because the audio transport is the flow's own browser-test socket |
400 | voice mode needs target_uuid: the flow whose agent should answer the call |
400 | target_uuid is required: it names the deployed agent to test. Runs are always played against a real flow, so there is no default to fall back on |
404 | {"detail": "scenario not found"} |
wait
Off by default, because a run takes tens of seconds. Leave it off and poll; turn it on for a script that has nothing else to do.
Modes
| Mode | Transport | Exercises |
|---|---|---|
text | LLM to LLM | The prompt, the routing, the knowledge |
voice | TTS into the flow's own WebSocket | Everything above, plus STT, VAD, barge-in, and synthesis |
The mode changes the transport, not what is under test — a text run still
drives a real deployed agent, which is why it needs target_uuid too.
Use text for regression runs over many scenarios; voice when the thing you
doubt is the audio path.
List runs
GET /api/nn-coach/runs — any member
Get a run
GET /api/nn-coach/runs/{uuid} — any member
Returns the transcript, the scores, and the verdict.
Run audio
GET /api/nn-coach/runs/{uuid}/audio — any member
The recording of a voice run.
Delete a run
DELETE /api/nn-coach/runs/{uuid} — manager or above
Statuses
| Status | Meaning |
|---|---|
queued | Accepted, not yet started |
running | In progress |
completed | Finished and scored |
failed | Something in the run failed |
timeout | Exceeded its time budget |
Verdicts
| Verdict | Meaning |
|---|---|
pass | The objective was met and the rubric was satisfied |
partial | Mixed — some criteria met |
fail | Neither |
Judging
Scoring can be delegated to nextneural_intel
via JUDGE_API_BASE_URL and JUDGE_API_TOKEN. The coach posts the transcript,
the scenario objective, the hidden state, the objections, and the rubric
fields; nextneural_intel scores it with the same extraction a real call receives.
That shared scoring path is the point: a trainee's practice score means the same thing as their score on a live call, rather than being a separate ladder that only correlates by accident.
The key the coach uses carries sales:simulations:judge and nothing else — it
can score simulations without being able to create call records.