Skip to main content

Runs

One practice call: a scenario played against a deployed agent, then scored.

Create a run​

POST /api/nn-coach/runs — any member

{
"scenario_uuid": "sc1a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8",
"target_kind": "flow",
"target_uuid": "f1g2h3i4-j5k6-7890-abcd-ef1234567890",
"mode": "text",
"agent_prompt": "You are a renewals specialist for Pragati AI…",
"agent_model": "anthropic/claude-sonnet-4",
"seed": 42,
"wait": false
}
FieldRequiredNotes
scenario_uuid✓

| target_kind | — | flow (the default), voice_agent, workflow, or human_rep | | target_uuid | ✓ | The deployed agent under test | | mode | ✓ | text or voice | | agent_prompt | — | The prompt the rep is played with — the thing under test | | agent_model | — | Model playing the rep | | proxy_model | — | Overrides the service default for this run | | tts_provider | — | Overrides the service default for this run | | agent_opens | — | Absent means "use what the scenario says" | | seed | — | For reproducible runs | | wait | — | Run inline and return the finished run |

An empty agent_prompt falls back to a generic sales agent — useful as a baseline to measure a real prompt against.

proxy_model and tts_provider are sent per run rather than stored, because the coach keeps no per-organization settings table. The caller holds the preference and states it each time.

StatusCause
400{"detail": "scenario_uuid is required"}
400mode must be "text" or "voice"
400voice mode can only call a flow, because the audio transport is the flow's own browser-test socket
400voice mode needs target_uuid: the flow whose agent should answer the call
400target_uuid is required: it names the deployed agent to test. Runs are always played against a real flow, so there is no default to fall back on
404{"detail": "scenario not found"}

wait​

Off by default, because a run takes tens of seconds. Leave it off and poll; turn it on for a script that has nothing else to do.

Modes​

ModeTransportExercises
textLLM to LLMThe prompt, the routing, the knowledge
voiceTTS into the flow's own WebSocketEverything above, plus STT, VAD, barge-in, and synthesis

The mode changes the transport, not what is under test — a text run still drives a real deployed agent, which is why it needs target_uuid too.

Use text for regression runs over many scenarios; voice when the thing you doubt is the audio path.

List runs​

GET /api/nn-coach/runs — any member

Get a run​

GET /api/nn-coach/runs/{uuid} — any member

Returns the transcript, the scores, and the verdict.

Run audio​

GET /api/nn-coach/runs/{uuid}/audio — any member

The recording of a voice run.

Delete a run​

DELETE /api/nn-coach/runs/{uuid} — manager or above

Statuses​

StatusMeaning
queuedAccepted, not yet started
runningIn progress
completedFinished and scored
failedSomething in the run failed
timeoutExceeded its time budget

Verdicts​

VerdictMeaning
passThe objective was met and the rubric was satisfied
partialMixed — some criteria met
failNeither

Judging​

Scoring can be delegated to nextneural_intel via JUDGE_API_BASE_URL and JUDGE_API_TOKEN. The coach posts the transcript, the scenario objective, the hidden state, the objections, and the rubric fields; nextneural_intel scores it with the same extraction a real call receives.

That shared scoring path is the point: a trainee's practice score means the same thing as their score on a live call, rather than being a separate ladder that only correlates by accident.

The key the coach uses carries sales:simulations:judge and nothing else — it can score simulations without being able to create call records.