Knowledge base
Documents the agent can quote from during a call. Retrieval is hybrid: lexical search always works, semantic search needs embeddings configured.
Check what search is available
GET /api/voice/knowledgebase/status
{ "semantic_search": true, "lexical_search": true }
Worth calling before offering an "enable for AI" control — without AWS_* /
BEDROCK_* configured, or without pgvector installed, semantic_search is
false and vectorization fails every time it is attempted.
List documents
GET /api/voice/knowledgebase
{
"documents": [
{
"uuid": "kb1a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8",
"title": "Refund policy 2026",
"description": "Current refund windows and exceptions",
"filename": "refund-policy.pdf",
"vectorization_status": "completed",
"created_at": "2026-09-02T10:11:12Z"
}
]
}
vectorization_status | Meaning |
|---|---|
pending | Uploaded, not yet embedded (the default) |
processing | Embedding in progress |
ready | Available to semantic search |
failed | Embedding failed; the message is stored with the row |
disabled | Explicitly withdrawn from AI use |
ready, not completedThese five are enforced by a database constraint. nextneural_assist uses a
different vocabulary for the same
concept — do not share a polling helper between the two.
Upload a document
POST /api/voice/knowledgebase — multipart/form-data
| Part | Required | Notes |
|---|---|---|
file | ✓ | The document. Max 10 MB |
title | — | Defaults to the filename |
description | — | Free text |
curl -X POST https://your-host/api/voice/knowledgebase \
-H "Authorization: Bearer <token>" \
-F "[email protected]" \
-F "title=Refund policy 2026"
201 Created returns the document record.
The 10 MB cap is not arbitrary — the whole document is held in memory while it is chunked. Split large manuals into sections, which also retrieves better.
| Status | Cause |
|---|---|
400 | {"detail": "a file is required"} |
400 | {"detail": "could not read the uploaded file"} |
Delete a document
DELETE /api/voice/knowledgebase/{uuid} → 204
Removes the document and its chunks and vectors.
Enable for AI
POST /api/voice/knowledgebase/{uuid}/enable-ai
Chunks the document and embeds it. Returns immediately — embedding runs in the background:
{ "vectorization_status": "processing" }
Poll the list endpoint until the status reaches ready or failed.
Disable for AI
POST /api/voice/knowledgebase/{uuid}/disable-ai
{ "vectorization_status": "disabled" }
The document stays, but is no longer searched. Use this rather than deleting when a policy is superseded but still worth keeping.
How retrieval works during a call
Retrieval runs before the caller's turn is added to the history, so the reference material sits next to the question it answers — which is where models attend to it most reliably.
- With a workflow: the router decides which sub-agent holds the turn, and
only that agent's
datasource_idsare searched. - Without one: plain retrieval runs across the enabled documents.
Retrieval is bounded to a 2-second budget. Past that the turn proceeds without context rather than leaving the caller in silence — a slow lookup should cost an unsupported answer, not a dead line.
Hybrid search combines lexical (BM25) and vector results. With embeddings unavailable, it runs lexical-only rather than failing.
Grounding
After a reply is generated it is checked against the passage it was given. A reply that is not grounded is logged, not blocked:
[call cal-…] ungrounded reply (no overlap): "We offer a 90-day refund window."
Holding every reply until it could be verified would cost the sentence-by-sentence latency that makes the agent feel responsive, and cutting a caller off mid-answer is worse than a rare unsupported sentence. Watch these log lines when tuning prompts.