Skip to main content

Knowledge base

Documents the agent can quote from during a call. Retrieval is hybrid: lexical search always works, semantic search needs embeddings configured.

Check what search is available​

GET /api/voice/knowledgebase/status

{ "semantic_search": true, "lexical_search": true }

Worth calling before offering an "enable for AI" control — without AWS_* / BEDROCK_* configured, or without pgvector installed, semantic_search is false and vectorization fails every time it is attempted.

List documents​

GET /api/voice/knowledgebase

{
"documents": [
{
"uuid": "kb1a2b3c-4d5e-6f70-8192-a3b4c5d6e7f8",
"title": "Refund policy 2026",
"description": "Current refund windows and exceptions",
"filename": "refund-policy.pdf",
"vectorization_status": "completed",
"created_at": "2026-09-02T10:11:12Z"
}
]
}
vectorization_statusMeaning
pendingUploaded, not yet embedded (the default)
processingEmbedding in progress
readyAvailable to semantic search
failedEmbedding failed; the message is stored with the row
disabledExplicitly withdrawn from AI use
Poll for ready, not completed

These five are enforced by a database constraint. nextneural_assist uses a different vocabulary for the same concept — do not share a polling helper between the two.

Upload a document​

POST /api/voice/knowledgebase — multipart/form-data

PartRequiredNotes
file✓The document. Max 10 MB
title—Defaults to the filename
description—Free text
curl -X POST https://your-host/api/voice/knowledgebase \
-H "Authorization: Bearer <token>" \
-F "[email protected]" \
-F "title=Refund policy 2026"

201 Created returns the document record.

The 10 MB cap is not arbitrary — the whole document is held in memory while it is chunked. Split large manuals into sections, which also retrieves better.

StatusCause
400{"detail": "a file is required"}
400{"detail": "could not read the uploaded file"}

Delete a document​

DELETE /api/voice/knowledgebase/{uuid} → 204

Removes the document and its chunks and vectors.

Enable for AI​

POST /api/voice/knowledgebase/{uuid}/enable-ai

Chunks the document and embeds it. Returns immediately — embedding runs in the background:

{ "vectorization_status": "processing" }

Poll the list endpoint until the status reaches ready or failed.

Disable for AI​

POST /api/voice/knowledgebase/{uuid}/disable-ai

{ "vectorization_status": "disabled" }

The document stays, but is no longer searched. Use this rather than deleting when a policy is superseded but still worth keeping.

How retrieval works during a call​

Retrieval runs before the caller's turn is added to the history, so the reference material sits next to the question it answers — which is where models attend to it most reliably.

  • With a workflow: the router decides which sub-agent holds the turn, and only that agent's datasource_ids are searched.
  • Without one: plain retrieval runs across the enabled documents.

Retrieval is bounded to a 2-second budget. Past that the turn proceeds without context rather than leaving the caller in silence — a slow lookup should cost an unsupported answer, not a dead line.

Hybrid search combines lexical (BM25) and vector results. With embeddings unavailable, it runs lexical-only rather than failing.

Grounding​

After a reply is generated it is checked against the passage it was given. A reply that is not grounded is logged, not blocked:

[call cal-…] ungrounded reply (no overlap): "We offer a 90-day refund window."

Holding every reply until it could be verified would cost the sentence-by-sentence latency that makes the agent feel responsive, and cutting a caller off mid-answer is worse than a rare unsupported sentence. Watch these log lines when tuning prompts.