Skip to main content

Knowledge base

What the live suggestions are drawn from. Unlike nextneural_converse, documents are submitted as text, not uploaded files.

List documents​

GET /api/kb/documents · scope kb:read

{
"search_available": true,
"documents": [
{
"id": 3,
"title": "Objection handling — price",
"description": "Responses to price objections, by segment",
"vectorization_status": "completed",
"vectorization_error": null,
"chunk_count": 18,
"size_chars": 9420,
"vectorized_at": "2026-09-02T10:14:00Z",
"created_at": "2026-09-02T10:11:12Z"
}
]
}

chunk_count and size_chars are worth watching: a document that produced one chunk is probably too short to retrieve usefully, and one that produced hundreds will dilute results.

Get a document​

GET /api/kb/documents/{id} · scope kb:read

Includes the full content.

Create a document​

POST /api/kb/documents · scope kb:write

{
"title": "Objection handling — price",
"content": "When a customer says the price is too high, first establish what they are comparing it to…",
"description": "Responses to price objections, by segment"
}

title and content are both required. Extract the text yourself before sending — there is no file parsing here.

Delete a document​

DELETE /api/kb/documents/{id} · scope kb:write → 204

Vectorize a document​

POST /api/kb/documents/{id}/vectorize · scope kb:write

{ "description": "Objection responses, updated for Q4 pricing" }

The body is optional — description updates the stored description as part of the same call, which is convenient when you are re-vectorizing after an edit.

Chunks and embeds the document. A document is not searchable until this succeeds — creating one is not enough.

Statuses​

vectorization_statusMeaning
noneCreated, not yet embedded (the default)
processingIn progress
doneAvailable to suggestions
errorFailed — see vectorization_error
These values differ from nextneural_converse's

nextneural_assist uses none / processing / done / error. nextneural_converse uses pending / processing / ready / failed / disabled for the same concept. Poll for the value the binary you are calling actually emits.

The vectorize call itself returns the terminal state directly:

{ "chunk_count": 18, "vectorization_status": "done" }
Starting a call with unvectorized documents is refused
{ "detail": "one or more selected documents aren't available for AI — enable them in Knowledgebase first" }

The check runs before the call is placed, so the failure arrives while the rep can still fix it — not mid-conversation.

Vectorization needs AWS_* / BEDROCK_* configured and pgvector available. Without them every attempt fails.

How suggestions use it​

Documents are selected per call through kb_document_ids when starting the call. Only those documents are searched.

That narrowing is the point. A rep on an insurance call should not be shown a property listing that happened to share some words — and with a large shared corpus, that is exactly what broad retrieval produces.

As the call is transcribed, each turn is matched against the selected documents and relevant passages are pushed to the rep's screen through the agent-assist socket.

Writing documents that retrieve well​

  • One topic per document. "Objection handling — price" retrieves better than "Sales playbook".
  • Lead with the answer. The first sentences carry most of the retrieval weight, and a rep reading mid-call has no time for preamble.
  • Use the customer's words, not internal jargon — the match is against what was actually said on the phone.
  • Keep them short. A rep can act on three sentences; they cannot read a page while someone is talking to them.