v1
latestOpenAPI 3.1.02026-07-263580186.3 KBAsk a question over your documents
Retrieval-augmented generation: searches your indexed corpus, then generates an LLM answer grounded in the retrieved passages.
Modes:
- stream=false (default): returns a single JSON response with results and answer.
- stream=true: returns Server-Sent Events — event: sources (retrieved chunks), event: token (answer tokens), event: done (stream complete), or event: error (generation failure).
Model: defaults to mistral-large-latest (flagship, best answer quality). Pass model=alfred-ft5 for the lighter, faster LightOn fine-tune. Company-specific custom models (custom-{company_id}-{uuid}) are also accepted. Any other value returns 422.
Relevance scoring: relevance scoring always runs in scoring_and_filtering mode — candidates are scored for relevance and only those above the quality threshold are used as context. score equals the relevance score (scores.relevance, 0–1). Results are returned in descending order of score. If the scoring model is temporarily unavailable, score falls back to the combined retrieval score (higher is better, no fixed upper bound) and scores.relevance is null.
Scoping: same rules as /api/v3/search — use workspace_id and/or tag_id to narrow results, or file_id to target specific files. file_id cannot be combined with workspace_id or tag_id (422).
Facet filtering: use content_type and attribute to narrow results by facet metadata. Content type uses colon-separated paths (e.g. legal:contract:nda). Repeated attribute entries are ANDed; values inside one entry are ORed with | (pipe, recommended). Example: attribute=fiscal_year:2024|2025&attribute=status:active → (fiscal_year 2024 OR 2025) AND (status active). Supports operators (>, >=, <, <=), prefix (name:prefix*), smart dates, and content-type scoping.
If the reranker is temporarily unavailable, results are returned in retrieval order and each result item includes a warnings array. Each warning has a code matching the degraded scores key (e.g. relevance) and a reason classifying the failure: model_not_found, timeout, service_error, or unknown. The warnings key is absent from result items when all pipeline steps succeed.
Billing: 1 search-with-generation credit per request.
Request body
Response
Synchronous mode (stream=false): complete answer with sources.
Streaming mode (stream=true): Server-Sent Events with event: sources, event: token, and event: done (or event: error).