v1

latestOpenAPI 3.1.02026-07-263580186.3 KB
Search

Search document chunks

Embedding → hybrid vector search → optional reranking, returning ranked chunks with provenance. No LLM generation is performed.

Billing: 1 retrieval credit per request.

Relevance scoring (relevance_scoring): controls the relevance scoring stage.

  • scoring_and_filtering (default): Score candidates for relevance and only return those above the quality threshold.
  • scoring_only: Score every candidate for relevance but return them all, even low-scoring ones. Useful for building your own filtering logic.
  • none: Skip the relevance scoring step and return all candidates unfiltered. Fastest option, useful when you handle scoring yourself.

Omit relevance_scoring for the default; send none to skip scoring. skip_rerank is deprecated — true maps to none, false to scoring_and_filtering.

Result ordering: results are returned in descending order of score. With scoring_and_filtering or scoring_only, score equals the relevance score (scores.relevance, 0–1). With none, score is the combined retrieval score (higher is better, no fixed upper bound).

If the scoring model is temporarily unavailable, results are returned in retrieval order and a warnings array is included. Each warning has a code matching the degraded scores key (e.g. relevance) and a reason classifying the failure: model_not_found, timeout, service_error, or unknown. The warnings key is absent when all pipeline steps succeed.

Scoping: use workspace_id and/or tag_id to narrow results, or file_id to target specific files. file_id cannot be combined with workspace_id or tag_id (422). A 403 is returned when filters resolve to no authorized resources. When no filters are provided, search runs across all documents authorized for the API key.

Facet filtering: use content_type and attribute to narrow results by facet metadata. Content type uses colon-separated paths (e.g. legal:contract:nda). Repeated attribute entries are ANDed; values inside one entry are ORed with | (pipe, recommended). Example: attribute=fiscal_year:2024|2025&attribute=status:active → (fiscal_year 2024 OR 2025) AND (status active). Supports operators (>, >=, <, <=), prefix (name:prefix*), smart dates, and content-type scoping.

Modes:

  • text (default): hybrid text search
  • vision: VLM-embedded page image search

Images: set include_image=true to receive a base64-encoded page image with each result. In text mode the image is fetched from the VisionChunk covering the chunk's start page (empty string if no vision index exists for that page).

Bounding boxes (PDF only): set include_bboxes=true to append a bboxes array to each result, giving the merged rectangles of the chunk's text on the source PDF (raw PDF points, top-left origin with y extending downward) so you can overlay highlights without re-locating the chunk. One rectangle per logical group; a chunk spanning two pages produces at least one rectangle per page. Available for PDF documents in text mode only — returns an empty list for non-PDF, vision-mode, or pre-v2.2.1 chunks. When include_bboxes=false (default) the bboxes key is omitted.

post/api/v3/search

Request body

content_typestring[]

Filter by content type path. Multiple values are OR. Exact-or-subtree matching by default (e.g. legal matches legal, legal:contract). Wildcards: *contract* (contains), legal:contract* (prefix).

attributestring[]

Filter by attribute value. Repeated attribute entries are ANDed; values inside one entry are ORed with | (pipe is the recommended OR delimiter — comma also works but can be ambiguous with multi-key values). Example: attribute=fiscal_year:2024|2025&attribute=status:active → (fiscal_year 2024 OR 2025) AND (status active). Formats: name (has any value), name:value (exact), name:>value / name:>=value (gt/gte), name:<value / name:<=value (lt/lte), name:prefix* (starts with, case-insensitive), name:*text* (contains, case-insensitive), name:a|b (OR). Smart dates: filing_date:2023 (year), filing_date:2023-06 (month). Type-aware: booleans (true/false), multi-select (membership check). Scoped: content_type(legal:compliance).regulation:AML.

querystring required

Natural-language search query. Maximum 1500 characters.

max_resultsinteger

Maximum number of chunks to return after reranking. Range: 1–50.

workspace_idinteger[]

Restrict search to these workspace IDs. Cannot combine with file_id.

tag_idinteger[]

Restrict to documents carrying any of these tag IDs (OR). Cannot combine with file_id.

file_idinteger[]

Restrict to specific file IDs. Cannot combine with workspace_id or tag_id.

mode'text' | 'vision'
  • text - text
  • vision - vision
relevance_scoring'none' | 'scoring_only' | 'scoring_and_filtering'
  • none - none
  • scoring_only - scoring_only
  • scoring_and_filtering - scoring_and_filtering
skip_rerankboolean

Deprecated — use relevance_scoring. true → relevance_scoring=none, false → relevance_scoring=scoring_and_filtering. Ignored when relevance_scoring is provided.

include_imageboolean

Append a base64-encoded page image to each result.

include_bboxesboolean

Append merged bounding boxes (in PDF points, top-left origin) to each result so callers can overlay chunk highlights on PDF pages. PDF documents in text mode only — non-PDF and vision-mode results always return an empty list. Omitted from the response entirely when false.

Response

Ranked search results. Empty array if no documents match.

explainobject

Scoring breakdown. Present only when explain=true and SEARCH_EXPLAIN_MODE is enabled.