v5

latestOpenAPI 3.1.02026-08-025631,1012.8 MB
Collection Documents

Get a document by ID.

Get a document by ID.

get/v1/collections/{collection_identifier}/documents/{document_id}

Path parameters

collection_identifierstring required

The ID of the collection.

The ID of the collection.

document_idstring required

The ID of the document to retrieve.

The ID of the document to retrieve.

Query parameters

return_presigned_urlsboolean

Generate fresh presigned download URLs for all blobs with S3 storage

Generate fresh presigned download URLs for all blobs with S3 storage

return_vectorsboolean nullable
return_vector_namesboolean

Include a '_vectors' field listing available vector names for this document (without actual embedding data)

Include a '_vectors' field listing available vector names for this document (without actual embedding data)

expandstring nullable

Comma-separated fields to resolve inline under an 'expanded' key. Reserved lineage keywords: 'parent' (immediate upstream document), 'root_object' (originating bucket object), 'ancestors' (full chain root → parent), 'children' (direct downstream documents, max 100). Any other value is treated as a user field path containing doc* references (supports dot-notation, e.g. 'items.product_id'). Max 50 unique references for non-lineage fields. Depth is limited to 1 (no recursive expansion).

Comma-separated fields to resolve inline under an 'expanded' key. Reserved lineage keywords: 'parent' (immediate upstream document), 'root_object' (originating bucket object), 'ancestors' (full chain root → parent), 'children' (direct downstream documents, max 100). Any other value is treated as a user field path containing doc* references (supports dot-notation, e.g. 'items.product_id'). Max 50 unique references for non-lineage fields. Depth is limited to 1 (no recursive expansion).

Response

Successful Response

document_idstring required

REQUIRED. Unique identifier for the document. Format: 'doc_' prefix + alphanumeric characters. Use for: API queries, references, filtering.

collection_idstring required

REQUIRED. ID of the collection this document belongs to. Format: 'col_' prefix + alphanumeric characters. Use for: Collection-scoped queries, filtering.

Example response

{
  "_internal": {
    "collection_id": "col_articles",
    "created_at": "2025-10-31T10:00:00Z",
    "document_id": "doc_f8966ff29c18e20c6b45e053",
    "internal_id": "org_abc123",
    "lineage": {
      "chain": [
        {
          "collection_id": "col_articles",
          "feature_extractor_id": "text_extractor_v1",
          "timestamp": "2025-10-31T10:00:00Z"
        }
      ],
      "path": "bkt_content/col_articles",
      "root_bucket_id": "bkt_content",
      "root_object_id": "obj_article_001",
      "source_object_id": "obj_article_001",
      "source_type": "bucket"
    },
    "metadata": {
      "ingestion_status": "COMPLETED"
    },
    "modality": "text",
    "namespace_id": "ns_xyz789",
    "source_blobs": [
      {
        "blob_id": "blob_text_001",
        "blob_property": "content",
        "blob_type": "text"
      }
    ],
    "updated_at": "2025-10-31T10:00:00Z"
  },
  "author": "Dr. Smith",
  "collection_id": "col_articles",
  "description": "Text document with _internal structure",
  "document_id": "doc_f8966ff29c18e20c6b45e053",
  "title": "AI in Healthcare"
}