v5

latestOpenAPI 3.1.02026-08-025631,1012.8 MB
Bucket Objects

Get Object

This endpoint retrieves an object by its ID from the specified bucket.

**Presigned URLs**: Set `return_presigned_urls=true` query parameter to generate fresh presigned download URLs
for all blobs with S3 storage (default: false). URLs are added to each blob's properties as
`presigned_url` and expire after 1 hour.

**Document count**: `document_count` (how many documents this object produced, via vector-store
lineage) is computed by default. It fans out to the vector partition, which on a cold/serverless
partition can be slow — pass `include_document_count=false` to skip it for latency-sensitive,
interactive views (e.g. an object-detail modal) that don't render the count. Even when requested,
the count is best-effort and bounded by a deadline: it returns `null` rather than blocking the
response if the partition is cold.
get/v1/buckets/{bucket_identifier}/objects/{object_identifier}

Path parameters

bucket_identifierstring required

The unique identifier of the bucket.

The unique identifier of the bucket.

object_identifierstring required

The unique identifier of the object.

The unique identifier of the object.

Query parameters

return_presigned_urlsboolean

Generate fresh presigned download URLs for all blobs with S3 storage

Generate fresh presigned download URLs for all blobs with S3 storage

include_document_countboolean

Compute document_count via vector-store lineage (default true). Pass false to skip the vector fan-out for latency-sensitive views that don't render the count.

Compute document_count via vector-store lineage (default true). Pass false to skip the vector fan-out for latency-sensitive views that don't render the count.

Response

Successful Response

object_idstring

Unique identifier for the object

bucket_idstring required

ID of the bucket this object belongs to

key_prefixstring nullable

Storage key/path of the object, this will be used to retrieve the object from the storage. It is similar to a file path. If not provided, it will be placed in the root of the bucket.

status'PENDING' | 'QUEUED' | 'IN_PROGRESS' | 'PROCESSING' | 'COMPLETED' | 'COMPLETED_WITH_ERRORS' | 'FAILED' | 'CANCELED' | 'INTERRUPTED' | 'UNKNOWN' | 'SKIPPED' | 'DRAFT' | 'ACTIVE' | 'ARCHIVED' | 'SUSPENDED'

Enumeration of task statuses for tracking asynchronous operations.

Task statuses indicate the current state of asynchronous operations like batch processing, object ingestion, clustering, and taxonomy execution.

Status Categories: Operation Statuses: Track progress of async operations Lifecycle Statuses: Track entity state (buckets, collections, namespaces)

Values: PENDING: Task is queued but has not started processing yet IN_PROGRESS: Task is currently being executed PROCESSING: Task is actively processing data (similar to IN_PROGRESS) COMPLETED: Task finished successfully with no errors COMPLETED_WITH_ERRORS: Task finished but some items failed (partial success) FAILED: Task encountered an error and could not complete CANCELED: Task was manually canceled by a user or system UNKNOWN: Task status could not be determined SKIPPED: Task was intentionally skipped DRAFT: Task is in draft state and not yet submitted

ACTIVE: Entity is active and operational (for buckets, collections, etc.)
ARCHIVED: Entity has been archived
SUSPENDED: Entity has been temporarily suspended

Terminal Statuses: COMPLETED, COMPLETED_WITH_ERRORS, FAILED, CANCELED are terminal statuses. Once a task reaches these states, it will not transition to another state.

Partial Success Handling: COMPLETED_WITH_ERRORS indicates that the operation completed but some documents/items failed. The task result includes: - List of successful items - List of failed items with error details - Success rate percentage This allows clients to handle partial success scenarios appropriately.

Polling Guidance: - Poll tasks in PENDING, QUEUED, IN_PROGRESS, or PROCESSING states - Stop polling when task reaches COMPLETED, COMPLETED_WITH_ERRORS, FAILED, or CANCELED - Use exponential backoff (1s → 30s) when polling

errorstring nullable

The error message if the object failed to process.

created_atstring date-time nullable

Timestamp when the object was created. Automatically populated by the system.

updated_atstring date-time nullable

Timestamp when the object was last updated. Automatically populated by the system.

document_countinteger nullable

Number of documents produced from this object across all collections. Populated on GET requests. Null on list responses (expensive query). Use this to check if an object has already been processed.

Example response

{
  "blobs": [
    {
      "blob_id": "blob_1",
      "data": {
        "num_pages": 5,
        "title": "Service Agreement 2024"
      },
      "key_prefix": "/contract-2024/content.pdf",
      "metadata": {
        "author": "John Doe",
        "department": "Legal"
      },
      "property": "content",
      "type": "PDF"
    }
  ],
  "bucket_id": "bkt_9xy8z7",
  "content_hash": "28a9f5e8...",
  "created_at": "2024-10-21T10:30:00Z",
  "key_prefix": "/contract-2024",
  "metadata": {
    "category": "contracts",
    "year": 2024
  },
  "object_id": "obj_123abc456def",
  "skip_duplicates": false,
  "status": "DRAFT",
  "updated_at": "2024-10-21T10:30:00Z"
}