v5

latestOpenAPI 3.1.02026-08-025631,1012.8 MB
Bucket Syncs

Create Sync Configuration

Create a sync configuration for automated storage ingestion.

Establishes automated synchronization between an external storage provider and a Mixpeek bucket. The sync monitors the source path and ingests files according to the specified mode and filters.

Supported Providers: google_drive, s3, snowflake, sharepoint, tigris

Built-in Robustness:

  • Dead Letter Queue (DLQ): Failed objects tracked with 3 retries
  • Idempotent ingestion: Deduplication prevents duplicate objects
  • Distributed locking: Prevents concurrent sync execution
  • Rate limit handling: Automatic backoff on 429 responses
  • Metrics: Duration, files synced/failed, batches created

Sync Modes (the request enum accepts exactly these two):

  • initial_only: Single bulk import, then sync stops
  • continuous: Polling-based monitoring (polling_interval_seconds, max 900)

(Docstring previously advertised one_time/scheduled, which the enum rejects with a 422 — modes here must match SyncCreateRequest.)

post/v1/buckets/{bucket_id}/syncs

Path parameters

bucket_idstring required

Request body

connection_idstring required

REQUIRED. Storage connection identifier to sync from. Must reference an existing connection created via POST /organizations/connections. The connection defines the storage provider and credentials. Supported providers: google_drive, s3, snowflake, sharepoint, tigris.

source_pathstring nullable

REQUIRED unless provider_filters.use_search_api is true (search-API syncs don't traverse a path; defaults to '/'). Source path within the storage provider to monitor and sync. Path format varies by provider: - s3/tigris: 'bucket-name/prefix' or 'bucket-name'. - google_drive: folder ID or path like '/Marketing/Assets'. - sharepoint: '/sites/SiteName/Shared Documents/folder'. - snowflake: 'DATABASE.SCHEMA.TABLE' or just 'TABLE' if defaults set.

sync_mode'initial_only' | 'continuous'

Supported sync modes for external storage ingestion.

file_filtersobject nullable

OPTIONAL. Filters to control which files are synced. When omitted, all files in source_path are synced. Supported filters: - include_patterns: Glob patterns to include (e.g., ['.mp4', '.mov']). - exclude_patterns: Glob patterns to exclude (e.g., ['.tmp', '.DS_Store']). - extensions: File extensions to include (e.g., ['.mp4', '.jpg']). - min_size_bytes: Minimum file size in bytes. - max_size_bytes: Maximum file size in bytes. - modified_after: ISO datetime, only sync files modified after this time. - mime_types: List of MIME types to include (e.g., ['video/', 'image/jpeg']).

polling_interval_secondsinteger

Interval in seconds between polling checks for new files. OPTIONAL. Defaults to 300 seconds (5 minutes). Must be between 30 and 86400 seconds (0.5 minutes to 1 day). Only applies to 'continuous' and 'scheduled' sync modes. Lower values mean faster detection but higher API usage.

batch_sizeinteger

Number of files to process in each batch during sync. OPTIONAL. Defaults to 50 files per batch. Must be between 1 and 100. Larger batches improve throughput but require more memory. Smaller batches provide more granular progress tracking.

skip_batch_submissionboolean

If True, sync objects to the bucket without creating or submitting batches for collection processing. Objects are created in the bucket but no tier processing is triggered. Useful for bulk data migration or when you want to manually control when processing occurs. OPTIONAL. Defaults to False (batches are created and submitted).

skip_duplicatesboolean

If True, skip files whose source ID already exists in the bucket. If False, replace existing objects when re-syncing. OPTIONAL. Defaults to True.

sync_fromstring date-time nullable

OPTIONAL. Seed the incremental watermark: only assets modified after this timestamp are ingested — the historical backlog is skipped. Use for a 'freshness lane': a second config on an already-backfilling source (with a distinct source_path label) that keeps NEW uploads landing within minutes while the backfill config churns. Omit for a full sync from the beginning.

provider_filtersobject nullable

OPTIONAL. Provider-specific pre-filters pushed down to the storage API call. Applied BEFORE file_filters (which are client-side). Each provider defines its own filter schema. Examples: - Iconik: {'collection_ids': ['col_abc']} - Google Drive: {'shared_drive_id': '0AH-Xabc123'} - S3: {'prefix': 'videos/'}

metadataobject nullable

Optional custom metadata to attach to the sync configuration. NOT REQUIRED. Arbitrary key-value pairs for tagging and organization. Common uses: project tags, environment labels, cost centers. Maximum 50 keys, values must be JSON-serializable.

Example request

{
  "batch_size": 50,
  "connection_id": "conn_s3_prod",
  "file_filters": {
    "extensions": [
      ".mp4",
      ".mov"
    ],
    "min_size_bytes": 1024
  },
  "metadata": {
    "project": "video-analysis"
  },
  "polling_interval_seconds": 300,
  "source_path": "media-bucket/videos/raw",
  "sync_mode": "continuous"
}

Response

Successful Response

sync_config_idstring

Unique identifier for the sync configuration.

bucket_idstring required

Target bucket identifier (e.g. 'bkt_marketing_assets').

connection_idstring required

Storage connection identifier (e.g. 'conn_abc123').

internal_idstring required

Organization internal identifier (multi-tenancy scope).

namespace_idstring required

Namespace identifier owning the bucket.

source_pathstring required

Source path in the external storage provider. Format varies by provider: s3/tigris='bucket/prefix', google_drive='folder_id', sharepoint='/sites/Name/Documents', snowflake='DB.SCHEMA.TABLE'.

sync_mode'initial_only' | 'continuous'

Supported sync modes for external storage ingestion.

polling_interval_secondsinteger

Polling interval in seconds (continuous mode). Up to 86400 (1 day) — slow intervals are a legitimate ops throttle (e.g. deliberately deprioritizing freshness lanes during a backfill), and the read model must accept any value the platform itself may have stored.

batch_sizeinteger

Number of files processed per sync batch.

create_object_on_confirmboolean

Whether objects should be created immediately after confirmation.

skip_duplicatesboolean

Skip files whose hashes already exist in the bucket.

skip_batch_submissionboolean

Sync-only mode: download and store files in the bucket without running them through the collection processing pipeline. Set to True during initial bulk ingestion, then flip to False to trigger processing once all files are synced.

status'PENDING' | 'QUEUED' | 'IN_PROGRESS' | 'PROCESSING' | 'COMPLETED' | 'COMPLETED_WITH_ERRORS' | 'FAILED' | 'CANCELED' | 'INTERRUPTED' | 'UNKNOWN' | 'SKIPPED' | 'DRAFT' | 'ACTIVE' | 'ARCHIVED' | 'SUSPENDED'

Enumeration of task statuses for tracking asynchronous operations.

Task statuses indicate the current state of asynchronous operations like batch processing, object ingestion, clustering, and taxonomy execution.

Status Categories: Operation Statuses: Track progress of async operations Lifecycle Statuses: Track entity state (buckets, collections, namespaces)

Values: PENDING: Task is queued but has not started processing yet IN_PROGRESS: Task is currently being executed PROCESSING: Task is actively processing data (similar to IN_PROGRESS) COMPLETED: Task finished successfully with no errors COMPLETED_WITH_ERRORS: Task finished but some items failed (partial success) FAILED: Task encountered an error and could not complete CANCELED: Task was manually canceled by a user or system UNKNOWN: Task status could not be determined SKIPPED: Task was intentionally skipped DRAFT: Task is in draft state and not yet submitted

ACTIVE: Entity is active and operational (for buckets, collections, etc.)
ARCHIVED: Entity has been archived
SUSPENDED: Entity has been temporarily suspended

Terminal Statuses: COMPLETED, COMPLETED_WITH_ERRORS, FAILED, CANCELED are terminal statuses. Once a task reaches these states, it will not transition to another state.

Partial Success Handling: COMPLETED_WITH_ERRORS indicates that the operation completed but some documents/items failed. The task result includes: - List of successful items - List of failed items with error details - Success rate percentage This allows clients to handle partial success scenarios appropriately.

Polling Guidance: - Poll tasks in PENDING, QUEUED, IN_PROGRESS, or PROCESSING states - Stop polling when task reaches COMPLETED, COMPLETED_WITH_ERRORS, FAILED, or CANCELED - Use exponential backoff (1s → 30s) when polling

is_activeboolean

Convenience flag used for filtering active syncs.

total_files_discoveredinteger

Cumulative count of files found in source across all runs.

total_files_syncedinteger

Cumulative count of successfully synced files.

total_files_failedinteger

Cumulative count of failed files (sent to DLQ after 3 retries).

total_bytes_syncedinteger

Cumulative bytes transferred across all runs.

created_atstring date-time

When sync configuration was created.

updated_atstring date-time

Last modification timestamp.

last_sync_atstring date-time nullable

When last successful sync completed. Used for incremental syncs.

per_shard_last_sync_atobject

Per-shard last-sync timestamps keyed by shard value (e.g. collection_id). When a new shard is added, its absence here forces a full scan even if the global last_sync_at is set.

next_sync_atstring date-time nullable

Scheduled time for next sync (continuous/scheduled modes).

created_by_user_idstring required

User identifier that created the sync configuration.

last_errorstring nullable

Most recent error message if sync attempts failed.

consecutive_failuresinteger
provider_filtersobject

Provider-specific pre-filters pushed down to the API call. The sync engine passes these to iter_objects() without interpretation. Each provider defines its own schema. Applied BEFORE file_filters. Examples: Iconik {'collection_ids': [...]}, Google Drive {'shared_drive_id': '...'}

source_typestring nullable

Storage provider type for API progress views (for example: s3, google_drive, iconik).

metadataobject

Arbitrary metadata supplied by the user.

locked_by_worker_idstring nullable

Worker ID that currently holds the lock for this sync

locked_atstring date-time nullable

Timestamp when lock was acquired

lock_expires_atstring date-time nullable

Timestamp when lock expires (for stale lock recovery)

pending_full_syncboolean

A full sweep was requested (trigger?full_sync=true) while a run held the lock. The finishing run dispatches it automatically on lock release.

pausedboolean

Whether sync is currently paused (user-controlled)

pause_reasonstring nullable

Reason for pause

paused_atstring date-time nullable

Timestamp when paused

paused_by_user_idstring nullable

User who paused the sync

max_objects_per_runinteger

Hard cap on objects per sync run (prevents runaway syncs)

max_batch_chunk_sizeinteger

Maximum objects per batch chunk

batch_chunk_sizeinteger

Number of objects per batch chunk (for concurrent processing)

current_sync_run_idstring nullable

UUID for current/last sync run

sync_run_counterinteger

Increments on each sync execution

batch_idsstring[]

List of batch IDs created by this sync

task_idsstring[]

List of task IDs for batches

batches_createdinteger

Total number of batches created

resume_enabledboolean

Whether resuming partial runs is enabled

resume_cursorstring nullable

Last page/cursor processed (for paginated APIs like Google Drive)

resume_last_primary_keystring nullable

Last primary key processed (for database syncs with stable ordering)

resume_objects_processedinteger

Count of objects processed in current/last run

resume_checkpoint_frequencyinteger

How often to checkpoint (in objects). Default: every 1000 objects

current_cursorstring nullable

Convenience mirror of the current resume cursor for API progress views.

sync_checkpointsobject

Per-(config, shard) high-water checkpoints keyed by shard key (e.g. collection_id for parallel fan-outs, 'pages_N_M' for page-range shards, 'default' for unsharded runs). Each entry holds: pass_id (lexicographically-ordered pass marker), cursor (provider cursor, e.g. JSON-encoded Iconik search_after), objects_processed (forward-only progress guard), modified_since (incremental filter frozen at pass start), completed_at (set when the shard drained its source — the next cycle wraps around to a fresh full pass only after the polling cadence elapses), and updated_at. A NEW job resumes each shard from its checkpoint instead of re-walking from page 1 (2026-06-11 re-scan treadmill).

scheduleobject nullable

Derived scheduling summary: mode, interval, next run, and last successful run.

sync_progressobject

Derived progress summary for API observability.

lockedboolean required

Whether a worker currently holds this sync's run lock.