---
title: "Create Upload"
method: POST
path: "/v1/buckets/{bucket_identifier}/uploads"
tags: ["Bucket Uploads"]
---

# Create Upload

`POST /v1/buckets/{bucket_identifier}/uploads`

Generate a presigned URL for direct S3 upload.

    This endpoint validates all requirements BEFORE generating the presigned URL,
    ensuring immediate feedback if something is wrong (bucket inactive, quota exceeded, etc.).

    **Duplicate Detection (Enabled by Default)**:
    - If `file_hash` provided and `skip_duplicates=true`: Checks for existing upload
    - If duplicate found: Returns existing upload (200 OK) with `is_duplicate=true`
    - If new file: Returns presigned URL (201 Created) with `is_duplicate=false`

    **Two-Step Flow**:
    1. Call this endpoint → Get presigned URL
    2. PUT file to presigned URL → Upload directly to S3
    3. Call confirm endpoint → Verify upload and create object

## Path parameters

- `bucket_identifier` string, required — The unique identifier of the bucket

## Request body

- CreateUploadRequest — Request to generate a presigned URL for direct S3 upload. ⚠️ ⚠️ ⚠️ THIS IS THE PRESIGNED URL SYSTEM ⚠️ ⚠️ ⚠️ This endpoint (POST /buckets/{id}/uploads) is the COMPLETE presigned URL system. It handles: - ✅ Presigned URL generation (S3 PUT URLs) - ✅ Upload tracking and status management - ✅ Validation (quotas, file size, content type, schema) - ✅ Duplicate detection - ✅ Auto object creation on confirmation - ✅ Returns upload_id for later reference DO NOT CREATE A NEW PRESIGNED UPLOAD ENDPOINT! If you need presigned URLs, use this existing system. If you think you need a new endpoint: 1. Check if this system already does it (it probably does) 2. Extend this system instead of creating redundancy 3. See api/buckets/uploads/services.py for implementation Integration Points: - Object creation: Use upload_id in CreateBlobRequest.upload_id field - See: shared/buckets/objects/blobs/models.py::CreateBlobRequest - See: api/buckets/objects/canonicalization.py::resolve_upload_reference() Workflow: 1. POST /buckets/{id}/uploads → Returns presigned_url + upload_id 2. PUT presigned_url with file content (client uploads directly to S3) 3. POST /uploads/{upload_id}/confirm → REQUIRED to finalize upload 4. Object is created automatically (default behavior) ⚠️ IMPORTANT: Step 3 (confirm) is REQUIRED! S3 presigned URLs have no callback mechanism - the API cannot detect when your upload to S3 completes. You MUST call the confirm endpoint to: - Verify the file exists in S3 - Validate integrity (ETag/size) - Create the bucket object - Mark upload as COMPLETED If you don't confirm: - Upload stays in PENDING status forever - No object is created - File exists in S3 but is orphaned - Presigned URL expires (default: 1 hour) Use Cases: - Simple: Upload → confirm → object created automatically (default) - Advanced: Upload multiple files with create_object_on_confirm=false, then POST /buckets/{id}/objects with all upload_ids to create one object Requirements: - filename: REQUIRED, will be validated (no path traversal) - content_type: REQUIRED, must be valid MIME type - bucket_id: Comes from URL path parameter, not request body - All other fields: OPTIONAL with sensible defaults Note: The bucket_id comes from the URL path (/v1/buckets/{bucket_id}/uploads), not from the request body. The bucket is validated before generating presigned URL.
  - `filename` string, required — Name of the file to upload. REQUIRED. Must be a valid filename without path traversal characters (../, \). The filename is used to derive the blob_property if not explicitly provided. Examples: 'product_video.mp4', 'thumbnail.jpg', 'transcript.txt'
  - `content_type` string, required — MIME type of the file. REQUIRED. Must be a valid MIME type (e.g., 'video/mp4', 'image/jpeg', 'application/pdf'). The presigned URL will enforce this content type during upload. Used to validate compatibility with bucket schema if create_object_on_confirm=true.
  - `file_size_bytes` integer, nullable — Expected file size in bytes. OPTIONAL but RECOMMENDED. If provided, will be validated against: 1. Tier-based file size limits (100MB free, 5GB pro, 50GB enterprise) 2. Storage quota availability 3. Actual uploaded file size during confirmation. If not provided, quota checking is skipped until confirmation.
  - `presigned_url_expiration` integer — How long the presigned URL is valid, in seconds. OPTIONAL, defaults to 3600 (1 hour). Valid range: 60 seconds (1 minute) to 86400 seconds (24 hours). After expiration, the URL cannot be used and you must request a new one. Recommendation: Use shorter expiration (300-900 seconds) for security-sensitive files, longer expiration (3600-7200 seconds) for large files that take time to upload.
  - `metadata` object — Custom metadata for tracking purposes. OPTIONAL. Stored with the upload record for filtering and analytics. Does NOT affect the created bucket object (use object_metadata for that). Common uses: campaign tracking, user identification, upload source.
  - `create_object_on_confirm` boolean — Whether to automatically create a bucket object when upload is confirmed. OPTIONAL, defaults to TRUE (object is created automatically). If true (default): - Bucket MUST have a schema defined - blob_property must exist in bucket schema - content_type must match schema field type - Validation happens BEFORE generating presigned URL - Object is created automatically on confirmation. If false: - Upload is confirmed but no object is created - Use this when combining multiple uploads into one object - Reference the upload_id later in POST /buckets/{id}/objects.
  - `object_metadata` object, nullable — Metadata to attach to the created bucket object. OPTIONAL. Only used if create_object_on_confirm=true. This metadata will be: 1. Validated against bucket schema (if keys match schema fields) 2. Attached to the bucket object 3. Passed to downstream documents in connected collections. Example: {'priority': 'high', 'category': 'products', 'tags': ['featured']}
  - `blob_property` string, nullable — Property name for the blob in the bucket object. OPTIONAL. Defaults to filename without extension (e.g., 'product_video.mp4' → 'product_video'). If create_object_on_confirm=true: - Must exist in bucket schema - Must be alphanumeric with underscores only - Will be validated BEFORE generating presigned URL. Common values: 'video', 'image', 'thumbnail', 'transcript', 'content'.
  - `blob_type` string, nullable — Type of blob. OPTIONAL. Defaults to type derived from content_type (e.g., 'video/mp4' → 'VIDEO'). Must be a valid BucketSchemaFieldType if provided. Valid values: IMAGE, VIDEO, AUDIO, TEXT, PDF, DOCUMENT, etc. If create_object_on_confirm=true, will be validated against bucket schema field type.
  - `file_hash` string, nullable — SHA256 hash of the file content for duplicate detection. OPTIONAL. If provided: - System checks for existing confirmed uploads with same hash - If duplicate found and skip_duplicates=true, returns existing upload - Hash will be validated against actual S3 ETag during confirmation. If not provided: - Hash is calculated from S3 ETag after upload - Duplicate detection only happens during confirmation. Use case: Pre-calculate hash client-side to avoid uploading duplicates. Format: 64-character hexadecimal string (SHA256).
  - `skip_duplicates` boolean — Skip upload if a file with the same hash already exists. OPTIONAL, defaults to TRUE. If true (default): - If file_hash provided: System checks MongoDB for existing completed upload with same hash - If duplicate found: Returns existing upload details WITHOUT generating new presigned URL - If file_hash NOT provided: Duplicate check happens during confirmation using S3 ETag - Saves bandwidth, storage, and upload time by reusing existing files. If false: - Always generates new presigned URL even if file already uploaded - Creates separate upload record for same file content - Useful when you need distinct upload tracking for identical files. Recommendation: Keep default (true) unless you specifically need multiple upload records for same file.

## Response `200`

Successful Response

- UploadResponse — Response containing presigned URL and upload tracking information. This response includes everything needed to: 1. Upload your file to S3 using the presigned_url 2. Track the upload status using upload_id 3. Confirm the upload using the confirmation endpoint The presigned_url is time-limited and specific to this upload. After uploading to S3, call POST /v1/buckets/{bucket_id}/uploads/{upload_id}/confirm.
  - `upload_id` string, required — Unique identifier for this upload. Auto-generated. ⚠️ NEXT STEP: After uploading to S3, you MUST confirm: POST /v1/uploads/{upload_id}/confirm Other operations: - Check status: GET /v1/uploads/{upload_id} - Cancel upload: DELETE /v1/uploads/{upload_id} Format: 'upl_' followed by 16 random characters.
  - `bucket_id` string, required — Target bucket ID where object will be created
  - `filename` string, required — Name of the file to upload
  - `content_type` string, required — MIME type enforced by the presigned URL
  - `file_size_bytes` integer, nullable — Expected file size in bytes if provided in request. Will be validated during confirmation.
  - `presigned_url` string, uri, nullable — Time-limited HTTPS URL for uploading directly to S3. **Step 1 - Upload to S3:** curl -X PUT '{presigned_url}' -H 'Content-Type: {content_type}' --upload-file {filename} **Step 2 - REQUIRED: Confirm the upload:** POST /v1/uploads/{upload_id}/confirm (S3 has no callback - you MUST call confirm to finalize) The URL includes authentication and expires after presigned_url_expiration seconds. S3 returns an ETag header on success - pass it to confirm for integrity validation. NOTE: This will be null if is_duplicate=true (duplicate found, no upload needed).
  - `presigned_url_expiration` integer, required — How long the presigned URL is valid, in seconds
  - `s3_key` string, required — Full S3 object key where the file will be stored. Format: {internal_id}/{namespace_id}/api_buckets_uploads_create/{upload_id}/{filename}. Used internally for verification and object creation.
  - `status` 'PENDING' | 'QUEUED' | 'IN_PROGRESS' | 'PROCESSING' | 'COMPLETED' | 'COMPLETED_WITH_ERRORS' | 'FAILED' | 'CANCELED' | 'INTERRUPTED' | 'UNKNOWN' | 'SKIPPED' | 'DRAFT' | 'ACTIVE' | 'ARCHIVED' | 'SUSPENDED' | 'DEACTIVATED', required — Enumeration of task statuses for tracking asynchronous operations. Task statuses indicate the current state of asynchronous operations like batch processing, object ingestion, clustering, and taxonomy execution. Status Categories: Operation Statuses: Track progress of async operations Lifecycle Statuses: Track entity state (buckets, collections, namespaces) Values: PENDING: Task is queued but has not started processing yet IN_PROGRESS: Task is currently being executed PROCESSING: Task is actively processing data (similar to IN_PROGRESS) COMPLETED: Task finished successfully with no errors COMPLETED_WITH_ERRORS: Task finished but some items failed (partial success) FAILED: Task encountered an error and could not complete CANCELED: Task was manually canceled by a user or system UNKNOWN: Task status could not be determined SKIPPED: Task was intentionally skipped DRAFT: Task is in draft state and not yet submitted ACTIVE: Entity is active and operational (for buckets, collections, etc.) ARCHIVED: Entity has been archived SUSPENDED: Entity has been temporarily suspended Terminal Statuses: COMPLETED, COMPLETED_WITH_ERRORS, FAILED, CANCELED are terminal statuses. Once a task reaches these states, it will not transition to another state. Partial Success Handling: COMPLETED_WITH_ERRORS indicates that the operation completed but some documents/items failed. The task result includes: - List of successful items - List of failed items with error details - Success rate percentage This allows clients to handle partial success scenarios appropriately. Polling Guidance: - Poll tasks in PENDING, QUEUED, IN_PROGRESS, or PROCESSING states - Stop polling when task reaches COMPLETED, COMPLETED_WITH_ERRORS, FAILED, or CANCELED - Use exponential backoff (1s → 30s) when polling
  - `metadata` object — Custom metadata for tracking
  - `create_object_on_confirm` boolean, required — Whether bucket object will be auto-created on confirmation
  - `object_metadata` object, nullable — Metadata for the bucket object (if create_object_on_confirm=true)
  - `blob_property` string, nullable — Property name for the blob in bucket object
  - `blob_type` string, nullable — Type of blob (IMAGE, VIDEO, etc.)
  - `file_hash` string, nullable — SHA256 hash of the file content. Set during confirmation from S3 metadata or provided in request. Used for duplicate detection.
  - `skip_duplicates` boolean — Whether duplicate detection was enabled for this upload
  - `is_duplicate` boolean — Whether this upload was identified as a duplicate of an existing file. If true: - duplicate_of_upload_id contains the original upload - presigned_url will be null (no upload needed) - You can use the original upload's S3 object. This saves bandwidth and storage costs.
  - `duplicate_of_upload_id` string, nullable — If skip_duplicates=true and duplicate found, this is the original upload_id. The response will reference the existing upload instead of creating a new one.
  - `skipped_unique_key` boolean — Whether this upload was skipped because the unique key already exists in the bucket. If true: - existing_object_id contains the ID of the existing object - presigned_url will be null (no upload needed) - No S3 upload is required This saves bandwidth and prevents duplicate objects.
  - `existing_object_id` string, nullable — If skipped_unique_key=true, this is the object_id of the existing object that has the same unique key values. The upload was skipped to prevent duplicates.
  - `message` string, nullable — Human-readable message about the upload. Provided when is_duplicate=true or other special conditions. Example: 'File already exists with the same content hash. No upload needed - returning existing upload.'
  - `created_at` string, date-time, required — When this upload record was created (ISO 8601 format)
  - `expires_at` string, date-time, required — When the presigned URL expires (ISO 8601 format). After this time: - The presigned URL cannot be used - Upload status will be marked as FAILED if not completed - The upload record will be auto-deleted 30 days later (MongoDB TTL)
  - `completed_at` string, date-time, nullable — When the upload was completed and verified (ISO 8601 format)
  - `verified_at` string, date-time, nullable — When S3 object existence was verified (ISO 8601 format)
  - `etag` string, nullable — S3 ETag from the uploaded object (set during confirmation)
  - `object_id` string, nullable — Created bucket object ID (if create_object_on_confirm was true)
  - `task_id` string, nullable — Task ID for async confirmation (if processed asynchronously)

## Other responses

- `400` — Bad Request
- `401` — Unauthorized
- `403` — Forbidden
- `404` — Not Found
- `422` — Validation Error
- `500` — Internal Server Error

---

[API](https://skmtc.net/mixpeek/apis/mixpeek-api.md) · [All operations](https://skmtc.net/mixpeek/apis/mixpeek-api/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/mixpeek/mixpeek-api/versions/220a3b263fda/schema)
