v1
latestOpenAPI 3.1.02026-07-26106201349.6 KBCreate Batch Inference Job
Path parameters
The Account Id
Query parameters
ID of the batch inference job.
Request body
The creation time of the batch inference job.
The time when the batch inference job will expire (stop running); any completed requests will have been written to the output dataset by then.
This is the job's effective execution deadline, derived by the server as create_time + the bounded run window (see max_job_duration). It is exposed so customers can read back the concrete deadline without recomputing it client-side. OUTPUT_ONLY: it is always computed server-side and any client-supplied value is ignored (previously this was a SUPERUSER_ONLY input that overlapped with max_job_duration; the two are now unified as one public input (duration) + one public derived deadline (this timestamp)).
The email address of the user who initiated this batch inference job.
JobState represents the state an asynchronous job can be in.
- JOB_STATE_PAUSED: Job is paused, typically due to account suspension or manual intervention.
- JOB_STATE_DELETED: Job has been deleted.
- JOB_STATE_ARCHIVED: User-facing state for jobs whose row is retained post-delete (e.g. RLOR trainers within the checkpoint retention window). The internal row is still in JOB_STATE_DELETED; the gateway translates it to ARCHIVED on public responses.
The name of the model to use for inference. This is required, except when continued_from_job_name is specified.
The name of the dataset used for inference. This is required, except when continued_from_job_name is specified.
The name of the dataset used for storing the results. This will also contain the error file.
Optional job-level system prompt. When set, it is injected as a leading system message into every input row that does NOT already begin with a system message (a row's own leading system message takes precedence). This lets callers avoid repeating a large, static system prompt on every row of the input dataset, shrinking the upload. Because the injected prefix is byte-identical across rows, prompt caching still applies.
The update time for the batch inference job.
The resource name of the batch inference job that this job continues from. Used for lineage tracking to understand job continuation chains.
The customer-requested wall-clock run window for the job: how long it may run before it is expired. This is the single public input that controls the job's lifetime. The server bounds it to [12h, 72h]; if unset it defaults to 24h. The resulting concrete deadline is surfaced as expire_time (= create_time + the bounded window) and the job is expired once that deadline passes. A duration (relative) is used rather than an absolute timestamp because the client does not know create_time at submit time. Customer-visible input.
True only while a job that has ALREADY started running is briefly re-acquiring capacity after a mid-run preemption/stockout (i.e. it regressed from RUNNING back to an internal PENDING/CREATING phase and is waiting to resume). This is intentionally a transient sub-status annotation on a job whose customer-facing state stays RUNNING — NOT a distinct state value: the job is still running (progress is saved, it auto-resumes) so introducing a new terminal-or-not enum value would force every state consumer (SDK/CLI/internal maps/billing) to special-case "still running." It drives the customer-facing "Briefly paused — waiting on capacity" card. It must NOT be set during first-time provisioning before the job has ever run, and is cleared once the job returns to RUNNING or reaches a terminal state. So "state=RUNNING, waiting_on_capacity=true" means: running, momentarily paused while it re-acquires capacity.
Response
A successful response.
The creation time of the batch inference job.
The time when the batch inference job will expire (stop running); any completed requests will have been written to the output dataset by then.
This is the job's effective execution deadline, derived by the server as create_time + the bounded run window (see max_job_duration). It is exposed so customers can read back the concrete deadline without recomputing it client-side. OUTPUT_ONLY: it is always computed server-side and any client-supplied value is ignored (previously this was a SUPERUSER_ONLY input that overlapped with max_job_duration; the two are now unified as one public input (duration) + one public derived deadline (this timestamp)).
The email address of the user who initiated this batch inference job.
JobState represents the state an asynchronous job can be in.
- JOB_STATE_PAUSED: Job is paused, typically due to account suspension or manual intervention.
- JOB_STATE_DELETED: Job has been deleted.
- JOB_STATE_ARCHIVED: User-facing state for jobs whose row is retained post-delete (e.g. RLOR trainers within the checkpoint retention window). The internal row is still in JOB_STATE_DELETED; the gateway translates it to ARCHIVED on public responses.
The name of the model to use for inference. This is required, except when continued_from_job_name is specified.
The name of the dataset used for inference. This is required, except when continued_from_job_name is specified.
The name of the dataset used for storing the results. This will also contain the error file.
Optional job-level system prompt. When set, it is injected as a leading system message into every input row that does NOT already begin with a system message (a row's own leading system message takes precedence). This lets callers avoid repeating a large, static system prompt on every row of the input dataset, shrinking the upload. Because the injected prefix is byte-identical across rows, prompt caching still applies.
The update time for the batch inference job.
The resource name of the batch inference job that this job continues from. Used for lineage tracking to understand job continuation chains.
The customer-requested wall-clock run window for the job: how long it may run before it is expired. This is the single public input that controls the job's lifetime. The server bounds it to [12h, 72h]; if unset it defaults to 24h. The resulting concrete deadline is surfaced as expire_time (= create_time + the bounded window) and the job is expired once that deadline passes. A duration (relative) is used rather than an absolute timestamp because the client does not know create_time at submit time. Customer-visible input.
True only while a job that has ALREADY started running is briefly re-acquiring capacity after a mid-run preemption/stockout (i.e. it regressed from RUNNING back to an internal PENDING/CREATING phase and is waiting to resume). This is intentionally a transient sub-status annotation on a job whose customer-facing state stays RUNNING — NOT a distinct state value: the job is still running (progress is saved, it auto-resumes) so introducing a new terminal-or-not enum value would force every state consumer (SDK/CLI/internal maps/billing) to special-case "still running." It drives the customer-facing "Briefly paused — waiting on capacity" card. It must NOT be set during first-time provisioning before the job has ever run, and is cleared once the job returns to RUNNING or reaches a terminal state. So "state=RUNNING, waiting_on_capacity=true" means: running, momentarily paused while it re-acquires capacity.