---
title: "List Batch Inference Jobs"
method: GET
path: "/v1/accounts/{account_id}/batchInferenceJobs"
tags: ["Gateway"]
---

# List Batch Inference Jobs

`GET /v1/accounts/{account_id}/batchInferenceJobs`

## Path parameters

- `account_id` string, required

## Query parameters

- `pageSize` integer
- `pageToken` string
- `filter` string
- `orderBy` string
- `readMask` string

## Response `200`

A successful response.

- GatewayListBatchInferenceJobsResponse
  - `batchInferenceJobs` GatewayBatchInferenceJob[]
    - `name` string
    - `displayName` string
    - `createTime` string, date-time — The creation time of the batch inference job.
    - `expireTime` string, date-time — The time when the batch inference job will expire (stop running); any completed requests will have been written to the output dataset by then. This is the job's effective execution deadline, derived by the server as create_time + the bounded run window (see max_job_duration). It is exposed so customers can read back the concrete deadline without recomputing it client-side. OUTPUT_ONLY: it is always computed server-side and any client-supplied value is ignored (previously this was a SUPERUSER_ONLY input that overlapped with max_job_duration; the two are now unified as one public input (duration) + one public derived deadline (this timestamp)).
    - `createdBy` string — The email address of the user who initiated this batch inference job.
    - `state` 'JOB_STATE_UNSPECIFIED' | 'JOB_STATE_CREATING' | 'JOB_STATE_RUNNING' | 'JOB_STATE_COMPLETED' | 'JOB_STATE_FAILED' | 'JOB_STATE_CANCELLED' | 'JOB_STATE_DELETING' | 'JOB_STATE_WRITING_RESULTS' | 'JOB_STATE_VALIDATING' | 'JOB_STATE_DELETING_CLEANING_UP' | 'JOB_STATE_PENDING' | 'JOB_STATE_EXPIRED' | 'JOB_STATE_RE_QUEUEING' | 'JOB_STATE_CREATING_INPUT_DATASET' | 'JOB_STATE_IDLE' | 'JOB_STATE_CANCELLING' | 'JOB_STATE_EARLY_STOPPED' | 'JOB_STATE_PAUSED' | 'JOB_STATE_DELETED' | 'JOB_STATE_ARCHIVED' — JobState represents the state an asynchronous job can be in. - JOB_STATE_PAUSED: Job is paused, typically due to account suspension or manual intervention. - JOB_STATE_DELETED: Job has been deleted. - JOB_STATE_ARCHIVED: User-facing state for jobs whose row is retained post-delete (e.g. RLOR trainers within the checkpoint retention window). The internal row is still in JOB_STATE_DELETED; the gateway translates it to ARCHIVED on public responses.
    - `status` GatewayStatus
      - `code` 'OK' | 'CANCELLED' | 'UNKNOWN' | 'INVALID_ARGUMENT' | 'DEADLINE_EXCEEDED' | 'NOT_FOUND' | 'ALREADY_EXISTS' | 'PERMISSION_DENIED' | 'UNAUTHENTICATED' | 'RESOURCE_EXHAUSTED' | 'FAILED_PRECONDITION' | 'ABORTED' | 'OUT_OF_RANGE' | 'UNIMPLEMENTED' | 'INTERNAL' | 'UNAVAILABLE' | 'DATA_LOSS' — - OK: Not an error; returned on success. HTTP Mapping: 200 OK - CANCELLED: The operation was cancelled, typically by the caller. HTTP Mapping: 499 Client Closed Request - UNKNOWN: Unknown error. For example, this error may be returned when a `Status` value received from another address space belongs to an error space that is not known in this address space. Also errors raised by APIs that do not return enough error information may be converted to this error. HTTP Mapping: 500 Internal Server Error - INVALID_ARGUMENT: The client specified an invalid argument. Note that this differs from `FAILED_PRECONDITION`. `INVALID_ARGUMENT` indicates arguments that are problematic regardless of the state of the system (e.g., a malformed file name). HTTP Mapping: 400 Bad Request - DEADLINE_EXCEEDED: The deadline expired before the operation could complete. For operations that change the state of the system, this error may be returned even if the operation has completed successfully. For example, a successful response from a server could have been delayed long enough for the deadline to expire. HTTP Mapping: 504 Gateway Timeout - NOT_FOUND: Some requested entity (e.g., file or directory) was not found. Note to server developers: if a request is denied for an entire class of users, such as gradual feature rollout or undocumented allowlist, `NOT_FOUND` may be used. If a request is denied for some users within a class of users, such as user-based access control, `PERMISSION_DENIED` must be used. HTTP Mapping: 404 Not Found - ALREADY_EXISTS: The entity that a client attempted to create (e.g., file or directory) already exists. HTTP Mapping: 409 Conflict - PERMISSION_DENIED: The caller does not have permission to execute the specified operation. `PERMISSION_DENIED` must not be used for rejections caused by exhausting some resource (use `RESOURCE_EXHAUSTED` instead for those errors). `PERMISSION_DENIED` must not be used if the caller can not be identified (use `UNAUTHENTICATED` instead for those errors). This error code does not imply the request is valid or the requested entity exists or satisfies other pre-conditions. HTTP Mapping: 403 Forbidden - UNAUTHENTICATED: The request does not have valid authentication credentials for the operation. HTTP Mapping: 401 Unauthorized - RESOURCE_EXHAUSTED: Some resource has been exhausted, perhaps a per-user quota, or perhaps the entire file system is out of space. HTTP Mapping: 429 Too Many Requests - FAILED_PRECONDITION: The operation was rejected because the system is not in a state required for the operation's execution. For example, the directory to be deleted is non-empty, an rmdir operation is applied to a non-directory, etc. Service implementors can use the following guidelines to decide between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`: (a) Use `UNAVAILABLE` if the client can retry just the failing call. (b) Use `ABORTED` if the client should retry at a higher level. For example, when a client-specified test-and-set fails, indicating the client should restart a read-modify-write sequence. (c) Use `FAILED_PRECONDITION` if the client should not retry until the system state has been explicitly fixed. For example, if an "rmdir" fails because the directory is non-empty, `FAILED_PRECONDITION` should be returned since the client should not retry unless the files are deleted from the directory. HTTP Mapping: 400 Bad Request - ABORTED: The operation was aborted, typically due to a concurrency issue such as a sequencer check failure or transaction abort. See the guidelines above for deciding between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`. HTTP Mapping: 409 Conflict - OUT_OF_RANGE: The operation was attempted past the valid range. E.g., seeking or reading past end-of-file. Unlike `INVALID_ARGUMENT`, this error indicates a problem that may be fixed if the system state changes. For example, a 32-bit file system will generate `INVALID_ARGUMENT` if asked to read at an offset that is not in the range [0,2^32-1], but it will generate `OUT_OF_RANGE` if asked to read from an offset past the current file size. There is a fair bit of overlap between `FAILED_PRECONDITION` and `OUT_OF_RANGE`. We recommend using `OUT_OF_RANGE` (the more specific error) when it applies so that callers who are iterating through a space can easily look for an `OUT_OF_RANGE` error to detect when they are done. HTTP Mapping: 400 Bad Request - UNIMPLEMENTED: The operation is not implemented or is not supported/enabled in this service. HTTP Mapping: 501 Not Implemented - INTERNAL: Internal errors. This means that some invariants expected by the underlying system have been broken. This error code is reserved for serious errors. HTTP Mapping: 500 Internal Server Error - UNAVAILABLE: The service is currently unavailable. This is most likely a transient condition, which can be corrected by retrying with a backoff. Note that it is not always safe to retry non-idempotent operations. See the guidelines above for deciding between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`. HTTP Mapping: 503 Service Unavailable - DATA_LOSS: Unrecoverable data loss or corruption. HTTP Mapping: 500 Internal Server Error
      - `message` string — A developer-facing error message in English.
    - `model` string — The name of the model to use for inference. This is required, except when continued_from_job_name is specified.
    - `inputDatasetId` string — The name of the dataset used for inference. This is required, except when continued_from_job_name is specified.
    - `outputDatasetId` string — The name of the dataset used for storing the results. This will also contain the error file.
    - `systemPrompt` string — Optional job-level system prompt. When set, it is injected as a leading system message into every input row that does NOT already begin with a system message (a row's own leading system message takes precedence). This lets callers avoid repeating a large, static system prompt on every row of the input dataset, shrinking the upload. Because the injected prefix is byte-identical across rows, prompt caching still applies.
    - `inferenceParameters` GatewayBatchInferenceJobInferenceParameters
      - `maxTokens` integer — Maximum number of tokens to generate per response.
      - `temperature` number, float — Sampling temperature, typically between 0 and 2.
      - `topP` number, float — Top-p sampling parameter, typically between 0 and 1.
      - `n` integer — Number of response candidates to generate per input.
      - `extraBody` string — Additional parameters for the inference request as a JSON string. For example: "{\"stop\": [\"\\n\"]}".
      - `topK` integer — Top-k sampling parameter, limits the token selection to the top k tokens.
    - `updateTime` string, date-time — The update time for the batch inference job.
    - `precision` 'PRECISION_UNSPECIFIED' | 'FP16' | 'FP8' | 'FP8_MM' | 'FP8_AR' | 'FP8_MM_KV_ATTN' | 'FP8_KV' | 'FP8_MM_V2' | 'FP8_V2' | 'FP8_MM_KV_ATTN_V2' | 'NF4' | 'FP4' | 'BF16' | 'FP4_BLOCKSCALED_MM' | 'FP4_MX_MOE'
    - `jobProgress` GatewayJobProgress — Progress of a job, e.g. RLOR, EVJ, BIJ etc.
      - `percent` integer — Progress percent, within the range from 0 to 100.
      - `epoch` integer — The epoch for which the progress percent is reported, usually starting from 0. This is optional for jobs that don't run in an epoch fasion, e.g. BIJ, EVJ.
      - `totalInputRequests` integer — Total number of input requests/rows in the job.
      - `totalProcessedRequests` integer — Total number of requests that have been processed (successfully or failed).
      - `successfullyProcessedRequests` integer — Number of requests that were processed successfully.
      - `failedRequests` integer — Number of requests that failed to process.
      - `outputRows` integer — Number of output rows generated.
      - `inputTokens` integer — Total number of input tokens processed.
      - `outputTokens` integer — Total number of output tokens generated.
      - `cachedInputTokenCount` integer — The number of input tokens that hit the prompt cache.
    - `continuedFromJobName` string — The resource name of the batch inference job that this job continues from. Used for lineage tracking to understand job continuation chains.
    - `placement` GatewayPlacement — The desired geographic region where the deployment must be placed. Exactly one field will be specified.
      - `region` 'REGION_UNSPECIFIED' | 'US_IOWA_1' | 'US_VIRGINIA_1' | 'US_VIRGINIA_2' | 'US_ILLINOIS_1' | 'AP_TOKYO_1' | 'US_ARIZONA_1' | 'US_TEXAS_1' | 'US_ILLINOIS_2' | 'EU_FRANKFURT_1' | 'US_TEXAS_2' | 'EU_ICELAND_1' | 'EU_ICELAND_2' | 'US_WASHINGTON_1' | 'US_WASHINGTON_2' | 'US_WASHINGTON_3' | 'AP_TOKYO_2' | 'US_CALIFORNIA_1' | 'US_UTAH_1' | 'US_ARIZONA_3' | 'US_GEORGIA_1' | 'US_GEORGIA_2' | 'US_WASHINGTON_4' | 'US_GEORGIA_3' | 'NA_BRITISHCOLUMBIA_1' | 'US_GEORGIA_4' | 'US_OHIO_1' | 'US_NEWYORK_1' | 'EU_NETHERLANDS_1' | 'US_WASHINGTON_5' | 'US_MINNESOTA_1' | 'US_CALIFORNIA_2' | 'NA_BRITISHCOLUMBIA_2' | 'AP_MALAYSIA_2' | 'US_OREGON_1' | 'NA_BRITISHCOLUMBIA_3' | 'AP_NEWSOUTHWALES_1'
      - `multiRegion` 'MULTI_REGION_UNSPECIFIED' | 'GLOBAL' | 'US' | 'EUROPE' | 'APAC'
      - `regions` GatewayRegion[]
    - `maxJobDuration` string — The customer-requested wall-clock run window for the job: how long it may run before it is expired. This is the single public input that controls the job's lifetime. The server bounds it to [12h, 72h]; if unset it defaults to 24h. The resulting concrete deadline is surfaced as expire_time (= create_time + the bounded window) and the job is expired once that deadline passes. A duration (relative) is used rather than an absolute timestamp because the client does not know create_time at submit time. Customer-visible input.
    - `lifecycle` BatchInferenceJobLifecycleTimestamps — Observed execution milestone timestamps for the job, grouped into one sub-message so the job's lifecycle reads as a single cohesive unit rather than a scatter of top-level time fields. All OUTPUT_ONLY.
      - `validatedTime` string, date-time — When dataset validation completed and the job left VALIDATING. Unset for jobs that skip validation or that fail before validating.
      - `runStartTime` string, date-time — When the runner first started processing (the job first became RUNNING).
      - `endTime` string, date-time — The terminal time of the job (when it reached COMPLETED, FAILED, EXPIRED, or CANCELLED).
    - `waitingOnCapacity` boolean — True only while a job that has ALREADY started running is briefly re-acquiring capacity after a mid-run preemption/stockout (i.e. it regressed from RUNNING back to an internal PENDING/CREATING phase and is waiting to resume). This is intentionally a transient sub-status annotation on a job whose customer-facing `state` stays RUNNING — NOT a distinct `state` value: the job is still running (progress is saved, it auto-resumes) so introducing a new terminal-or-not enum value would force every state consumer (SDK/CLI/internal maps/billing) to special-case "still running." It drives the customer-facing "Briefly paused — waiting on capacity" card. It must NOT be set during first-time provisioning before the job has ever run, and is cleared once the job returns to RUNNING or reaches a terminal state. So "state=RUNNING, waiting_on_capacity=true" means: running, momentarily paused while it re-acquires capacity.
  - `nextPageToken` string — A token, which can be sent as `page_token` to retrieve the next page. If this field is omitted, there are no subsequent pages.
  - `totalSize` integer — The total number of batch inference jobs.

---

[API](https://skmtc.net/fireworks/apis/fireworks-ai-anthropic-compatible-messages-api.md) · [All operations](https://skmtc.net/fireworks/apis/fireworks-ai-anthropic-compatible-messages-api/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/fireworks/fireworks-ai-anthropic-compatible-messages-api/versions/954d6bc5d922/schema)
