v1

latestOpenAPI 3.1.02026-07-26106201349.6 KB
Gateway

Get Deployment

get/v1/accounts/{account_id}/deployments/{deployment_id}

Path parameters

account_idstring required

The Account Id

deployment_idstring required

The Deployment Id

Query parameters

readMaskstring

The fields to be returned in the response. If empty or "*", all fields will be returned.

Response

A successful response.

namestring
displayNamestring

Human-readable display name of the deployment. e.g. "My Deployment" Must be fewer than 64 characters long.

descriptionstring

Description of the deployment.

createTimestring date-time

The creation time of the deployment.

expireTimestring date-time

Deprecated: This field is deprecated and no longer causes auto-deletion. The time at which this deployment will automatically be deleted.

purgeTimestring date-time

The time at which the resource will be hard deleted.

deleteTimestring date-time

The time at which the resource will be soft deleted.

state'STATE_UNSPECIFIED' | 'CREATING' | 'READY' | 'DELETING' | 'FAILED' | 'UPDATING' | 'DELETED'
  • CREATING: The deployment is still being created.
  • READY: The deployment is ready to be used.
  • DELETING: The deployment is being deleted.
  • FAILED: The deployment failed to be created. See the status field for additional details on why it failed.
  • UPDATING: There are in-progress updates happening with the deployment.
  • DELETED: The deployment is soft-deleted.
annotationsobject

Annotations to identify deployment properties. Key/value pairs may be used by external tools or other services. The "image-tag-reason" key is redacted from API responses for non-superuser principals.

minReplicaCountinteger

The minimum number of replicas. If not specified, the default is 0.

maxReplicaCountinteger

The maximum number of replicas. If not specified, the default is max(min_replica_count, 1). May be set to 0 to downscale the deployment to 0.

maxWithRevocableReplicaCountinteger

max_with_revocable_replica_count is max replica count including revocable capacity. The max revocable capacity will be max_with_revocable_replica_count - max_replica_count.

desiredReplicaCountinteger

The desired number of replicas for this deployment. This represents the target replica count that the system is trying to achieve.

replicaCountinteger
baseModelstring required
acceleratorCountinteger

The number of accelerators used per replica. If not specified, the default is the estimated minimum required by the base model.

acceleratorType'ACCELERATOR_TYPE_UNSPECIFIED' | 'NVIDIA_A100_80GB' | 'NVIDIA_H100_80GB' | 'AMD_MI300X_192GB' | 'NVIDIA_A10G_24GB' | 'NVIDIA_A100_40GB' | 'NVIDIA_L4_24GB' | 'NVIDIA_H200_141GB' | 'NVIDIA_B200_180GB' | 'AMD_MI325X_256GB' | 'AMD_MI350X_288GB' | 'NVIDIA_B300_288GB' | 'NVIDIA_GB200' | 'NVIDIA_GB300'
precision'PRECISION_UNSPECIFIED' | 'FP16' | 'FP8' | 'FP8_MM' | 'FP8_AR' | 'FP8_MM_KV_ATTN' | 'FP8_KV' | 'FP8_MM_V2' | 'FP8_V2' | 'FP8_MM_KV_ATTN_V2' | 'NF4' | 'FP4' | 'BF16' | 'FP4_BLOCKSCALED_MM' | 'FP4_MX_MOE'
maxConcurrencyPerReplicainteger

The maximum number of concurrent (in-flight) requests a single replica will accept before shedding load. Requests that arrive while a replica is already at this limit are rejected early with HTTP 429 instead of queueing — a per-replica admission gate for controlling tail latency. When unset (0), the platform default is used.

clusterstring

If set, this deployment is deployed to a cloud-premise cluster.

enableAddonsboolean

If true, PEFT addons are enabled for this deployment.

draftTokenCountinteger

The number of candidate tokens to generate per step for speculative decoding. Default is the base model's draft_token_count. Set CreateDeploymentRequest.disable_speculative_decoding to false to disable this behavior.

draftModelstring

The draft model name for speculative decoding. e.g. accounts/fireworks/models/my-draft-model If empty, speculative decoding using a draft model is disabled. Default is the base model's default_draft_model. Set CreateDeploymentRequest.disable_speculative_decoding to false to disable this behavior.

ngramSpeculationLengthinteger

The length of previous input sequence to be considered for N-gram speculation.

enableSessionAffinityboolean

Whether to apply sticky routing based on user field. Serverless will be set to true when creating deployment.

directRouteApiKeysstring[]

The set of API keys used to access the direct route deployment. If direct routing is not enabled, this field is unused.

numPeftDeviceCachedinteger
directRouteType'DIRECT_ROUTE_TYPE_UNSPECIFIED' | 'INTERNET' | 'GCP_PRIVATE_SERVICE_CONNECT' | 'AWS_PRIVATELINK'
directRouteHandlestring

The handle for calling a direct route. The meaning of the handle depends on the direct route type of the deployment: INTERNET -> The host name for accessing the deployment GCP_PRIVATE_SERVICE_CONNECT -> The service attachment name used to create the PSC endpoint. AWS_PRIVATELINK -> The service name used to create the VPC endpoint.

deploymentTemplatestring

The name of the deployment template to use for this deployment. Only available to enterprise accounts.

region'REGION_UNSPECIFIED' | 'US_IOWA_1' | 'US_VIRGINIA_1' | 'US_VIRGINIA_2' | 'US_ILLINOIS_1' | 'AP_TOKYO_1' | 'US_ARIZONA_1' | 'US_TEXAS_1' | 'US_ILLINOIS_2' | 'EU_FRANKFURT_1' | 'US_TEXAS_2' | 'EU_ICELAND_1' | 'EU_ICELAND_2' | 'US_WASHINGTON_1' | 'US_WASHINGTON_2' | 'US_WASHINGTON_3' | 'AP_TOKYO_2' | 'US_CALIFORNIA_1' | 'US_UTAH_1' | 'US_ARIZONA_3' | 'US_GEORGIA_1' | 'US_GEORGIA_2' | 'US_WASHINGTON_4' | 'US_GEORGIA_3' | 'NA_BRITISHCOLUMBIA_1' | 'US_GEORGIA_4' | 'US_OHIO_1' | 'US_NEWYORK_1' | 'EU_NETHERLANDS_1' | 'US_WASHINGTON_5' | 'US_MINNESOTA_1' | 'US_CALIFORNIA_2' | 'NA_BRITISHCOLUMBIA_2' | 'AP_MALAYSIA_2' | 'US_OREGON_1' | 'NA_BRITISHCOLUMBIA_3' | 'AP_NEWSOUTHWALES_1'
maxContextLengthinteger

The maximum context length supported by the model (context window). If set to 0 or not specified, the model's default maximum context length will be used.

updateTimestring date-time

The update time for the deployment.

disableDeploymentSizeValidationboolean

Whether the deployment size validation is disabled.

enableHotLoadboolean

Whether to use hot load for this deployment.

hotLoadBucketType'BUCKET_TYPE_UNSPECIFIED' | 'MINIO' | 'S3' | 'NEBIUS' | 'FW_HOSTED'
enableHotReloadLatestAddonboolean

Allows up to 1 addon at a time to be loaded, and will merge it into the base model.

deploymentShapestring

The name of the deployment shape that this deployment is using. On the server side, this will be replaced with the deployment shape version name.

activeModelVersionstring

The model version that is currently active and applied to running replicas of a deployment.

targetModelVersionstring

The target model version that is being rolled out to the deployment. In a ready steady state, the target model version is the same as the active model version.

hotLoadBucketUrlstring
pricingPlanIdstring

Optional pricing plan ID for custom billing configuration. If set, this deployment will use the pricing plan's billing rules instead of default billing behavior.

hotLoadTrainerJobstring
hotLoadTransitionType'HOT_LOAD_TRANSITION_TYPE_UNSPECIFIED' | 'ASYNC' | 'SYNC'
preemptibleboolean

When true, this deployment runs as preemptible.