---
title: "Update Deployment"
method: PATCH
path: "/v1/accounts/{account_id}/deployments/{deployment_id}"
tags: ["Gateway"]
---

# Update Deployment

`PATCH /v1/accounts/{account_id}/deployments/{deployment_id}`

## Path parameters

- `account_id` string, required
- `deployment_id` string, required

## Query parameters

- `skipShapeValidation` boolean

## Request body

- object
  - `displayName` string — Human-readable display name of the deployment. e.g. "My Deployment" Must be fewer than 64 characters long.
  - `description` string — Description of the deployment.
  - `createTime` string, date-time — The creation time of the deployment.
  - `expireTime` string, date-time — Deprecated: This field is deprecated and no longer causes auto-deletion. The time at which this deployment will automatically be deleted.
  - `purgeTime` string, date-time — The time at which the resource will be hard deleted.
  - `deleteTime` string, date-time — The time at which the resource will be soft deleted.
  - `state` 'STATE_UNSPECIFIED' | 'CREATING' | 'READY' | 'DELETING' | 'FAILED' | 'UPDATING' | 'DELETED' — - CREATING: The deployment is still being created. - READY: The deployment is ready to be used. - DELETING: The deployment is being deleted. - FAILED: The deployment failed to be created. See the `status` field for additional details on why it failed. - UPDATING: There are in-progress updates happening with the deployment. - DELETED: The deployment is soft-deleted.
  - `status` GatewayStatus
    - `code` 'OK' | 'CANCELLED' | 'UNKNOWN' | 'INVALID_ARGUMENT' | 'DEADLINE_EXCEEDED' | 'NOT_FOUND' | 'ALREADY_EXISTS' | 'PERMISSION_DENIED' | 'UNAUTHENTICATED' | 'RESOURCE_EXHAUSTED' | 'FAILED_PRECONDITION' | 'ABORTED' | 'OUT_OF_RANGE' | 'UNIMPLEMENTED' | 'INTERNAL' | 'UNAVAILABLE' | 'DATA_LOSS' — - OK: Not an error; returned on success. HTTP Mapping: 200 OK - CANCELLED: The operation was cancelled, typically by the caller. HTTP Mapping: 499 Client Closed Request - UNKNOWN: Unknown error. For example, this error may be returned when a `Status` value received from another address space belongs to an error space that is not known in this address space. Also errors raised by APIs that do not return enough error information may be converted to this error. HTTP Mapping: 500 Internal Server Error - INVALID_ARGUMENT: The client specified an invalid argument. Note that this differs from `FAILED_PRECONDITION`. `INVALID_ARGUMENT` indicates arguments that are problematic regardless of the state of the system (e.g., a malformed file name). HTTP Mapping: 400 Bad Request - DEADLINE_EXCEEDED: The deadline expired before the operation could complete. For operations that change the state of the system, this error may be returned even if the operation has completed successfully. For example, a successful response from a server could have been delayed long enough for the deadline to expire. HTTP Mapping: 504 Gateway Timeout - NOT_FOUND: Some requested entity (e.g., file or directory) was not found. Note to server developers: if a request is denied for an entire class of users, such as gradual feature rollout or undocumented allowlist, `NOT_FOUND` may be used. If a request is denied for some users within a class of users, such as user-based access control, `PERMISSION_DENIED` must be used. HTTP Mapping: 404 Not Found - ALREADY_EXISTS: The entity that a client attempted to create (e.g., file or directory) already exists. HTTP Mapping: 409 Conflict - PERMISSION_DENIED: The caller does not have permission to execute the specified operation. `PERMISSION_DENIED` must not be used for rejections caused by exhausting some resource (use `RESOURCE_EXHAUSTED` instead for those errors). `PERMISSION_DENIED` must not be used if the caller can not be identified (use `UNAUTHENTICATED` instead for those errors). This error code does not imply the request is valid or the requested entity exists or satisfies other pre-conditions. HTTP Mapping: 403 Forbidden - UNAUTHENTICATED: The request does not have valid authentication credentials for the operation. HTTP Mapping: 401 Unauthorized - RESOURCE_EXHAUSTED: Some resource has been exhausted, perhaps a per-user quota, or perhaps the entire file system is out of space. HTTP Mapping: 429 Too Many Requests - FAILED_PRECONDITION: The operation was rejected because the system is not in a state required for the operation's execution. For example, the directory to be deleted is non-empty, an rmdir operation is applied to a non-directory, etc. Service implementors can use the following guidelines to decide between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`: (a) Use `UNAVAILABLE` if the client can retry just the failing call. (b) Use `ABORTED` if the client should retry at a higher level. For example, when a client-specified test-and-set fails, indicating the client should restart a read-modify-write sequence. (c) Use `FAILED_PRECONDITION` if the client should not retry until the system state has been explicitly fixed. For example, if an "rmdir" fails because the directory is non-empty, `FAILED_PRECONDITION` should be returned since the client should not retry unless the files are deleted from the directory. HTTP Mapping: 400 Bad Request - ABORTED: The operation was aborted, typically due to a concurrency issue such as a sequencer check failure or transaction abort. See the guidelines above for deciding between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`. HTTP Mapping: 409 Conflict - OUT_OF_RANGE: The operation was attempted past the valid range. E.g., seeking or reading past end-of-file. Unlike `INVALID_ARGUMENT`, this error indicates a problem that may be fixed if the system state changes. For example, a 32-bit file system will generate `INVALID_ARGUMENT` if asked to read at an offset that is not in the range [0,2^32-1], but it will generate `OUT_OF_RANGE` if asked to read from an offset past the current file size. There is a fair bit of overlap between `FAILED_PRECONDITION` and `OUT_OF_RANGE`. We recommend using `OUT_OF_RANGE` (the more specific error) when it applies so that callers who are iterating through a space can easily look for an `OUT_OF_RANGE` error to detect when they are done. HTTP Mapping: 400 Bad Request - UNIMPLEMENTED: The operation is not implemented or is not supported/enabled in this service. HTTP Mapping: 501 Not Implemented - INTERNAL: Internal errors. This means that some invariants expected by the underlying system have been broken. This error code is reserved for serious errors. HTTP Mapping: 500 Internal Server Error - UNAVAILABLE: The service is currently unavailable. This is most likely a transient condition, which can be corrected by retrying with a backoff. Note that it is not always safe to retry non-idempotent operations. See the guidelines above for deciding between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`. HTTP Mapping: 503 Service Unavailable - DATA_LOSS: Unrecoverable data loss or corruption. HTTP Mapping: 500 Internal Server Error
    - `message` string — A developer-facing error message in English.
  - `annotations` object — Annotations to identify deployment properties. Key/value pairs may be used by external tools or other services. The "image-tag-reason" key is redacted from API responses for non-superuser principals.
  - `minReplicaCount` integer — The minimum number of replicas. If not specified, the default is 0.
  - `maxReplicaCount` integer — The maximum number of replicas. If not specified, the default is max(min_replica_count, 1). May be set to 0 to downscale the deployment to 0.
  - `maxWithRevocableReplicaCount` integer — max_with_revocable_replica_count is max replica count including revocable capacity. The max revocable capacity will be max_with_revocable_replica_count - max_replica_count.
  - `desiredReplicaCount` integer — The desired number of replicas for this deployment. This represents the target replica count that the system is trying to achieve.
  - `replicaCount` integer
  - `autoscalingPolicy` GatewayAutoscalingPolicy
    - `scaleUpWindow` string — The duration the autoscaler will wait before scaling up a deployment after observing increased load. Default is 30s. Must be less than or equal to 1 hour.
    - `scaleDownWindow` string — The duration the autoscaler will wait before scaling down a deployment after observing decreased load. Default is 10m. Must be less than or equal to 1 hour.
    - `scaleToZeroWindow` string — The duration after which there are no requests that the deployment will be scaled down to zero replicas, if min_replica_count==0. Default is 1h. This must be at least 5 minutes.
    - `loadTargets` object
    - `scalingSchedules` object — Named scaling schedules that override min_replica_count on a time-based cron schedule. When multiple schedules are active simultaneously, the highest min_replica_count across all active schedules is used ("max wins"). When no schedule is active, the deployment's base min_replica_count applies. Maximum 5 schedules per deployment.
  - `baseModel` string, required
  - `acceleratorCount` integer — The number of accelerators used per replica. If not specified, the default is the estimated minimum required by the base model.
  - `acceleratorType` 'ACCELERATOR_TYPE_UNSPECIFIED' | 'NVIDIA_A100_80GB' | 'NVIDIA_H100_80GB' | 'AMD_MI300X_192GB' | 'NVIDIA_A10G_24GB' | 'NVIDIA_A100_40GB' | 'NVIDIA_L4_24GB' | 'NVIDIA_H200_141GB' | 'NVIDIA_B200_180GB' | 'AMD_MI325X_256GB' | 'AMD_MI350X_288GB' | 'NVIDIA_B300_288GB' | 'NVIDIA_GB200' | 'NVIDIA_GB300'
  - `precision` 'PRECISION_UNSPECIFIED' | 'FP16' | 'FP8' | 'FP8_MM' | 'FP8_AR' | 'FP8_MM_KV_ATTN' | 'FP8_KV' | 'FP8_MM_V2' | 'FP8_V2' | 'FP8_MM_KV_ATTN_V2' | 'NF4' | 'FP4' | 'BF16' | 'FP4_BLOCKSCALED_MM' | 'FP4_MX_MOE'
  - `maxConcurrencyPerReplica` integer — The maximum number of concurrent (in-flight) requests a single replica will accept before shedding load. Requests that arrive while a replica is already at this limit are rejected early with HTTP 429 instead of queueing — a per-replica admission gate for controlling tail latency. When unset (0), the platform default is used.
  - `cluster` string — If set, this deployment is deployed to a cloud-premise cluster.
  - `enableAddons` boolean — If true, PEFT addons are enabled for this deployment.
  - `draftTokenCount` integer — The number of candidate tokens to generate per step for speculative decoding. Default is the base model's draft_token_count. Set CreateDeploymentRequest.disable_speculative_decoding to false to disable this behavior.
  - `draftModel` string — The draft model name for speculative decoding. e.g. accounts/fireworks/models/my-draft-model If empty, speculative decoding using a draft model is disabled. Default is the base model's default_draft_model. Set CreateDeploymentRequest.disable_speculative_decoding to false to disable this behavior.
  - `ngramSpeculationLength` integer — The length of previous input sequence to be considered for N-gram speculation.
  - `enableSessionAffinity` boolean — Whether to apply sticky routing based on `user` field. Serverless will be set to true when creating deployment.
  - `directRouteApiKeys` string[] — The set of API keys used to access the direct route deployment. If direct routing is not enabled, this field is unused.
  - `numPeftDeviceCached` integer
  - `directRouteType` 'DIRECT_ROUTE_TYPE_UNSPECIFIED' | 'INTERNET' | 'GCP_PRIVATE_SERVICE_CONNECT' | 'AWS_PRIVATELINK'
  - `directRouteHandle` string — The handle for calling a direct route. The meaning of the handle depends on the direct route type of the deployment: INTERNET -> The host name for accessing the deployment GCP_PRIVATE_SERVICE_CONNECT -> The service attachment name used to create the PSC endpoint. AWS_PRIVATELINK -> The service name used to create the VPC endpoint.
  - `deploymentTemplate` string — The name of the deployment template to use for this deployment. Only available to enterprise accounts.
  - `autoTune` GatewayAutoTune
    - `longPrompt` boolean — If true, this deployment is optimized for long prompt lengths.
  - `placement` GatewayPlacement — The desired geographic region where the deployment must be placed. Exactly one field will be specified.
    - `region` 'REGION_UNSPECIFIED' | 'US_IOWA_1' | 'US_VIRGINIA_1' | 'US_VIRGINIA_2' | 'US_ILLINOIS_1' | 'AP_TOKYO_1' | 'US_ARIZONA_1' | 'US_TEXAS_1' | 'US_ILLINOIS_2' | 'EU_FRANKFURT_1' | 'US_TEXAS_2' | 'EU_ICELAND_1' | 'EU_ICELAND_2' | 'US_WASHINGTON_1' | 'US_WASHINGTON_2' | 'US_WASHINGTON_3' | 'AP_TOKYO_2' | 'US_CALIFORNIA_1' | 'US_UTAH_1' | 'US_ARIZONA_3' | 'US_GEORGIA_1' | 'US_GEORGIA_2' | 'US_WASHINGTON_4' | 'US_GEORGIA_3' | 'NA_BRITISHCOLUMBIA_1' | 'US_GEORGIA_4' | 'US_OHIO_1' | 'US_NEWYORK_1' | 'EU_NETHERLANDS_1' | 'US_WASHINGTON_5' | 'US_MINNESOTA_1' | 'US_CALIFORNIA_2' | 'NA_BRITISHCOLUMBIA_2' | 'AP_MALAYSIA_2' | 'US_OREGON_1' | 'NA_BRITISHCOLUMBIA_3' | 'AP_NEWSOUTHWALES_1'
    - `multiRegion` 'MULTI_REGION_UNSPECIFIED' | 'GLOBAL' | 'US' | 'EUROPE' | 'APAC'
    - `regions` GatewayRegion[]
  - `region` 'REGION_UNSPECIFIED' | 'US_IOWA_1' | 'US_VIRGINIA_1' | 'US_VIRGINIA_2' | 'US_ILLINOIS_1' | 'AP_TOKYO_1' | 'US_ARIZONA_1' | 'US_TEXAS_1' | 'US_ILLINOIS_2' | 'EU_FRANKFURT_1' | 'US_TEXAS_2' | 'EU_ICELAND_1' | 'EU_ICELAND_2' | 'US_WASHINGTON_1' | 'US_WASHINGTON_2' | 'US_WASHINGTON_3' | 'AP_TOKYO_2' | 'US_CALIFORNIA_1' | 'US_UTAH_1' | 'US_ARIZONA_3' | 'US_GEORGIA_1' | 'US_GEORGIA_2' | 'US_WASHINGTON_4' | 'US_GEORGIA_3' | 'NA_BRITISHCOLUMBIA_1' | 'US_GEORGIA_4' | 'US_OHIO_1' | 'US_NEWYORK_1' | 'EU_NETHERLANDS_1' | 'US_WASHINGTON_5' | 'US_MINNESOTA_1' | 'US_CALIFORNIA_2' | 'NA_BRITISHCOLUMBIA_2' | 'AP_MALAYSIA_2' | 'US_OREGON_1' | 'NA_BRITISHCOLUMBIA_3' | 'AP_NEWSOUTHWALES_1'
  - `maxContextLength` integer — The maximum context length supported by the model (context window). If set to 0 or not specified, the model's default maximum context length will be used.
  - `updateTime` string, date-time — The update time for the deployment.
  - `disableDeploymentSizeValidation` boolean — Whether the deployment size validation is disabled.
  - `enableHotLoad` boolean — Whether to use hot load for this deployment.
  - `hotLoadBucketType` 'BUCKET_TYPE_UNSPECIFIED' | 'MINIO' | 'S3' | 'NEBIUS' | 'FW_HOSTED'
  - `enableHotReloadLatestAddon` boolean — Allows up to 1 addon at a time to be loaded, and will merge it into the base model.
  - `deploymentShape` string — The name of the deployment shape that this deployment is using. On the server side, this will be replaced with the deployment shape version name.
  - `activeModelVersion` string — The model version that is currently active and applied to running replicas of a deployment.
  - `targetModelVersion` string — The target model version that is being rolled out to the deployment. In a ready steady state, the target model version is the same as the active model version.
  - `replicaStats` GatewayReplicaStats
    - `pendingSchedulingReplicaCount` integer — Number of replicas waiting to be scheduled to a node.
    - `downloadingModelReplicaCount` integer — Number of replicas downloading model weights.
    - `initializingReplicaCount` integer — Number of replicas initializing the model server.
    - `readyReplicaCount` integer — Number of replicas that are ready and serving traffic.
    - `revocableReplicaCount` integer
    - `effectiveReplicaCount` number, float — The effective number of replicas currently serving traffic, including fractional progress toward a replica that is not yet fully ready. Equals ready_replica_count when no replica is partially ready.
  - `hotLoadBucketUrl` string
  - `pricingPlanId` string — Optional pricing plan ID for custom billing configuration. If set, this deployment will use the pricing plan's billing rules instead of default billing behavior.
  - `hotLoadTrainerJob` string
  - `hotLoadTransitionType` 'HOT_LOAD_TRANSITION_TYPE_UNSPECIFIED' | 'ASYNC' | 'SYNC'
  - `preemptible` boolean — When true, this deployment runs as preemptible.

## Response `200`

A successful response.

- GatewayDeployment
  - `name` string
  - `displayName` string — Human-readable display name of the deployment. e.g. "My Deployment" Must be fewer than 64 characters long.
  - `description` string — Description of the deployment.
  - `createTime` string, date-time — The creation time of the deployment.
  - `expireTime` string, date-time — Deprecated: This field is deprecated and no longer causes auto-deletion. The time at which this deployment will automatically be deleted.
  - `purgeTime` string, date-time — The time at which the resource will be hard deleted.
  - `deleteTime` string, date-time — The time at which the resource will be soft deleted.
  - `state` 'STATE_UNSPECIFIED' | 'CREATING' | 'READY' | 'DELETING' | 'FAILED' | 'UPDATING' | 'DELETED' — - CREATING: The deployment is still being created. - READY: The deployment is ready to be used. - DELETING: The deployment is being deleted. - FAILED: The deployment failed to be created. See the `status` field for additional details on why it failed. - UPDATING: There are in-progress updates happening with the deployment. - DELETED: The deployment is soft-deleted.
  - `status` GatewayStatus
    - `code` 'OK' | 'CANCELLED' | 'UNKNOWN' | 'INVALID_ARGUMENT' | 'DEADLINE_EXCEEDED' | 'NOT_FOUND' | 'ALREADY_EXISTS' | 'PERMISSION_DENIED' | 'UNAUTHENTICATED' | 'RESOURCE_EXHAUSTED' | 'FAILED_PRECONDITION' | 'ABORTED' | 'OUT_OF_RANGE' | 'UNIMPLEMENTED' | 'INTERNAL' | 'UNAVAILABLE' | 'DATA_LOSS' — - OK: Not an error; returned on success. HTTP Mapping: 200 OK - CANCELLED: The operation was cancelled, typically by the caller. HTTP Mapping: 499 Client Closed Request - UNKNOWN: Unknown error. For example, this error may be returned when a `Status` value received from another address space belongs to an error space that is not known in this address space. Also errors raised by APIs that do not return enough error information may be converted to this error. HTTP Mapping: 500 Internal Server Error - INVALID_ARGUMENT: The client specified an invalid argument. Note that this differs from `FAILED_PRECONDITION`. `INVALID_ARGUMENT` indicates arguments that are problematic regardless of the state of the system (e.g., a malformed file name). HTTP Mapping: 400 Bad Request - DEADLINE_EXCEEDED: The deadline expired before the operation could complete. For operations that change the state of the system, this error may be returned even if the operation has completed successfully. For example, a successful response from a server could have been delayed long enough for the deadline to expire. HTTP Mapping: 504 Gateway Timeout - NOT_FOUND: Some requested entity (e.g., file or directory) was not found. Note to server developers: if a request is denied for an entire class of users, such as gradual feature rollout or undocumented allowlist, `NOT_FOUND` may be used. If a request is denied for some users within a class of users, such as user-based access control, `PERMISSION_DENIED` must be used. HTTP Mapping: 404 Not Found - ALREADY_EXISTS: The entity that a client attempted to create (e.g., file or directory) already exists. HTTP Mapping: 409 Conflict - PERMISSION_DENIED: The caller does not have permission to execute the specified operation. `PERMISSION_DENIED` must not be used for rejections caused by exhausting some resource (use `RESOURCE_EXHAUSTED` instead for those errors). `PERMISSION_DENIED` must not be used if the caller can not be identified (use `UNAUTHENTICATED` instead for those errors). This error code does not imply the request is valid or the requested entity exists or satisfies other pre-conditions. HTTP Mapping: 403 Forbidden - UNAUTHENTICATED: The request does not have valid authentication credentials for the operation. HTTP Mapping: 401 Unauthorized - RESOURCE_EXHAUSTED: Some resource has been exhausted, perhaps a per-user quota, or perhaps the entire file system is out of space. HTTP Mapping: 429 Too Many Requests - FAILED_PRECONDITION: The operation was rejected because the system is not in a state required for the operation's execution. For example, the directory to be deleted is non-empty, an rmdir operation is applied to a non-directory, etc. Service implementors can use the following guidelines to decide between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`: (a) Use `UNAVAILABLE` if the client can retry just the failing call. (b) Use `ABORTED` if the client should retry at a higher level. For example, when a client-specified test-and-set fails, indicating the client should restart a read-modify-write sequence. (c) Use `FAILED_PRECONDITION` if the client should not retry until the system state has been explicitly fixed. For example, if an "rmdir" fails because the directory is non-empty, `FAILED_PRECONDITION` should be returned since the client should not retry unless the files are deleted from the directory. HTTP Mapping: 400 Bad Request - ABORTED: The operation was aborted, typically due to a concurrency issue such as a sequencer check failure or transaction abort. See the guidelines above for deciding between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`. HTTP Mapping: 409 Conflict - OUT_OF_RANGE: The operation was attempted past the valid range. E.g., seeking or reading past end-of-file. Unlike `INVALID_ARGUMENT`, this error indicates a problem that may be fixed if the system state changes. For example, a 32-bit file system will generate `INVALID_ARGUMENT` if asked to read at an offset that is not in the range [0,2^32-1], but it will generate `OUT_OF_RANGE` if asked to read from an offset past the current file size. There is a fair bit of overlap between `FAILED_PRECONDITION` and `OUT_OF_RANGE`. We recommend using `OUT_OF_RANGE` (the more specific error) when it applies so that callers who are iterating through a space can easily look for an `OUT_OF_RANGE` error to detect when they are done. HTTP Mapping: 400 Bad Request - UNIMPLEMENTED: The operation is not implemented or is not supported/enabled in this service. HTTP Mapping: 501 Not Implemented - INTERNAL: Internal errors. This means that some invariants expected by the underlying system have been broken. This error code is reserved for serious errors. HTTP Mapping: 500 Internal Server Error - UNAVAILABLE: The service is currently unavailable. This is most likely a transient condition, which can be corrected by retrying with a backoff. Note that it is not always safe to retry non-idempotent operations. See the guidelines above for deciding between `FAILED_PRECONDITION`, `ABORTED`, and `UNAVAILABLE`. HTTP Mapping: 503 Service Unavailable - DATA_LOSS: Unrecoverable data loss or corruption. HTTP Mapping: 500 Internal Server Error
    - `message` string — A developer-facing error message in English.
  - `annotations` object — Annotations to identify deployment properties. Key/value pairs may be used by external tools or other services. The "image-tag-reason" key is redacted from API responses for non-superuser principals.
  - `minReplicaCount` integer — The minimum number of replicas. If not specified, the default is 0.
  - `maxReplicaCount` integer — The maximum number of replicas. If not specified, the default is max(min_replica_count, 1). May be set to 0 to downscale the deployment to 0.
  - `maxWithRevocableReplicaCount` integer — max_with_revocable_replica_count is max replica count including revocable capacity. The max revocable capacity will be max_with_revocable_replica_count - max_replica_count.
  - `desiredReplicaCount` integer — The desired number of replicas for this deployment. This represents the target replica count that the system is trying to achieve.
  - `replicaCount` integer
  - `autoscalingPolicy` GatewayAutoscalingPolicy
    - `scaleUpWindow` string — The duration the autoscaler will wait before scaling up a deployment after observing increased load. Default is 30s. Must be less than or equal to 1 hour.
    - `scaleDownWindow` string — The duration the autoscaler will wait before scaling down a deployment after observing decreased load. Default is 10m. Must be less than or equal to 1 hour.
    - `scaleToZeroWindow` string — The duration after which there are no requests that the deployment will be scaled down to zero replicas, if min_replica_count==0. Default is 1h. This must be at least 5 minutes.
    - `loadTargets` object
    - `scalingSchedules` object — Named scaling schedules that override min_replica_count on a time-based cron schedule. When multiple schedules are active simultaneously, the highest min_replica_count across all active schedules is used ("max wins"). When no schedule is active, the deployment's base min_replica_count applies. Maximum 5 schedules per deployment.
  - `baseModel` string, required
  - `acceleratorCount` integer — The number of accelerators used per replica. If not specified, the default is the estimated minimum required by the base model.
  - `acceleratorType` 'ACCELERATOR_TYPE_UNSPECIFIED' | 'NVIDIA_A100_80GB' | 'NVIDIA_H100_80GB' | 'AMD_MI300X_192GB' | 'NVIDIA_A10G_24GB' | 'NVIDIA_A100_40GB' | 'NVIDIA_L4_24GB' | 'NVIDIA_H200_141GB' | 'NVIDIA_B200_180GB' | 'AMD_MI325X_256GB' | 'AMD_MI350X_288GB' | 'NVIDIA_B300_288GB' | 'NVIDIA_GB200' | 'NVIDIA_GB300'
  - `precision` 'PRECISION_UNSPECIFIED' | 'FP16' | 'FP8' | 'FP8_MM' | 'FP8_AR' | 'FP8_MM_KV_ATTN' | 'FP8_KV' | 'FP8_MM_V2' | 'FP8_V2' | 'FP8_MM_KV_ATTN_V2' | 'NF4' | 'FP4' | 'BF16' | 'FP4_BLOCKSCALED_MM' | 'FP4_MX_MOE'
  - `maxConcurrencyPerReplica` integer — The maximum number of concurrent (in-flight) requests a single replica will accept before shedding load. Requests that arrive while a replica is already at this limit are rejected early with HTTP 429 instead of queueing — a per-replica admission gate for controlling tail latency. When unset (0), the platform default is used.
  - `cluster` string — If set, this deployment is deployed to a cloud-premise cluster.
  - `enableAddons` boolean — If true, PEFT addons are enabled for this deployment.
  - `draftTokenCount` integer — The number of candidate tokens to generate per step for speculative decoding. Default is the base model's draft_token_count. Set CreateDeploymentRequest.disable_speculative_decoding to false to disable this behavior.
  - `draftModel` string — The draft model name for speculative decoding. e.g. accounts/fireworks/models/my-draft-model If empty, speculative decoding using a draft model is disabled. Default is the base model's default_draft_model. Set CreateDeploymentRequest.disable_speculative_decoding to false to disable this behavior.
  - `ngramSpeculationLength` integer — The length of previous input sequence to be considered for N-gram speculation.
  - `enableSessionAffinity` boolean — Whether to apply sticky routing based on `user` field. Serverless will be set to true when creating deployment.
  - `directRouteApiKeys` string[] — The set of API keys used to access the direct route deployment. If direct routing is not enabled, this field is unused.
  - `numPeftDeviceCached` integer
  - `directRouteType` 'DIRECT_ROUTE_TYPE_UNSPECIFIED' | 'INTERNET' | 'GCP_PRIVATE_SERVICE_CONNECT' | 'AWS_PRIVATELINK'
  - `directRouteHandle` string — The handle for calling a direct route. The meaning of the handle depends on the direct route type of the deployment: INTERNET -> The host name for accessing the deployment GCP_PRIVATE_SERVICE_CONNECT -> The service attachment name used to create the PSC endpoint. AWS_PRIVATELINK -> The service name used to create the VPC endpoint.
  - `deploymentTemplate` string — The name of the deployment template to use for this deployment. Only available to enterprise accounts.
  - `autoTune` GatewayAutoTune
    - `longPrompt` boolean — If true, this deployment is optimized for long prompt lengths.
  - `placement` GatewayPlacement — The desired geographic region where the deployment must be placed. Exactly one field will be specified.
    - `region` 'REGION_UNSPECIFIED' | 'US_IOWA_1' | 'US_VIRGINIA_1' | 'US_VIRGINIA_2' | 'US_ILLINOIS_1' | 'AP_TOKYO_1' | 'US_ARIZONA_1' | 'US_TEXAS_1' | 'US_ILLINOIS_2' | 'EU_FRANKFURT_1' | 'US_TEXAS_2' | 'EU_ICELAND_1' | 'EU_ICELAND_2' | 'US_WASHINGTON_1' | 'US_WASHINGTON_2' | 'US_WASHINGTON_3' | 'AP_TOKYO_2' | 'US_CALIFORNIA_1' | 'US_UTAH_1' | 'US_ARIZONA_3' | 'US_GEORGIA_1' | 'US_GEORGIA_2' | 'US_WASHINGTON_4' | 'US_GEORGIA_3' | 'NA_BRITISHCOLUMBIA_1' | 'US_GEORGIA_4' | 'US_OHIO_1' | 'US_NEWYORK_1' | 'EU_NETHERLANDS_1' | 'US_WASHINGTON_5' | 'US_MINNESOTA_1' | 'US_CALIFORNIA_2' | 'NA_BRITISHCOLUMBIA_2' | 'AP_MALAYSIA_2' | 'US_OREGON_1' | 'NA_BRITISHCOLUMBIA_3' | 'AP_NEWSOUTHWALES_1'
    - `multiRegion` 'MULTI_REGION_UNSPECIFIED' | 'GLOBAL' | 'US' | 'EUROPE' | 'APAC'
    - `regions` GatewayRegion[]
  - `region` 'REGION_UNSPECIFIED' | 'US_IOWA_1' | 'US_VIRGINIA_1' | 'US_VIRGINIA_2' | 'US_ILLINOIS_1' | 'AP_TOKYO_1' | 'US_ARIZONA_1' | 'US_TEXAS_1' | 'US_ILLINOIS_2' | 'EU_FRANKFURT_1' | 'US_TEXAS_2' | 'EU_ICELAND_1' | 'EU_ICELAND_2' | 'US_WASHINGTON_1' | 'US_WASHINGTON_2' | 'US_WASHINGTON_3' | 'AP_TOKYO_2' | 'US_CALIFORNIA_1' | 'US_UTAH_1' | 'US_ARIZONA_3' | 'US_GEORGIA_1' | 'US_GEORGIA_2' | 'US_WASHINGTON_4' | 'US_GEORGIA_3' | 'NA_BRITISHCOLUMBIA_1' | 'US_GEORGIA_4' | 'US_OHIO_1' | 'US_NEWYORK_1' | 'EU_NETHERLANDS_1' | 'US_WASHINGTON_5' | 'US_MINNESOTA_1' | 'US_CALIFORNIA_2' | 'NA_BRITISHCOLUMBIA_2' | 'AP_MALAYSIA_2' | 'US_OREGON_1' | 'NA_BRITISHCOLUMBIA_3' | 'AP_NEWSOUTHWALES_1'
  - `maxContextLength` integer — The maximum context length supported by the model (context window). If set to 0 or not specified, the model's default maximum context length will be used.
  - `updateTime` string, date-time — The update time for the deployment.
  - `disableDeploymentSizeValidation` boolean — Whether the deployment size validation is disabled.
  - `enableHotLoad` boolean — Whether to use hot load for this deployment.
  - `hotLoadBucketType` 'BUCKET_TYPE_UNSPECIFIED' | 'MINIO' | 'S3' | 'NEBIUS' | 'FW_HOSTED'
  - `enableHotReloadLatestAddon` boolean — Allows up to 1 addon at a time to be loaded, and will merge it into the base model.
  - `deploymentShape` string — The name of the deployment shape that this deployment is using. On the server side, this will be replaced with the deployment shape version name.
  - `activeModelVersion` string — The model version that is currently active and applied to running replicas of a deployment.
  - `targetModelVersion` string — The target model version that is being rolled out to the deployment. In a ready steady state, the target model version is the same as the active model version.
  - `replicaStats` GatewayReplicaStats
    - `pendingSchedulingReplicaCount` integer — Number of replicas waiting to be scheduled to a node.
    - `downloadingModelReplicaCount` integer — Number of replicas downloading model weights.
    - `initializingReplicaCount` integer — Number of replicas initializing the model server.
    - `readyReplicaCount` integer — Number of replicas that are ready and serving traffic.
    - `revocableReplicaCount` integer
    - `effectiveReplicaCount` number, float — The effective number of replicas currently serving traffic, including fractional progress toward a replica that is not yet fully ready. Equals ready_replica_count when no replica is partially ready.
  - `hotLoadBucketUrl` string
  - `pricingPlanId` string — Optional pricing plan ID for custom billing configuration. If set, this deployment will use the pricing plan's billing rules instead of default billing behavior.
  - `hotLoadTrainerJob` string
  - `hotLoadTransitionType` 'HOT_LOAD_TRANSITION_TYPE_UNSPECIFIED' | 'ASYNC' | 'SYNC'
  - `preemptible` boolean — When true, this deployment runs as preemptible.

---

[API](https://skmtc.net/fireworks/apis/fireworks-ai-anthropic-compatible-messages-api.md) · [All operations](https://skmtc.net/fireworks/apis/fireworks-ai-anthropic-compatible-messages-api/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/fireworks/fireworks-ai-anthropic-compatible-messages-api/versions/954d6bc5d922/schema)
