v32

latestOpenAPI 3.1.0raw.githubusercontent.com2026-05-151,1412,2144.5 MB
Everywhere Inference

Stop inference deployment

This operation shuts down an inference deployment, making it unavailable for handling requests. The deployment will scale down to 0 replicas, overriding any minimum replica settings.

  • Once stopped, the deployment will not process any inference requests or SQS messages.
  • It will not restart automatically and must be started manually.
  • While stopped, the deployment will not incur any charges.
post/cloud/v3/inference/{project_id}/deployments/{deployment_name}/stop

Path parameters

project_idinteger required

Project ID

Example:1

Project ID

deployment_namestring required

Inference instance name.

Example:my-instance

Inference instance name.

Response

No Content