v32

latestOpenAPI 3.1.0raw.githubusercontent.com2026-05-151,1412,2144.5 MB
Everywhere Inference Apps

Create inference application deployment

Creates a new application deployment based on a selected catalog application. Specify the desired deployment name, target regions, and configuration for each component. The platform will provision the necessary resources and initialize the application accordingly.

post/cloud/v3/inference/applications/{project_id}/deployments

Path parameters

project_idinteger required

Project ID

Example:1

Project ID

Request body

api_keysstring[]

List of API keys for the application

application_namestring required

Identifier of the application from the catalog

components_configurationobject required

Mapping of component names to their configuration (e.g., "model": {...})

namestring required

Desired name for the new deployment

regionsinteger[] required

Geographical regions where the deployment should be created

Example request

{
  "api_keys": [
    "key1",
    "key2"
  ],
  "application_name": "demo-app",
  "components_configuration": {
    "model": {
      "exposed": true,
      "flavor": "inference-16vcpu-232gib-1xh100-80gb",
      "scale": {
        "max": 1,
        "min": 1
      }
    }
  },
  "regions": [
    1,
    2
  ]
}

Response

OK

tasksstring[] required

List of task IDs representing asynchronous operations. Use these IDs to monitor operation progress:

  • GET /v1/tasks/{task_id} - Check individual task status and details Poll task status until completion (FINISHED/ERROR) before proceeding with dependent operations.
idstring required

Example response

{
  "tasks": [
    "d478ae29-dedc-4869-82f0-96104425f565"
  ]
}