v2
latestOpenAPI 3.1.02026-08-07139158192.4 KBCreate a self-hosted serving deployment (platform admin)
Records the operator's intent to run N replicas of an inference engine for one public catalog model. The spec is validated hard at admission (the engine must exist and the replica must plan) so a row that could never dispatch is rejected here instead of failing later. The new deployment starts in status pending; the serving reconcile loop admits replicas, stages weights, probes readiness, and registers the pool backend from there.
Request body
Initial routing weight; defaults to 0 (standby)
Must resolve to a known engine (e.g. vllm, sglang)
Engine container image override
Launch and health-check replicas without ever publishing their pool as a catalog backend. Requires backendWeight 0.
Exclusive text-request byte ceiling for this deployment's routing lane; 0 or absent = unbounded
Inclusive text-request byte floor for this deployment's routing lane; 0 or absent = unbounded
Defaults to chat_completions at registration
Stable logical capacity lane, e.g. rtx5090 or h100. Deployments in different node pools keep independent routing weights.
Org the managed replicas run (and meter) under
Catalog model the pool backend registers on. Must already be seeded, with pricing
Must lie within replicasMin..replicasMax
Must equal the catalog backend ref the gateway rewrites to
Staged-checkpoint reference; omit to let the engine load from its own source
Response
Created
Live catalog routing weight of this deployment's backend. Only present when backendRegistered is true. May differ from backendWeight (the row mirror) while a rollout ramp is in flight.
Live catalog health of this deployment's backend (the kill switch). Only present when backendRegistered is true.
Whether this deployment's pool:// backend is currently present on the catalog model row. Absent when the live catalog state could not be read; false until the first replica passes readiness.
Routing weight of the catalog backend. 0 = standby (blue/green starts here)
Inference engine, e.g. vllm or sglang
Engine container image override; absent = the engine's default
True when replicas and their private endpoint pool are for direct operator testing only. Isolated deployments never register a catalog backend, never participate in blue/green rollout, and cannot receive gateway traffic.
Exclusive text-request byte ceiling for routing eligibility; 0 or absent = unbounded. Routing hint only, not a context limit.
Inclusive text-request byte floor for routing eligibility; 0 or absent = unbounded. Routing hint only, not a context limit.
Stable logical capacity lane within the public model. Blue/green only replaces versions in the same node pool; empty preserves legacy model-wide grouping.
Org that owns the managed replicas (and their metering)
Endpoint pool the deployment's replicas register in
Catalog model this deployment serves capacity for
Model id the engine serves; must equal the catalog backend ref
Human-readable cause for degraded/failed states
Generation marker for blue/green within one public model and node pool
Staged-checkpoint reference; absent = the engine loads from its own source