---
title: "Per-replica runtime view of a serving deployment (platform admin)"
method: GET
path: "/v1/admin/inference/deployments/{id}/replicas"
tags: ["Internal"]
---

# Per-replica runtime view of a serving deployment (platform admin)

`GET /v1/admin/inference/deployments/{id}/replicas`

Joins each replica VM row with its pool-member state and the last scraped engine load, so an operator can see where every replica is stuck (placement, weights staging, image staging, readiness) and what traffic it is taking, without walking Nomad and Redis by hand.

## Path parameters

- `id` string, required

## Response `200`

The deployment's replicas

- ServingReplicaList
  - `affinityGuardUtil` number, double — prefix_affinity overflow bound; absent = the gateway default
  - `items` ServingReplica[], required
    - `capacity` integer — In-flight bound the engine was launched with (= pool member capacity)
    - `createdAt` string
    - `gpuCount` integer
    - `hostname` string — run.* hostname the gateway dials; the pool-member key
    - `load` ServingReplicaLoad — The replica's last scraped engine load, as published to the gateways. Absent fields mean the engine did not expose the metric; a fabricated zero is never reported.
      - `atMs` integer, required — Publish time, unix milliseconds; readings expire after ~45s
      - `hasKvUtil` boolean
      - `hasWaiting` boolean
      - `kvUtil` number — KV-cache utilization 0..1; present only when hasKvUtil
      - `running` integer, required — Requests currently executing on the engine
      - `waiting` integer — Requests queued on the engine; present only when hasWaiting
    - `memberCapacity` integer — Capacity currently advertised on the pool member; present only when a member exists
    - `memberState` 'active' | 'draining' | 'out' | 'none', required — Pool membership: none = not (yet) a member of the deployment's pool
    - `memberWeight` number, double — Ranking weight on the pool member (the warmup/canary dial); absent = full share (1.0)
    - `name` string
    - `nodeId` string — Node the replica placed on; absent before placement
    - `nomadAllocId` string
    - `provisioningStage` string
    - `status` string, required — VM lifecycle status (deploying, running, stopped, terminated, failed); "unreadable" when the row could not be read this request
    - `statusReason` string — Defer/failure reason from dispatch (weights staging, image staging, placement, ...)
    - `vmId` string, required
  - `policy` 'p2c_util' | 'prefix_affinity', required — The pool's member-selection policy (p2c_util when the pool row does not exist yet)
  - `poolId` string, required — The deployment's endpoint pool id
  - `strayMembers` string[] — Pool member hostnames no current replica owns (pending the reconcile loop's stray sweep). Usually empty.

## Other responses

- `401` — Missing or invalid API key
- `403` — API key lacks the required scope
- `404` — Resource not found
- `503` — A required integration (e.g. payments) is not configured

---

[API](https://skmtc.net/openrelay/apis/openrelay-api.md) · [All operations](https://skmtc.net/openrelay/apis/openrelay-api/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/openrelay/openrelay-api/versions/32905b8f44f8/schema)
