---
title: "Set the deployment pool's member-selection routing (platform admin)"
method: POST
path: "/v1/admin/inference/deployments/{id}/pool-routing"
tags: ["Internal"]
---

# Set the deployment pool's member-selection routing (platform admin)

`POST /v1/admin/inference/deployments/{id}/pool-routing`

Sets the endpoint pool's member-selection policy and/or the prefix_affinity load guard. This edits the POOL (deployments sharing a pool share the setting) and reaches every gateway task in ~2s via the config generation stamp; setting policy back to p2c_util is the instant rollback. Returns the refreshed replica view.

## Path parameters

- `id` string, required

## Request body

- ServingPoolRoutingRequest — Pool routing knobs; only the provided fields change. policy switches the member-selection algorithm; affinityGuardUtil bounds how full a conversation's home member may run before prefix_affinity overflows to load-only P2C (0 clears it back to the gateway default).
  - `affinityGuardUtil` number, double
  - `policy` 'p2c_util' | 'prefix_affinity'

## Response `200`

Updated

- ServingReplicaList
  - `affinityGuardUtil` number, double — prefix_affinity overflow bound; absent = the gateway default
  - `items` ServingReplica[], required
    - `capacity` integer — In-flight bound the engine was launched with (= pool member capacity)
    - `createdAt` string
    - `gpuCount` integer
    - `hostname` string — run.* hostname the gateway dials; the pool-member key
    - `load` ServingReplicaLoad — The replica's last scraped engine load, as published to the gateways. Absent fields mean the engine did not expose the metric; a fabricated zero is never reported.
      - `atMs` integer, required — Publish time, unix milliseconds; readings expire after ~45s
      - `hasKvUtil` boolean
      - `hasWaiting` boolean
      - `kvUtil` number — KV-cache utilization 0..1; present only when hasKvUtil
      - `running` integer, required — Requests currently executing on the engine
      - `waiting` integer — Requests queued on the engine; present only when hasWaiting
    - `memberCapacity` integer — Capacity currently advertised on the pool member; present only when a member exists
    - `memberState` 'active' | 'draining' | 'out' | 'none', required — Pool membership: none = not (yet) a member of the deployment's pool
    - `memberWeight` number, double — Ranking weight on the pool member (the warmup/canary dial); absent = full share (1.0)
    - `name` string
    - `nodeId` string — Node the replica placed on; absent before placement
    - `nomadAllocId` string
    - `provisioningStage` string
    - `status` string, required — VM lifecycle status (deploying, running, stopped, terminated, failed); "unreadable" when the row could not be read this request
    - `statusReason` string — Defer/failure reason from dispatch (weights staging, image staging, placement, ...)
    - `vmId` string, required
  - `policy` 'p2c_util' | 'prefix_affinity', required — The pool's member-selection policy (p2c_util when the pool row does not exist yet)
  - `poolId` string, required — The deployment's endpoint pool id
  - `strayMembers` string[] — Pool member hostnames no current replica owns (pending the reconcile loop's stray sweep). Usually empty.

## Other responses

- `400` — The request is invalid
- `401` — Missing or invalid API key
- `403` — API key lacks the required scope
- `404` — Resource not found
- `409` — The request conflicts with existing state
- `503` — A required integration (e.g. payments) is not configured

---

[API](https://skmtc.net/openrelay/apis/openrelay-api.md) · [All operations](https://skmtc.net/openrelay/apis/openrelay-api/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/openrelay/openrelay-api/versions/32905b8f44f8/schema)
