v4

latestOpenAPI 3.1.02026-07-31155218294.8 KB
Dedicated Models

Deploy Llm Presets

DeepInfra presets and mirrored vLLM recipes for hf_repo_id, told apart by source; empty when none. Filter by gpu/engine/source.

get/deploy/llm/presets

Query parameters

hf_repo_idstring required
gpu'L4-24GB' | 'L40S-48GB' | 'A100-80GB' | 'H100-80GB' | 'H200-141GB' | 'B200-180GB' | 'B300-270GB' | 'RTXPRO6000-96GB' | 'other'
enginestring nullable
sourcestring nullable

Headers

xi-api-keystring nullable
x-api-keystring nullable

Response

Successful Response

idstring required

Preset id.

sourcestring

Config source.

enginestring

Inference engine.

gpu_configsstring[] required

Allowed Nx<GPU> configs.

extra_argsstring[] nullable

Raw engine flags; vLLM recipes only.

labelstring

Short display name (e.g. "Throughput-optimized").