List all models and their properties
Query parameters
Number of records to skip for pagination. When both offset and limit are omitted, the full list is returned
Number of records to skip for pagination. When both offset and limit are omitted, the full list is returned
Maximum number of records to return (max 1000). When both offset and limit are omitted, the full list is returned
Maximum number of records to return (max 1000). When both offset and limit are omitted, the full list is returned
Filter models by use case category
Filter models by use case category
Filter models by supported parameter (comma-separated)
Filter models by supported parameter (comma-separated)
Filter models by output modality. Accepts a comma-separated list of modalities (text, image, audio, embeddings) or "all" to include all models. Defaults to "text".
Filter models by output modality. Accepts a comma-separated list of modalities (text, image, audio, embeddings) or "all" to include all models. Defaults to "text".
Sort the returned models server-side. Prefer this over fetching the full list and sorting client-side. Options: pricing-low-to-high, pricing-high-to-low (average prompt/completion price), context-high-to-low (context length), throughput-high-to-low, latency-low-to-high (recent median performance), most-popular, top-weekly (tokens processed in the last week), newest (creation date), intelligence-high-to-low, coding-high-to-low, agentic-high-to-low (Artificial Analysis indices), design-arena-elo-high-to-low (best Design Arena ELO across arenas). Models without a score for the chosen benchmark are placed last. When omitted, the existing default ordering is preserved.
Sort the returned models server-side. Prefer this over fetching the full list and sorting client-side. Options: pricing-low-to-high, pricing-high-to-low (average prompt/completion price), context-high-to-low (context length), throughput-high-to-low, latency-low-to-high (recent median performance), most-popular, top-weekly (tokens processed in the last week), newest (creation date), intelligence-high-to-low, coding-high-to-low, agentic-high-to-low (Artificial Analysis indices), design-arena-elo-high-to-low (best Design Arena ELO across arenas). Models without a score for the chosen benchmark are placed last. When omitted, the existing default ordering is preserved.
Return results as RSS feed
Return results as RSS feed
Use chat links in RSS feed items
Use chat links in RSS feed items
Free-text search by model name or slug.
Free-text search by model name or slug.
Filter models by input modality. Comma-separated list of: text, image, audio, file.
Filter models by input modality. Comma-separated list of: text, image, audio, file.
Minimum context length (tokens). Models with smaller context are excluded.
Minimum context length (tokens). Models with smaller context are excluded.
Minimum prompt price in $/M tokens.
Minimum prompt price in $/M tokens.
Maximum prompt price in $/M tokens.
Maximum prompt price in $/M tokens.
Filter models by architecture/model family (e.g. GPT, Claude, Gemini, Llama).
Filter models by architecture/model family (e.g. GPT, Claude, Gemini, Llama).
Filter models by the organization that created the model. Comma-separated list of author slugs.
Filter models by the organization that created the model. Comma-separated list of author slugs.
Filter models by hosting provider. Comma-separated list of provider names.
Filter models by hosting provider. Comma-separated list of provider names.
Filter by distillation capability. "true" returns only distillable models, "false" excludes them.
Filter by distillation capability. "true" returns only distillable models, "false" excludes them.
When set to "true", return only models with zero data retention endpoints.
When set to "true", return only models with zero data retention endpoints.
Filter to models with endpoints in the given data region ("eu" or "us").
Filter to models with endpoints in the given data region ("eu" or "us").
Minimum completion (output) price in $/M tokens.
Minimum completion (output) price in $/M tokens.
Maximum completion (output) price in $/M tokens.
Maximum completion (output) price in $/M tokens.
Minimum model age in days since its creation date.
Minimum model age in days since its creation date.
Maximum model age in days since its creation date.
Maximum model age in days since its creation date.
Minimum Artificial Analysis intelligence index.
Minimum Artificial Analysis intelligence index.
Maximum Artificial Analysis intelligence index.
Maximum Artificial Analysis intelligence index.
Minimum Artificial Analysis coding index.
Minimum Artificial Analysis coding index.
Maximum Artificial Analysis coding index.
Maximum Artificial Analysis coding index.
Minimum Artificial Analysis agentic index.
Minimum Artificial Analysis agentic index.
Maximum Artificial Analysis agentic index.
Maximum Artificial Analysis agentic index.
Minimum tool-calling success rate, as a fraction in [0, 1] (e.g. 0.9 = 90% of requests finishing with a tool_calls finish reason).
Minimum tool-calling success rate, as a fraction in [0, 1] (e.g. 0.9 = 90% of requests finishing with a tool_calls finish reason).
Maximum tool-calling success rate, as a fraction in [0, 1].
Maximum tool-calling success rate, as a fraction in [0, 1].
Response
Returns a list of models or RSS feed
Example response
{
"data": [
{
"architecture": {
"input_modalities": [
"text"
],
"instruct_type": "chatml",
"modality": "text->text",
"output_modalities": [
"text"
],
"tokenizer": "GPT"
},
"canonical_slug": "openai/gpt-4",
"context_length": 8192,
"created": 1692901234,
"default_parameters": null,
"description": "GPT-4 is a large multimodal model that can solve difficult problems with greater accuracy.",
"expiration_date": null,
"id": "openai/gpt-4",
"knowledge_cutoff": null,
"links": {
"details": "/api/v1/models/openai/gpt-4/endpoints"
},
"name": "GPT-4",
"per_request_limits": null,
"pricing": {
"completion": "0.00006",
"image": "0",
"prompt": "0.00003",
"request": "0"
},
"supported_parameters": [
"temperature",
"top_p",
"max_tokens",
"frequency_penalty",
"presence_penalty"
],
"supported_voices": null,
"top_provider": {
"context_length": 8192,
"is_moderated": true,
"max_completion_tokens": 4096
}
}
],
"links": {
"next": "/api/v1/models?offset=500&limit=500"
},
"total_count": 150
}