List Benchmarks
Unified benchmark endpoint that aggregates scores from multiple benchmark sources (Artificial Analysis, Design Arena, and OpenRouter's own tau-bench, GPQA, and web-search evals). Filter by source to reproduce the exact shapes from the legacy per-source endpoints, or use task_type to find models suited for specific workloads. Use task_type=search (or a search_* benchmark_type) for OpenRouter's search benchmarks, which publish each model's highest-scoring eligible evaluation configuration with same-configuration runs combined by task-weighted mean. Authenticate with any valid OpenRouter API key. Rate-limited to 30 requests/minute per key and 500 requests/day per account.
Query parameters
Benchmark source to query. Determines the shape of the returned items. When omitted, returns results from all sources.
Benchmark source to query. Determines the shape of the returned items. When omitted, returns results from all sources.
Filter results by task type. For Artificial Analysis, maps to the corresponding index. For Design Arena, maps to the matching category. search returns OpenRouter search benchmark results only.
Filter results by task type. For Artificial Analysis, maps to the corresponding index. For Design Arena, maps to the matching category. search returns OpenRouter search benchmark results only.
Return results for one exact OpenRouter benchmark. A search_* value narrows the response to search results only; a classic value narrows the OpenRouter items and leaves other sources' items as they are.
Return results for one exact OpenRouter benchmark. A search_* value narrows the response to search results only; a classic value narrows the OpenRouter items and leaves other sources' items as they are.
Search benchmarks only: include the published lane configuration whitelist in each search item. Defaults to false. The whitelist is limited to agent turn count, reasoning effort, and temperature so future harness configuration changes do not change the public contract.
Search benchmarks only: include the published lane configuration whitelist in each search item. Defaults to false. The whitelist is limited to agent turn count, reasoning effort, and temperature so future harness configuration changes do not change the public contract.
OpenRouter search benchmarks only: filter by the search engine used.
OpenRouter search benchmarks only: filter by the search engine used.
OpenRouter search benchmarks only: filter by the request surface the lane ran on.
OpenRouter search benchmarks only: filter by the request surface the lane ran on.
Design Arena only: arena to query. Defaults to models when source is design-arena.
Design Arena only: arena to query. Defaults to models when source is design-arena.
Design Arena only: category within the arena (e.g. codecategories, uicomponent, gamedev, 3d, dataviz, image, video, svg). When omitted, returns all categories.
Design Arena only: category within the arena (e.g. codecategories, uicomponent, gamedev, 3d, dataviz, image, video, svg). When omitted, returns all categories.
Maximum number of items to return. When omitted, all matching results are returned.
Maximum number of items to return. When omitted, all matching results are returned.
Response
Benchmark results filtered by the specified source and optional task type.
Example response
{
"data": [
{
"agentic_index": 58.3,
"coding_index": 65.8,
"display_name": "GPT-4o",
"intelligence_index": 71.2,
"model_permaslug": "openai/gpt-4o",
"pricing": {
"completion": "0.00001",
"prompt": "0.0000025"
},
"source": "artificial-analysis"
},
{
"accuracy": 0.72,
"accuracy_stddev": 0.03,
"avg_cost_per_task": 0.002,
"benchmark_type": "gpqa_diamond",
"display_name": "GPT-4o",
"last_run_timestamp": "2026-06-03T12:00:00Z",
"model_permaslug": "openai/gpt-4o",
"source": "openrouter",
"total_tasks": 300
}
],
"meta": {
"as_of": "2026-06-03T12:00:00Z",
"citation": null,
"model_count": 1,
"source": null,
"source_url": null,
"task_type": null,
"version": "v1"
}
}