v126

latestOpenAPI 3.1.0MITraw.githubusercontent.com2026-08-04957391.2 MB
Benchmarks

List Benchmarks

Unified benchmark endpoint that aggregates scores from multiple benchmark sources (Artificial Analysis, Design Arena, and OpenRouter's own tau-bench and GPQA evals). Filter by source to reproduce the exact shapes from the legacy per-source endpoints, or use task_type to find models suited for specific workloads. Authenticate with any valid OpenRouter API key. Rate-limited to 30 requests/minute per key and 500 requests/day per account.

get/benchmarks

Query parameters

source'artificial-analysis' | 'design-arena' | 'openrouter'

Benchmark source to query. Determines the shape of the returned items. When omitted, returns results from all sources.

Example:artificial-analysis

Benchmark source to query. Determines the shape of the returned items. When omitted, returns results from all sources.

task_type'coding' | 'intelligence' | 'agentic'

Filter results by task type. For Artificial Analysis, maps to the corresponding index. For Design Arena, maps to the matching category.

Example:coding

Filter results by task type. For Artificial Analysis, maps to the corresponding index. For Design Arena, maps to the matching category.

arena'models' | 'builders' | 'agents'

Design Arena only: arena to query. Defaults to models when source is design-arena.

Example:models

Design Arena only: arena to query. Defaults to models when source is design-arena.

categorystring

Design Arena only: category within the arena (e.g. codecategories, uicomponent, gamedev, 3d, dataviz, image, video, svg). When omitted, returns all categories.

Example:codecategories

Design Arena only: category within the arena (e.g. codecategories, uicomponent, gamedev, 3d, dataviz, image, video, svg). When omitted, returns all categories.

max_resultsinteger

Maximum number of items to return. When omitted, all matching results are returned.

Example:50

Maximum number of items to return. When omitted, all matching results are returned.

Response

Benchmark results filtered by the specified source and optional task type.

Example response

{
  "data": [
    {
      "agentic_index": 58.3,
      "coding_index": 65.8,
      "display_name": "GPT-4o",
      "intelligence_index": 71.2,
      "model_permaslug": "openai/gpt-4o",
      "pricing": {
        "completion": "0.00001",
        "prompt": "0.0000025"
      },
      "source": "artificial-analysis"
    },
    {
      "accuracy": 0.72,
      "accuracy_stddev": 0.03,
      "avg_cost_per_task": 0.002,
      "benchmark_type": "gpqa_diamond",
      "display_name": "GPT-4o",
      "last_run_timestamp": "2026-06-03T12:00:00Z",
      "model_permaslug": "openai/gpt-4o",
      "source": "openrouter",
      "total_tasks": 300
    }
  ],
  "meta": {
    "as_of": "2026-06-03T12:00:00Z",
    "citation": null,
    "model_count": 1,
    "source": null,
    "source_url": null,
    "task_type": null,
    "version": "v1"
  }
}