latestOpenAPI 3.1.0raw.githubusercontent.com2026-08-21161224637.5 KB

8978a4e735fa

Evals

Grade an eval run with LLM as Judge

Spends. Runs the goal-completion judge over the finished run, scoring each case's final answer against its expected output.

202: scheduled, not done. Read the grades from the run detail's judges.goalCompletion rather than re-requesting — a second POST only spends again.

A run's grading config is pinned when the run is created, so turning the judge on for the suite does not reach an already-recorded run: enable: true is what grades one, and it changes nothing beyond that run. Omitting model and threshold clears any override a previous request left on the run.

post/projects/{projectId}/eval-runs/{runId}/judge

Request body

forceboolean

Re-grade a run that already has a result. SPENDS AGAIN. Must be a real boolean — "false" is rejected rather than read as consent.

enableboolean

Grade this run even though the judge was off when it ran. A per-RUN answer, not a suite edit: grading reads the config pinned when the run was created, so enabling the judge on the suite does not reach an already-recorded run.

modelstring

Judge model for this run only.

thresholdnumber

Pass threshold for this run only.

Response

Scheduled.

runIdstring required
projectIdstring required
status'pending' required