Evals
Grade an eval run with LLM as Judge
Spends. Runs the goal-completion judge over the finished run, scoring each case's final answer against its expected output.
202: scheduled, not done. Read the grades from the run detail's judges.goalCompletion rather than re-requesting — a second POST only spends again.
A run's grading config is pinned when the run is created, so turning the judge on for the suite does not reach an already-recorded run: enable: true is what grades one, and it changes nothing beyond that run. Omitting model and threshold clears any override a previous request left on the run.
post/projects/{projectId}/eval-runs/{runId}/judge
Request body
Response
Scheduled.