OpenAPI 3.1.02026-08-205711,1272.8 MB
c768d4d28e4c
Retriever Evaluations
Score predictions vs ground truth (stateless)
Compute quality metrics (Precision@K, Recall@K, F1@K, F2@K) for precomputed predictions vs ground truth, with NO retriever, namespace, or persistence. Each item's predicted and ground_truth are treated as SETS. Send a single pair (predicted + ground_truth) or an items batch. Returns per-item scores and the macro-average across items. Auth only (no X-Namespace). This is the stateless counterpart to the retriever-scoped evaluation runs — use it to dogfood F1/F2 on extraction/generation benchmarks.
post/v1/evaluations/score
Request body
Response
Successful Response