GET/1.0/caliper/evals/:id/runs/:runId
Returns one run of an eval: its lifecycle status, the overall score once scoring settles, item totals, and every item with its per-criterion scores. This is the poll half of the CI loop. Submit a run, keep its id, and read this endpoint until status is DONE or FAILED.
Required scope: caliper:evals:read (scopes reference).
Path parameters
idstringRequiredThe eval's id. The key's workspace must own the eval.
runIdstringRequiredThe run's id, from the 202 response of submit a run. A run that belongs to a different eval is a 404.
Response (200)
idstringThe run's id.
evalIdstringThe eval this run belongs to.
status"PENDING" | "RUNNING" | "SCORING" | "DONE" | "FAILED"Lifecycle state. External runs move PENDING → SCORING → DONE. Poll until DONE or FAILED.
overallScorenumber | nullAggregate across items and criteria, normalized to 0–1. null until the run reaches DONE. This is the number to gate a build on.
totalItemsintegerHow many items the run holds.
completedItemsintegerHow many have finished scoring. Rises as the judge works through the items.
failedItemsintegerHow many failed scoring. The run still reaches DONE unless every item failed.
triggeredBy"ci" | "ui" | "workbench_publish" | "agent"What started the run. Runs submitted through this API carry what the request said, usually ci.
labelstring | nullThe label the request set, for example "PR #1234".
errorMessagestring | nullSet when status is FAILED. Per-item failures live on the item rows below.
startedAtstring | nullISO-8601 timestamp, or null while pending.
completedAtstring | nullISO-8601 timestamp once the run settles.
createdAtstringISO-8601 timestamp.
itemsarrayOne entry per submitted item.
id— the run item's id.itemId— the dataset item it scored, matchingdataset.items[].idon read an eval.status—PENDING|SCORING|DONE|FAILED.actualOutput— the output your script submitted.scores—{ [criterionId]: { value, reasoning, judgeModel } }once scored,nullbefore. Criterion ids matchrubricSnapshot.criteria[].idon the eval.errorMessage— why this item failed, when it did.
Example
curl https://api.zerowidth.ai/1.0/caliper/evals/eval_8a2c91.../runs/run_4f1d20... \
-H "Authorization: Bearer zw_live_abc12345_..."{
"id": "run_4f1d20...",
"evalId": "eval_8a2c91...",
"status": "DONE",
"triggeredBy": "ci",
"label": "PR #1234",
"overallScore": 0.87,
"totalItems": 2,
"completedItems": 2,
"failedItems": 0,
"errorMessage": null,
"startedAt": "2026-09-02T21:40:02.000Z",
"completedAt": "2026-09-02T21:40:19.000Z",
"createdAt": "2026-09-02T21:40:01.000Z",
"items": [
{
"id": "ri_01...",
"itemId": "item_001",
"status": "DONE",
"actualOutput": "We accept refunds within 30 days.",
"scores": {
"crit_tone": { "value": 5, "reasoning": "Warm and direct.", "judgeModel": "caliper-judge" },
"crit_policy": { "value": 4, "reasoning": "States the window; omits the receipt requirement.", "judgeModel": "caliper-judge" }
},
"errorMessage": null
}
]
}Polling
Scoring takes a few seconds per item. Poll every few seconds with a ceiling of a few minutes for a typical dataset, and treat FAILED as a build failure with errorMessage in the log. Reading the eval's latestScore instead of this endpoint is ambiguous when two runs overlap, so gate on the run you submitted.
Errors
| status | code | when |
|---|---|---|
| 401 | auth_missing / auth_invalid | No or unknown bearer key |
| 403 | scope_missing | Key lacks caliper:evals:read |
| 403 | forbidden | Personal keys aren't accepted here; use a workspace key |
| 404 | not_found | The eval or run doesn't exist, isn't this eval's, or isn't visible to the key's workspace |