Skip to main content

GET/1.0/caliper/evals/:id/runs/:runId

Returns one run of an eval: its lifecycle status, the overall score once scoring settles, item totals, and every item with its per-criterion scores. This is the poll half of the CI loop. Submit a run, keep its id, and read this endpoint until status is DONE or FAILED.

Required scope: caliper:evals:read (scopes reference).

Path parameters

idstringRequired

The eval's id. The key's workspace must own the eval.

runIdstringRequired

The run's id, from the 202 response of submit a run. A run that belongs to a different eval is a 404.

Response (200)

idstring

The run's id.

evalIdstring

The eval this run belongs to.

status"PENDING" | "RUNNING" | "SCORING" | "DONE" | "FAILED"

Lifecycle state. External runs move PENDING → SCORING → DONE. Poll until DONE or FAILED.

overallScorenumber | null

Aggregate across items and criteria, normalized to 0–1. null until the run reaches DONE. This is the number to gate a build on.

totalItemsinteger

How many items the run holds.

completedItemsinteger

How many have finished scoring. Rises as the judge works through the items.

failedItemsinteger

How many failed scoring. The run still reaches DONE unless every item failed.

triggeredBy"ci" | "ui" | "workbench_publish" | "agent"

What started the run. Runs submitted through this API carry what the request said, usually ci.

labelstring | null

The label the request set, for example "PR #1234".

errorMessagestring | null

Set when status is FAILED. Per-item failures live on the item rows below.

startedAtstring | null

ISO-8601 timestamp, or null while pending.

completedAtstring | null

ISO-8601 timestamp once the run settles.

createdAtstring

ISO-8601 timestamp.

itemsarray

One entry per submitted item.

  • id — the run item's id.
  • itemId — the dataset item it scored, matching dataset.items[].id on read an eval.
  • status — PENDING | SCORING | DONE | FAILED.
  • actualOutput — the output your script submitted.
  • scores — { [criterionId]: { value, reasoning, judgeModel } } once scored, null before. Criterion ids match rubricSnapshot.criteria[].id on the eval.
  • errorMessage — why this item failed, when it did.

Example

curl https://api.zerowidth.ai/1.0/caliper/evals/eval_8a2c91.../runs/run_4f1d20... \
  -H "Authorization: Bearer zw_live_abc12345_..."

Polling

Scoring takes a few seconds per item. Poll every few seconds with a ceiling of a few minutes for a typical dataset, and treat FAILED as a build failure with errorMessage in the log. Reading the eval's latestScore instead of this endpoint is ambiguous when two runs overlap, so gate on the run you submitted.

Errors

statuscodewhen
401auth_missing / auth_invalidNo or unknown bearer key
403scope_missingKey lacks caliper:evals:read
403forbiddenPersonal keys aren't accepted here; use a workspace key
404not_foundThe eval or run doesn't exist, isn't this eval's, or isn't visible to the key's workspace
3 min read