Skip to main content

POST/1.0/caliper/feedback

Record one piece of feedback — a thumb, a comment, or both — against a response your product showed someone. The feedback lands as an item in a Caliper dataset, so the raw material for your first eval accumulates while people use the thing.

Call it from wherever your AI actually runs.

Required scope: caliper:datasets:write (scopes reference). Workspace-kind keys only.

Returns 201 Created with the dataset the feedback landed in.

Request body

Say what the feedback is about with either flowUuid or target — exactly one. Sending both, or neither, is a 400.

flowUuidstring

A ZeroWidth flow's uuid — the same id you run it by at POST /1.0/flows/:flowUuid/runs.

Feedback lands in that flow's own dataset, alongside anything its in-product thumbs collected, so you get one view of how it's doing rather than two halves. The internal id from the editor URL is not accepted here, same as everywhere else on the API.

targetstring

For a system we don't host: your own name for it — "support-assistant", "checkout-copilot".

This names the dataset. Every call with the same target accumulates into one, created the first time you post, which is what makes the feedback reviewable as a set later. A target that happens to look like a ZeroWidth id still can't write into a flow's dataset — the two are kept apart.

inputanyRequired

What the person asked. A string, or an object — if it carries a recognizable messages array, the item is stored as a chat transcript rather than a single question, and the whole conversation is visible in Caliper.

outputanyRequired

What your AI answered. Same shapes as input.

rating"thumbs_up" | "thumbs_down" | null

The verdict. null clears a rating you sent earlier. Omit it entirely to record the exchange without a verdict — useful for saving an interesting case that nobody has judged yet.

commentstring

What the person said in their own words. Worth capturing even when the rating is positive; it's usually the part that tells you what to change.

sourceSubIdstring

Your own id for this exchange, stored alongside the item so you can match a row back to the conversation it came from. Every call records a new row, so if someone flips a thumb you'll see both.

contextany

Anything else worth keeping with the case: the user's plan, the retrieved documents, the model you called. Stored alongside the item and visible when reviewing.

timestampstring

When the exchange actually happened, ISO-8601 with a timezone. Defaults to now — set it when you're backfilling, so timelines reflect real event time rather than when you imported.

intent"captured" | "expected" | "input_only"

What the saved item is for. "captured" (the default) stores the answer as what your system said — ready for a human to rate. "expected" stores it as the right answer, making the item a test case a future version has to match. "input_only" keeps just the question, for a case whose correct answer nobody has written yet.

The source field is always recorded as external_api for key-authenticated calls, whatever the body says.

Example

curl -X POST https://api.zerowidth.ai/1.0/caliper/feedback \
  -H "Authorization: Bearer $ZEROWIDTH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "target": "support-assistant",
    "sourceSubId": "conv_8123",
    "rating": "thumbs_down",
    "comment": "Told me to check a page that does not exist.",
    "input": { "question": "How do I cancel mid-cycle?" },
    "output": { "answer": "Visit Settings → Billing → Cancel." },
    "context": { "plan": "pro", "model": "claude-sonnet-4" }
  }'

Response (201)

datasetIdstring

The dataset the feedback landed in. Open it in Caliper to review what's accumulated.

datasetCreatedboolean

true on the first call for a given subject — the dataset was created by this request.

itemIdstring

The item's id within the dataset.

Where this leads

A dataset of real thumbs-down cases is the most useful test set you can have, because every item is a failure someone actually hit. Once enough has collected:

  1. Open the dataset in Caliper and read what came in.
  2. Write a rubric for what a good answer would have been.
  3. Run an eval against the set, so the next version is scored on the cases that went wrong before.

See datasets and the quickstart.

4 min read