Sources
A source is an app sending its agent's conversations to Caliper: what came in, what the agent answered, and every model call, tool call, and lookup in between. Evals tell you how your agent does on cases you chose. A source shows you what it does on the conversations that actually arrive.
Sources appear on their own. The first time a conversation arrives under a new name, Caliper starts a source for it, and it shows up under Sources in Caliper, in your workspace's file browser, and in search.
Three ways to send
- From your code. Send each conversation to
POST /1.0/caliper/tracesafter your agent replies. The OpenAI-style message list you already have is enough; tool calls in it become steps. - With OpenTelemetry. If your agent framework exports OpenTelemetry spans, point its exporter at Caliper. See Send traces with OpenTelemetry.
- From a Workbench flow. Publish a flow and every run of the published version arrives under the flow's name. Runs you start while testing and runs from Caliper evals stay out.
Sending needs a workspace API key with the Send agent traces scope (caliper:traces:write). That scope can only send, so it's safe to put in your agent's environment. Whatever you leave out of what you send isn't stored.
Over time
Open a source to see the last 7, 30, or 90 days:
- Traces, tool calls per trace, average time, cost per trace, and failures, each with a bar per day.
- Tools: for each tool, the share of that day's conversations that called it. This is where you see that your agent started reaching for web search twice as often last Tuesday.
- Lookups: the documents your agent's searches landed on most, by day, and how many lookups found nothing.
- Conversations: the conversations themselves, newest first.
Click any bar, tool, or document to narrow the conversation list to what's behind it. What you've picked stays in the page's address, so you can send someone the exact view.
Open a conversation to see what came in, what went out, and every step laid out on the conversation's own timeline. Open a step to see a tool's arguments and result, the model's reasoning, or what a lookup searched for and found.
Save conversations to a dataset
Save to dataset takes the conversations you've narrowed to and puts them in a new or existing dataset. Choose what they're for:
- Rate these replies: each item keeps what your agent said, tool calls included, for people to rate in a review.
- Score new versions against them: the agent's reply becomes the expected answer, for an eval.
- Write the right answers: only the questions are kept, for a spec.
Live datasets
Set New conversations to keep adding, and new conversations that match keep arriving in the dataset as they happen: every one, one in 10, or one in 100. A live dataset stops on its own once it has added 1,000 conversations. You'll find a source's live datasets on its page, each with a Stop button; what they've already added stays.
How long conversations are kept
Conversations are kept for the number of days your plan includes (see pricing). A source's page says how long. The daily totals behind the charts are kept for good, so a chart keeps its history after the conversations behind it are gone. Conversations you've saved into a dataset are copies and stay with the dataset.