For product teams
The feature shipped. Whether it worked is still an opinion.
Set the standard before you ship, in your team’s own words. Then score every change against it. When something slips, you see it on the next run.
against the run before
Tone held. Right urgency has been sliding since the model changed underneath it, and the feature still demos fine.
A Caliper per-criterion trend for a ticket-triage eval: six runs down the side, four criteria across the top, each cell the score that criterion got in that run. Tone and Nothing invented hold above 90 percent throughout, and Names the policy sits around 70 percent the whole time. Right urgency falls from 88 percent in the first run to 41 percent in the sixth, which drags the overall mean down with it.
The three ways an AI feature goes wrong after launch
It demos well and fails on the long tail
The cases that embarrass you never show up in the demo. They arrive from an angry customer with three problems in one message, and you meet them in public unless you go looking first.
It works, then quietly stops
Nothing in your release notes changed, but the model underneath did, or someone rewrote a prompt. Quality drifts and the first signal is a support ticket.
Nobody can say whether it paid off
Six months in, support says it's better and sales says it's worse, and the metric everyone quotes was chosen after the numbers came in.
When the AI is in your product
Build it where your team can argue with it, and score it against what they agreed
The feature your customers touch gets designed where your whole team can see it, and measured the way the rest of your software is.
- 01
Design the behavior where your team can see it
The feature's logic is a flow on a canvas, not a prompt buried in a pull request. Designers and support leads can read it, argue with it, and try it the way a customer would.

- 02
Agree on what good means before launch
Your experts judge real cases. Their answers become a dataset and a rubric — the bar the feature has to clear, written down while you can still change it cheaply.

- 03
A dropping score stops the release
Caliper scores every pull request and every model you try against that bar, and a drop blocks the merge the way a broken test does. Your customers never see the version that slipped.
Swap triage model to the new release#482- build1m 12s
- unit tests284 passed
- caliper / ticket-triage0.74 of 0.80
Caliper · 47 casesBelow threshold- Right queue
- 0.94
- Tone
- 0.89
- Right urgency
- 0.41 ↓0.51
Eleven cases that passed last release now fail, and nine of them are the ones your support leads flagged as urgent. The merge is blocked.
- 04
Let real use feed the next version
Thumbs and comments from your own product land in the same dataset, so the cases you argue about next quarter come from customers rather than from a workshop.

The experts setting that bar don’t need accounts, and won’t be learning anything — they answer from a link. That’s the page to send them.
When the AI changes how you work
Pick the change on evidence, and settle it in public
- 01
Find where the time actually goes
Interview the people running a process and map what they tell you. What you work on next comes off that map, instead of off whatever the loudest team asked for.

- 02
Say what you expect, before you roll it out
Name the numbers that should move and the ones that mustn't suffer, and write them down before launch. Nobody can renegotiate the verdict once the results are in.

- 03
Settle it honestly, and keep the lesson
The numbers come in from wherever they already live, and you settle the call as confirmed, missed or mixed. You write the lesson then, while you still remember why, and it stays attached to the decision.

What would you have to see to call it a success?
If that question is easier to answer after launch than before it, that’s the thing worth fixing first. It’s usually a short conversation.