Skip to main content

For product teams

The feature shipped. Whether it worked is still an opinion.

Set the standard before you ship, in your team’s own words. Then score every change against it. When something slips, you see it on the next run.

ticket triage · per-criterion trend6 runs
74%↘ -2 points
against the run before

Tone held. Right urgency has been sliding since the model changed underneath it, and the feature still demos fine.

A Caliper per-criterion trend for a ticket-triage eval: six runs down the side, four criteria across the top, each cell the score that criterion got in that run. Tone and Nothing invented hold above 90 percent throughout, and Names the policy sits around 70 percent the whole time. Right urgency falls from 88 percent in the first run to 41 percent in the sixth, which drags the overall mean down with it.

Caliper
Workbench
Compass
Ledger

The three ways an AI feature goes wrong after launch

It demos well and fails on the long tail

The cases that embarrass you never show up in the demo. They arrive from an angry customer with three problems in one message, and you meet them in public unless you go looking first.

It works, then quietly stops

Nothing in your release notes changed, but the model underneath did, or someone rewrote a prompt. Quality drifts and the first signal is a support ticket.

Nobody can say whether it paid off

Six months in, support says it's better and sales says it's worse, and the metric everyone quotes was chosen after the numbers came in.

When the AI is in your product

Build it where your team can argue with it, and score it against what they agreed

The feature your customers touch gets designed where your whole team can see it, and measured the way the rest of your software is.

  1. 01

    Design the behavior where your team can see it

    The feature's logic is a flow on a canvas, not a prompt buried in a pull request. Designers and support leads can read it, argue with it, and try it the way a customer would.

    A Workbench flow open on the canvas with the test panel beside it, mid-conversation — readable enough that a non-engineer could follow the path
  2. 02

    Agree on what good means before launch

    Your experts judge real cases. Their answers become a dataset and a rubric — the bar the feature has to clear, written down while you can still change it cheaply.

    A contributor writing what a good answer looks like, one scenario at a time, from a share link with no account
  3. 03

    A dropping score stops the release

    Caliper scores every pull request and every model you try against that bar, and a drop blocks the merge the way a broken test does. Your customers never see the version that slipped.

    Swap triage model to the new release#482
    • build1m 12s
    • unit tests284 passed
    • caliper / ticket-triage0.74 of 0.80
    Caliper · 47 casesBelow threshold
    Right queue
    0.94
    Tone
    0.89
    Right urgency
    0.41 ↓0.51

    Eleven cases that passed last release now fail, and nine of them are the ones your support leads flagged as urgent. The merge is blocked.

  4. 04

    Let real use feed the next version

    Thumbs and comments from your own product land in the same dataset, so the cases you argue about next quarter come from customers rather than from a workshop.

    A Caliper dataset filling with feedback from a live product — thumbs and comments on real answers, with the source visible

The experts setting that bar don’t need accounts, and won’t be learning anything — they answer from a link. That’s the page to send them.

When the AI changes how you work

Pick the change on evidence, and settle it in public

  1. 01

    Find where the time actually goes

    Interview the people running a process and map what they tell you. What you work on next comes off that map, instead of off whatever the loudest team asked for.

    A Compass workflow page with its people, systems and pain points connected — the shape of a real process, not a tidy demo one
  2. 02

    Say what you expect, before you roll it out

    Name the numbers that should move and the ones that mustn't suffer, and write them down before launch. Nobody can renegotiate the verdict once the results are in.

    The Ledger expectation form mid-write: a reach claim and a hold claim on the same decision, before any evidence exists
  3. 03

    Settle it honestly, and keep the lesson

    The numbers come in from wherever they already live, and you settle the call as confirmed, missed or mixed. You write the lesson then, while you still remember why, and it stays attached to the decision.

    A Ledger metric with its readings, the feed they arrive from, and a decision watching it against a target

What would you have to see to call it a success?

If that question is easier to answer after launch than before it, that’s the thing worth fixing first. It’s usually a short conversation.