Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to tell whether a change caused the number to move

Reading time
6 minutes
Assumes
You shipped something and want to know if it worked
Updated
Sep 6, 2026

Why before-and-after is so weak

A pre-post comparison attributes everything that happened in the window to your change. Everything else that happened in that window is now inside your result.

Four confounders account for most of the damage.

Trend. The metric was already moving. Half the improvements teams claim are a continuation of a line that was already going that way.

Seasonality. Volume, mix, and behavior change by month, by quarter, by holiday. Comparing January to December measures the calendar.

Co-interventions. Almost nobody ships one thing. If you also retrained the team, rewrote the templates, and changed the routing, the AI is one of four candidate causes and probably not the largest.

Composition. The population changed. More enterprise customers, a different acquisition channel, a segment that churned. The metric moves without anything about your process changing.

None of this makes pre-post useless. It makes it a starting point that needs strengthening, and there are cheap ways to strengthen it.

The cheapest strong evidence

Hold something back. One team, one region, one segment, one queue left on the old way for a month. It is almost always rejected as too complicated and it is worth more than every statistical adjustment on this page combined.

Run a holdout even if it's imperfect

A partial, imperfect holdout beats a perfect before-and-after.

The objections are usually fairness ("why should one team not get the new thing") and complexity. Both are real and both are smaller than the cost of not knowing whether the thing works. Frame it as a staged rollout — which is operationally sensible anyway — and the fairness objection mostly evaporates.

Where a true holdout is impossible, look for a natural one. A region that gets it later. A case type outside scope. A period when it was switched off. These aren't randomized and they're far better than nothing, as long as you're honest that the groups may differ.

Use the pre-period shape, not just its level

If you can't hold anything back, at least compare against the right counterfactual.

Take enough pre-period data to see the trend — twelve weeks, not two. Fit the trend and extend it forward. Your comparison is now against where the metric was heading rather than against where it was, and that difference often accounts for most of the apparent effect.

Look at the variation too. If the metric swings 15% week to week, an 8% improvement is inside the noise and no amount of careful framing changes that. Better to say the effect was too small to detect than to present it and have somebody else notice.

Write the confounders down in advance

The move that changes how the eventual argument goes: list the things that could explain a change, before the change ships.

Seasonality. A planned launch. A pricing change. A hiring wave. Anything you know is coming that touches the same metric.

Written in advance, these are context and they make your analysis more credible. Written afterwards, they're excuses, and everyone in the room can tell the difference. This is the same mechanism that makes pre-registration work — the value is entirely in the timing.

Look at more than one number

Triangulation substitutes for control reasonably well.

If your change should reduce handle time, it should probably also show up in agent-reported effort, in the count of transfers, or in how long the queue is at peak. If handle time moved and none of the others did, the effect is either smaller than it looks or is being produced by something other than what you think.

Check the metric that should not have moved as well. A change to refund handling that also moved your billing metrics is telling you the effect isn't what you named.

Segment, and look for the shape you'd expect

A real causal effect usually has structure. A confound often doesn't.

If your change should help most on complex cases, check whether it did. If it should have no effect on a segment it doesn't touch, check that segment stayed flat. Where the effect appears exactly where the mechanism predicts and nowhere else, that's meaningfully stronger evidence than an aggregate move.

Where the effect is uniform across every segment including ones your change doesn't touch, you're probably looking at something else.

Common mistake

Bundling several changes and attributing the result to the AI component. If you shipped the model, new templates, and a process change together, you cannot separate them, and the honest claim is about the bundle. Teams that claim the component get caught when the next project ships the AI alone.

Say how confident you are, and in what

The output of this work is not a number. It's a claim with a strength attached.

"Handle time fell 11%, against a pre-period trend that was already falling about 3%, with no holdout, in a quarter where we also changed the templates. Our best estimate is 4 to 8 percentage points attributable, and we'd want a holdout before extending this to the rest of the org."

That paragraph is more useful than a confident single figure, because it lets somebody decide how much weight to put on it. It also survives the meeting, which a confident figure with a hidden confound does not.

Before you claim causation

0 of 6 checked