Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to run a design review on an AI feature

Reading time
6 minutes
Assumes
You have a working draft to get feedback on
Updated
Sep 6, 2026

Reviewing an artifact you can't see

Design review works for interfaces because the artifact is visible. Everyone looks at the same screen and reacts.

An AI feature has no such surface. What you're reviewing is a behavior across a space of inputs, and any single output is one sample from it. Show a stakeholder one good answer and they'll approve a feature that fails on a case neither of you tried. Show them one bad answer and they'll reject something that works.

The review has to be structured around the space rather than a sample, which changes both what you put in front of people and what you ask them for.

Bring cases, not screenshots

The unit of review is a case somebody recognizes from their own work. Ten of those produce more usable feedback than an hour of discussion about the concept.

Choose reviewers who own the consequences

The useful reviewers are the ones who will handle the fallout: the person whose queue receives the output, the one who fields the complaint, the one accountable if it's wrong.

That's a different list from the project stakeholders. Someone who has never worked a ticket will comment on tone. Someone who works tickets daily will tell you the answer is technically correct and will generate a follow-up question every time, which is the finding you needed.

Three or four reviewers is enough, and they should disagree with each other. A review panel that agrees on everything is telling you about panel composition rather than about the feature.

Give them real cases, including the hard ones

Assemble ten to fifteen cases before the review, and choose them deliberately.

Include cases that currently go wrong, because the feature's handling of failure is the part most worth reviewing. Include boundary cases where the right answer is genuinely arguable. Include cases where the correct behavior is to decline — reviewers routinely have strong opinions here and rarely get asked.

Run them beforehand and bring the actual output. A review where the feature is driven live becomes a demo, and demos get demo feedback.

Ask about the output, not the concept

The question that produces nothing: "what do you think of this?" You'll get opinions about AI.

The questions that produce something are all specific and all about a case in front of them. Would you send this? What would you change? What's missing that a customer would need? What would you do with this case?

That last one is the highest-yield question in the set, because the reviewer answers from their own practice, and the gap between what they'd do and what the system did is the actual finding.

Sort feedback by what it actually is

Reviewer feedback arrives mixed together, and the sorting is what makes it actionable. Four kinds, each going somewhere different.

A definition. "We don't call it that, we call it a service credit." This is missing knowledge — it belongs in the material the feature draws on, not in a prompt tweak.

A criterion. "It should always say what happens next." This is a standard. It belongs in your rubric, so every future version is checked against it, rather than being fixed once and forgotten.

An expectation. "If this works, escalations should drop." This is a prediction. Record it before you ship, and settle it later.

A configuration or behavior change. "It shouldn't offer a callback outside business hours." This is the only kind that's a direct edit.

Most teams treat all four as the fourth, which is why review feedback produces a long list of prompt edits and no durable improvement. A criterion turned into a one-off fix will regress the next time somebody rewrites the prompt.

Ask one clarifying question when you need to

Reviewer feedback is often ambiguous in a way that matters. "This is too formal" could mean the register, the length, the greeting, or the absence of the reviewer's name.

Ask once, in the moment, while they're still in the case. A clarification asked a week later in a follow-up email gets a reconstructed answer, and usually gets no answer at all.

Keep it to one question. A reviewer who is asked three follow-ups per piece of feedback learns to leave less feedback, which is the opposite of what you want.

Close the loop visibly

The single biggest determinant of whether reviewers engage a second time is whether anything happened after the first.

Go back with what changed, what didn't, and why. The "didn't" cases matter most. A reviewer whose suggestion was declined with a reason stays engaged. One whose suggestion vanished without comment stops answering, and you won't find out why.

Then run the same cases again and show the difference. That's the artifact worth producing: not a changelog, but the same ten cases before and after, which lets reviewers see the effect of their own feedback.

Common mistake

Treating a design review as an approval gate. Reviewers who believe they're approving become conservative and vague, because approval carries responsibility. Frame it as improving a draft that will ship in some form, and the feedback gets concrete.

Before you run the review

0 of 6 checked