Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to write acceptance criteria for an AI feature

Reading time
6 minutes
Assumes
You're specifying an AI feature for someone to build
Updated
Sep 6, 2026

Why the usual format breaks

Conventional acceptance criteria are binary and per-case: given this input, the system does that. It works because software is deterministic — check it once and it holds.

An AI feature has no such property. The same question asked twice can produce different wording, different structure, occasionally a different conclusion. "The assistant answers refund questions correctly" is either untestable or trivially false, depending on how literally you read it.

Teams respond in one of two unhelpful ways. They write criteria so loose that anything passes, or they demand perfection on a class of case where perfection isn't available and the feature never ships.

The workable form is a threshold on a sample: on this set of cases, at least this share meet this standard, and these specific cases must never fail.

The shape of an AI acceptance criterion

A named set of cases, a written standard, a percentage, and a short list of absolute failures. All four, or it isn't testable.

Separate the rate from the floor

Two different kinds of criterion, and conflating them is the most common mistake.

Rate criteria apply to quality across a distribution. "At least 85% of responses in the acceptance set meet the rubric standard." These acknowledge variance and set a bar above it. The number should come from what a competent human achieves on the same set — if your support team gets 90% right, demanding 99% of a model is a decision to never ship.

Floor criteria are absolute and few. "Never states a policy that doesn't exist." "Never promises a refund." "Never produces advice in a regulated category." These aren't percentages, they're conditions, and a single violation in the acceptance set blocks release.

Keep the floor list short — five or fewer. A long absolute list is a wish, and it will either be quietly relaxed under deadline or block the feature forever.

Name the acceptance set in the criteria

A percentage means nothing without the set it applies to. "85% correct" on easy cases and on adversarial cases are different features.

Write the set into the criterion: what it contains, how many items, and who agreed it's representative. Freeze it before the build starts. A set assembled after the feature exists will be shaped, unconsciously and unavoidably, by what the feature does well.

The set should include failure modes deliberately — boundary cases, cases that should be refused, cases in the messy shapes real users produce. A set of well-formed happy-path questions will report high numbers and predict nothing.

Write the standard as behavior

Your criteria depend on a rubric, and a vague rubric makes a precise-looking percentage meaningless.

Criteria have to name behaviors that can be checked. Not "the response is helpful" but "the response answers the question asked, cites the policy it relies on, and offers a next step where one exists." The test is whether two people scoring the same output would produce the same verdict.

This is the part of specification work that transfers directly into evaluation. Time spent making the standard precise isn't overhead on the acceptance criteria — it's the thing that makes them mean anything, and it's reusable for every future version.

Include criteria for the unhappy paths

Three that get left out and cause most post-launch surprise.

What happens when it doesn't know. Specify the behavior — decline, hand off, ask a clarifying question — and make it a floor criterion. Without it you have accepted a feature that invents answers under uncertainty.

What happens when the input is malformed. Empty, truncated, in another language, containing instructions. Real traffic contains all of these in week one.

What happens when a dependency fails. Retrieval returns nothing, the model times out, a tool errors. Specify the user-visible behavior, because "undefined" is what ships otherwise.

Common mistake

Accepting on a demo. A demo is a sample of one, chosen by the person who built it, on an input they picked. It is evidence that the feature can work, which is not the question acceptance criteria exist to answer.

Say what happens after acceptance

An AI feature that met its criteria in March can drift by June — the model changes, the corpus grows, real traffic differs from your set.

Write the ongoing terms into the acceptance criteria. How often the set is rerun. Who looks at the result. What score triggers a response, and what that response is. This turns acceptance from a gate you pass once into a standard you hold, which is the only version that means anything for a system that changes underneath you.

Include the review cadence for the set itself. Real traffic will contain shapes your acceptance set doesn't, and the set should absorb them — in deliberate batches, treated as a version boundary, so scores stay comparable.

Before you hand the criteria over

0 of 6 checked