Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to evaluate an AI feature when you have no test data

Reading time
6 minutes
Assumes
You're building something new with nothing to test against
Updated
Sep 6, 2026

Four sources, and you'll use several

You need thirty to fifty items to start, with a standard attached to each. That's less than it sounds and there are four ways to get there.

Adjacent human work. Whatever your feature will do, somebody currently does it some other way. Support tickets and their replies. Emails your team has sent. Documents somebody wrote by hand. This is the highest-quality source available before launch, because the standard is already attached — a human decided what good looked like at the time.

Written cases from the people who know. Sit with two domain experts for an hour and have them write the cases they'd worry about. This is faster than it sounds and produces exactly the edge cases you'd otherwise discover in production.

Synthetic variations. Take real cases and vary them systematically — different phrasing, different customer types, a language you support, a malformed version. Cheap, and the only source that reliably covers weird inputs.

Whatever traffic you can get. An internal alpha, a friendly customer, your own team using it for a week. Small volumes of real traffic beat large volumes of anything else for representativeness.

Start with the adjacent work

The email your support lead sent last Tuesday is a test case with a human-approved answer already attached. Most teams have thousands of these and overlook them because they don't look like a dataset.

Mine the existing work properly

The adjacent-work source is the best one and the one most often done badly.

Pull a sample that isn't just the recent stuff — recency correlates with whatever you changed most recently. Deliberately include the cases that were escalated, the ones that took longest, and the ones where a customer came back.

Be careful about what the human output actually represents. A support reply is what one person sent under time pressure, not necessarily the ideal answer. Use it as a captured output to be rated rather than as an expected output to be matched, at least until somebody has reviewed it. The distinction matters: one asks "how good was this," the other asks "does the new version match this standard."

Strip identifying details before it becomes a test fixture. A dataset outlives the project and gets shared more widely than the original records.

Get experts to write the hard cases

An hour with two people who do the work produces better edge cases than a week of your own imagination.

Ask for specific things rather than "some examples." The case you'd escalate. The one where the policy is ambiguous. The one where somebody is angry and technically wrong. The one where the right answer is no. The one a new starter always gets wrong.

Have them write the expected output too, and have the other expert check it. Disagreement between two experts on the same case is not a problem to resolve quickly — it's the discovery that your standard is undefined, and it's much cheaper to find now.

Use synthetic data for coverage, not for the core

Synthetic cases are good at one thing: filling gaps you can enumerate.

Take twenty real cases and generate variations along dimensions you care about — terse versus verbose, formal versus casual, a supported second language, missing information, an angry framing. You now have coverage of input shapes your real sample happened not to contain.

What synthetic data cannot do is tell you what real inputs look like. Generated cases share a distribution with whatever generated them, which is smoother and more grammatical than reality. A suite that is mostly synthetic will report scores that don't survive contact with users.

Keep it as a minority of the set and label which items are synthetic, so you can report scores with and without them.

Once you have traffic, replace as fast as you can

Everything above is scaffolding. Real traffic is better than all of it and you should start swapping the moment you have any.

Sample production continuously. Prioritize the cases that went wrong, the ones users rephrased or repeated, and anything where somebody complained. Add them in deliberate batches, treated as a version boundary so scores stay comparable across the change.

Retire synthetic items as real equivalents arrive. Keep expert-written edge cases indefinitely — those tend to stay rare in traffic and remain the hardest part of the suite.

Be explicit about what a pre-launch score means

The number you get from a bootstrapped set is real and it is narrower than it looks. Say so when you report it.

It tells you the feature handles the cases you thought of. It does not tell you how it performs on real traffic, because your set isn't drawn from real traffic. It's a floor, not an estimate.

The honest framing: "on 45 cases assembled from past support replies and expert review, 82% met the standard. This set doesn't represent real traffic distribution and we'll rebuild it from production within a month of launch." That's a sentence somebody can make a decision with.

Common mistake

Building the set by running the feature and keeping what it did well. It happens by accident — you're testing, you save the interesting ones, and the interesting ones are the ones that worked. Assemble the set from cases first, then run, never the other way round.

Before you call it evaluated

0 of 6 checked