Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to move an AI prototype into production

Reading time
6 minutes
Assumes
You have a working prototype
Updated
Sep 6, 2026

The prototype was answering a different question

A prototype answers "can this work." Production answers "does this work reliably, at volume, for inputs nobody anticipated, when a dependency is down, at a cost we can sustain, in a way we can fix at 3am."

That's not one gap, it's several, and they're mostly independent of model quality. Teams that treat the remaining work as "polish" underestimate it by about a factor of three, then ship late while claiming the model was the problem.

The reframe

The prototype proved the idea. Everything from here is about failure — what happens when inputs are weird, dependencies break, volume spikes, and quality drifts. None of it is visible while you're demoing the happy path.

Handle the inputs you didn't design for

Prototypes are tested on inputs their builder chose. Production traffic is chosen by everyone else.

You need defined behavior for empty input, enormous input, input in another language, input that's off-topic, and input containing instructions aimed at your system. Right now, in your prototype, most of these produce something undefined and occasionally something embarrassing.

Decide the behavior for each, then test that it happens. This is the single biggest source of week-one incidents and the cheapest thing on this list to fix.

Define what happens when something breaks

Every external dependency will fail. The model will time out, retrieval will return nothing, a tool will error, a rate limit will hit.

For each, decide: retry, fall back, or fail visibly. Then decide what the user sees, because the default is a spinner that never resolves or a stack trace.

Two specifics worth naming. Bound your retries — an unbounded retry loop turns a cheap operation expensive during exactly the incident you're already having. And make failure legible: "we couldn't reach that system, try again shortly" is a better product than a generic error and dramatically better than silence.

Get the cost model right at real volume

Prototype costs are misleading, because prototype volume is misleading and prototype inputs are shorter than real ones.

Work out cost per operation at production input sizes, multiply by realistic volume, and add retries. Then ask what happens at ten times that, because the scenario that produces a surprising bill is usually a loop or a spike rather than steady growth.

Put a ceiling in place before launch. A spend cap that stops the system is better than one you discover was needed.

Instrument what you'll need at 3am

The prototype logs nothing because you were watching it run. In production you're not.

Log the input, the output, which version produced it, timing per step, and cost. Trace a single request end to end. Without that, a report of "it gave a weird answer yesterday" is unanswerable, and you'll spend the incident reconstructing what happened instead of fixing it.

Alert on the things that indicate trouble rather than on individual errors: error rate, latency, cost per hour, and the volume in your uncertain lane. A sudden change in the last one is often the earliest signal that something upstream shifted.

Pin the version, and know how to roll back

A prototype runs the latest of everything. Production needs to know exactly what produced a given output, and needs to be able to go back.

Pin the version the running system uses, so a published change is a deliberate act. Keep the previous version deployable. And record which version produced each output, because the first question during an incident is whether the behavior changed and the second is what changed.

Model versions belong in that pinning too. A provider updating a model underneath you is a change to your system that you didn't make and can't see without recording what you called.

Decide the human's role before launch

Almost every production AI system has a person in it somewhere. Decide where deliberately rather than discovering it during an incident.

Who sees the uncertain cases, and how quickly. Who gets told when quality drops. What the manual path is when the system is unavailable — this one is routinely forgotten, and a support process with no fallback becomes an outage rather than a degradation.

Then check the volume assumption. A review step that works at prototype volume and needs six hours a day at production volume is not a design, it's a problem you've deferred.

Ship narrow, then widen

The prototype covered a broad case because breadth is what makes a demo convincing. Launch narrow, because narrow is what makes an incident survivable.

Pick a segment, a subset of case types, or a percentage of traffic. Run it with the pre-registered expectation you wrote before you started. Widen when the numbers hold and you've seen what the failures look like.

The temptation is to launch broadly because the prototype handled breadth. The prototype handled breadth on inputs you chose.

Common mistake

Treating launch as the end of evaluation. Your acceptance set was assembled before you had real traffic, which means it's missing the shapes real users produce. Sample production, add what surprises you, and treat those additions as a version boundary so scores stay comparable.

Before you turn it on for real users

0 of 7 checked