Home / Insights / When a pilot should not go to production

When a pilot should not go to production.

The three checks we run before an AI product is handed over.

14 August 2026 · 7 min read · Alex Novak

Most AI pilots impress in a demo and fail in production. The demo runs on ten hand-picked examples with someone watching. Production runs on ten thousand messy ones with nobody watching. The gap between those two is where budgets disappear.

Before we hand anything over, it passes three checks. A pilot that fails any of them stays a pilot, and we say so.

One: it has an evaluation suite, not a vibe

An evaluation suite is a fixed set of real historical cases with known correct answers, run automatically on every change. For the Nordwind claims workbench that was 1,200 historical claims. Without it you cannot tell whether a prompt change helped or quietly broke a category of edge case.

If a team cannot tell you their accuracy number and how it was measured, they do not have one.

Two: it knows what it does not know

Every output carries a confidence signal, and anything below the threshold routes to a person with the context attached. The threshold is a business decision, not an engineering one: how often are you willing to be wrong, and what does being wrong cost?

Systems that answer everything with equal confidence are the ones that erode trust fastest. One bad answer with no hedge undoes fifty good ones.

Three: someone owns it on Monday

Handover is not a zip file. It is documentation, a runbook, an on-call owner, and a person on the client side who has changed a prompt, run the evaluation and deployed it while we watched.

We build the 30-day support window in for exactly this. If nobody on your side has touched it by day 30, it is not adopted, and we would rather fix that than close the invoice.

AN
Alex Novak
Delivery director · Eazetech

Scopes the audit week, writes the plans and owns handover. Product lead on the Nordwind claims workbench and the Coastline order run.

Related reading

How to price the hours before you automate them.
Jun 2026 · 6 min · P. Shah
All insights

Related services

AI development

Agents and LLM features built with the evaluation suite and the confidence thresholds described here.

Read more

Software development

The product and interface work around the model, which is usually what decides adoption.

Read more

AI automation

The same three checks applied to workflows rather than products, on n8n, Make, Zapier or code.

Read more

Have a project where this applies?

Send the workflow or the pilot that is stuck. Alex Novak or one of the delivery team reads it and replies within one business day, with a rough estimate attached.

Get an estimate