HomeServicesPortfolioAboutContactBlogCareers
Book a call
AI Engineering

How to Run an AI Pilot That Produces a Real Go/No-Go Number

September 2026 · ISTRALLEN Team

A pilot replaces the biggest guess, not "does it work"

The point of an AI pilot isn't to see whether the technology functions in a vague sense. It's to replace the single number the business case rests on — usually a containment rate, a resolution rate, a detection rate, or an uplift — with something you measured instead of assumed. Everything about how to run an AI pilot follows from that.

Pick one narrow slice

One intent group ("where's my order" calls). One store region. One query type. One transaction segment. Not the whole thing. A narrow slice is faster to stand up, cheaper to run, and easier to measure cleanly — and it's enough to get the number you need.

Define go/no-go before you start

Write down the metric and the threshold that make this a yes: "at least 40% of status calls contained, with CSAT at or above 4.2." Get the people who own the budget and the operation to sign off on it. Deciding the bar after you see the result is how pilots turn into rationalisations.

Measure against a control

A holdout slice that stays on the old process, over the same period — or a clean pre-period if you can't split traffic. Without a control you'll credit the pilot for seasonality, a promo, or a staffing change.

Run it long enough

Cover one full cycle of whatever you're measuring: a restock cycle for shelf availability, a chargeback window for fraud, a promo for search. Usually four to eight weeks. A one-week readout is noise.

Cost the run side too

Get the real per-unit run cost from actual bills during the pilot — per conversation, per minute, per search, per transaction — not a vendor quote. That number scales into the full deployment and it's the one people underestimate.

A worked pilot

A lender wants a voice agent for first-line calls. The pilot: six weeks, one intent group — application-status calls only. A holdout keeps a fraction of those calls on the existing queue. The go/no-go, agreed with ops and finance before the start: at least 40% of status calls contained end to end, with CSAT no lower than 4.2, and a per-minute run cost from the provider's actual bills that keeps the projected full-scope cost under the labour it displaces. At the end you have three real numbers — containment, CSAT, cost per minute — each with a range, and a decision that isn't a guess.

The output

A number with a confidence range and its assumptions stated, that maps directly onto the full-scope investment decision. The pipeline write-ups in our portfolio show the kind of measured figures each build was scoped against before the full commitment.

Don't let the pilot become the product

A pilot is built to answer a question, not to run forever. It's fine for it to be rougher than a production system — hard-coded scope, a thinner escalation path, manual monitoring. The risk is a "successful" pilot quietly staying in production because nobody scheduled the real build. Decide up front what happens the day the pilot ends: scale it properly, or turn it off.

Where this stops being right

  • A small, reversible change you can just ship behind a flag and watch.
  • A slice that isn't representative of the whole — then the number doesn't generalise, and you should say so rather than extrapolate.
  • A pilot that costs more than the information is worth — rare, but a very expensive pilot for a cheap decision isn't worth running.

FAQ

How long should a pilot run? Long enough to span one full cycle of what you're measuring — a chargeback window, a restock cycle, a promo. Usually four to eight weeks.

What's the one thing to decide up front? The go/no-go metric and its threshold, signed off by the people who own the budget and the operation.

Pilot or shadow mode? Shadow mode measures a component safely with no behaviour change. A pilot measures a real deployment on a narrow slice with a decision attached. Often you do shadow first to prove it's safe, then a pilot to prove it's worth building at full scope.

ISTRALLEN scopes builds against a measured pilot slice before the full commitment — see what we do.

See it in production
Services → Portfolio →
← All articles