Every AI project I have seen stall had a successful prototype behind it. That is not a coincidence, and it is not because the prototypes were dishonest. It is because almost nobody writes down what counts as passing before the test starts, so when the demo runs, the room decides afterwards what it was meant to prove. Under those conditions a prototype always passes. It has never in the history of the practice failed.
The fix is unglamorous. Before anyone opens a laptop, you settle eight things, and the pass threshold is the one that matters most and gets skipped most.
Start with the assumption. Ask the room: if this build fails, what will have been the reason? There is normally one belief the whole thing rests on. That the system can read our inputs accurately enough. That people will actually use it. That the data is clean enough to work from. Name it, and then name the belief underneath it, which is usually that a person will catch what the system gets wrong.
That second one is where prototypes go quietly wrong. A reviewer who rubber-stamps produces excellent acceptance numbers and worse output than the process you had before. Your metrics look wonderful right up until the first challenged result. So the prototype has to be designed to detect that, and the standard design does not.
The way you detect it is a seeded-error arm. Alongside the real cases, you plant a handful of outputs containing a deliberate, plausible mistake, and you measure how many the reviewers catch. Make it an independent gate: fail it and the prototype fails regardless of how good the acceptance score was. It is the single most informative thing you can add to a pilot and it costs almost nothing, because the errors are cheap to construct and you were running the test anyway.
Then the rest. Scope, in and out, so you are testing one path rather than an impression. Named participants who do the work, not a demo audience. The data it runs on, drawn as a random sample rather than hand-picked, because hand-picked cases are how you accidentally test the easy ones. A start date and an end date, fixed before it begins. And the scoring method, which means who scores it, whether they are independent of the people who built or edited it, and whether the measurement is instrumented rather than self-reported. If the people whose hours the build saves are also the people scoring whether it worked, you have a result you cannot use.
Then two decisions nobody wants to make in advance, which is exactly why you make them in advance. What we do if it fails, stated as kill, narrow and retest once, or re-scope. And whether the code is disposable or becomes phase one, said out loud before the first line is generated, because that decision silently gets made by default and then everybody lives with it for three years.
Last, cost the test itself, separately from the build. A stage-gated process approves an increment, not a programme. The committee should be releasing the price of the experiment, against the full estimate, rather than approving a whole build at the point of least information. That single change does more for capital discipline than any amount of governance.
One more thing on honesty. If you test forty cases and thirty pass, that is a seventy-five percent result with a confidence interval running from roughly fifty-nine to eighty-seven percent. It is evidence, not proof, and it should be treated as a decision to continue rather than a measurement of accuracy. Say that in the write-up. It costs you nothing and it stops the number being quoted for the next two years as if it were established fact.
None of this is about being rigorous for its own sake. It is about the difference between spending a small amount to learn something and spending a large amount to confirm what you already wanted to believe. Write the number down first. Define it, measure it, own it, close it. A test you cannot fail is not a test.
Related KeyDelta Services

Russ Reeder
Founder & CEO, KeyDelta | Forbes Technology Council
30+ years scaling technology companies as a CEO, COO, and operator across Oracle, GoDaddy, OVHcloud, Infrascale, Netrix Global, and XTIUM. Founder of Rightsline (Disney+, Hulu, Sony). Forbes Technology Council member. HBS Executive Education. Russ advises CEOs, PE-backed leadership, and management teams on execution clarity through the VOOCS operating system.
More on Execution
Twelve people saving twenty minutes a day is not a headcount
The most common benefit in an AI business case is also the most commonly overstated. Here is the question that separates hours you can bank from hours that just make people less swamped.
Read article →ExecutionOne payback number hides the argument
Give a committee a single payback figure and you have told them what to believe instead of what to decide. Three numbers at three levels of evidence make the real argument visible.
Read article →Want to discuss these ideas?
If your team is navigating execution challenges, we should talk.
Book a 30-minute call