The Production Gap

The Production Gap

There's a meeting happening in hundreds of companies this quarter, and I've sat in more versions of it than I can count. It usually goes something like this.

Twelve months ago, the board asked for an AI strategy. The company responded the way most companies did: it launched pilots. A copilot for the service team. A document tool for legal. A forecasting model for the planners. The demos were very impressive. People love them.

And then someone in the room, usually someone from finance, asks the uncomfortable question: which line of the P&L did this change?

Silence.

And then comes the part almost everyone gets wrong. The room concludes that the technology isn't ready yet. The models need another year, the vendors oversold, and the prudent move is, obviously, to wait.

That conclusion is almost always backwards.

In most failed AI initiatives, the model is the single best-performing component of the entire effort. What failed is everything the organization was supposed to build around it. I call this the Production Gap: the distance between a demo that works and a system that changes how your business runs and shows up in a number your CFO cares about.

Most companies never cross it. Almost none of them fail for the reason they think.

The best-documented failure in modern business

The Production Gap is no longer an anecdote. It may be the most consistently replicated finding in enterprise technology.

In that famous research piece from last year, MIT's Project NANDA studied more than 300 enterprise AI deployments and found that 95% of generative AI pilots deliver no measurable P&L impact. In another one, RAND, analyzing 65 initiatives over three years, put the failure rate at just over 80%. S&P Global found that by mid-2025, 42% of companies had abandoned most of their AI initiatives, up from 17% only a year earlier. And the abandoned projects weren't cheap experiments: the average one cost $7.2 million and ran for 14 months before shutdown. Long enough to burn real money and real careers. Short enough to leave nothing behind.

On the other side of the gap, the largest annual surveys of AI adoption find only about 6% of companies qualify as AI high performers: organizations getting meaningful earnings impact at scale.

When every major research house measures the same thing and finds the same shape, you are not looking at bad luck or just bad vendors. You are looking at structure.

The tell everyone misses

Here's the detail that should change how you read every AI status update you receive: the failure never shows up in the demo.

The demo works. It always works. The failure shows up nine to fourteen months later, in the messy data, the security review, the workflow nobody changed - when there's an actual workflow that most people follow - and the team that quietly went back to the old way of doing things.

Think about what that means. If technology were the constraint, the failure would appear there. It doesn't. Which tells you the constraint is somewhere else entirely.

You can think of a pilot as a concept car. Any automaker can hand-build one beautiful car for the auto show. Building 300,000 of them is a completely different ballgame. It requires supply chains, quality control, dealer networks, and service, and mastering it is what actually constitutes the business. For three years now, companies have been celebrating concept cars and wondering why there's no revenue.

So if the model isn't the problem, what is? Every major research group that has dissected the failures, from MIT to Gartner, finds the same three causes, in the same order, every time.

Reason one: The work was never redesigned

The most common failure is also the least visible one: AI was dropped into processes and products, many of which were designed decades before AI existed, and at that point, AI was just an afterthought bolted on to show the board that something was happening.

The tool drafts the contract in four minutes instead of four hours. Then the contract sits for three days waiting for the same two approvals it always did. Measured end-to-end, the cycle time barely moved. So nothing moved in the P&L either. You didn't transform the process. You gave a 1995 process a very fast typist.

The research on this point is unambiguous: the single biggest driver of AI's earnings impact is not model selection, not infrastructure, not even data quality. It is workflow redesign. The companies getting real earnings impact are roughly three times more likely than everyone else to have fundamentally redesigned how the work flows before deploying the technology.

The symptom a board member can check in one question: since we deployed AI, which process map has changed? Which handoff was eliminated? Which role was redefined? If the answer is none, the pilot was just a very expensive ornament.

Reason two: The pilot was built as a demo, and demos don't survive contact with your company

A demo lives in a sandbox: clean data, a friendly user, no edge cases, no auditor in the room. Production lives in your company: the ERP with the fields sales never filled in, the security review, the compliance sign-off, the exception that arrives at 4:52 p.m. on the last day of the quarter.

I've watched this up close: a company running hundreds of AI initiatives at once, demos everywhere, and a production record close to zero. Architectures designed as if decades of legacy systems didn't exist. Systems that fell over in production because they hit the model vendor's rate limits with no fallback plan. The model never gave a wrong answer. Nobody had planned for what happens when an entire company starts asking at once.

The distance between those two environments is enormous, and it is almost always underestimated, especially by internal teams building for the first time. Only about 12% of enterprise data is actually AI-ready, according to Informatica, some other sources go even lower at 7%. Gartner expects 60% of AI projects that lack AI-ready data to be abandoned through 2026. Most of the real work is the unglamorous engineering around the model: data plumbing, integration, permissions, exception handling, and reliability. It's also the part with no applause or crazy wow factors.

This is where the build-versus-partner question stops being philosophical. MIT's data shows external partnerships reach deployment about twice as often as internal builds: 67% versus 33%. Not because outside teams have smarter engineers, but because they've already made the expensive mistakes on someone else's budget.

The symptom: eight months in (the S&P Global average from prototype to production), and everyone still calls it a pilot.

Reason three: Nobody uses it

Two years ago, the executive fear was that the technology wouldn't work. In 2026, the fear has inverted: the technology works, and nobody uses it.

This is now the top-ranked blocker in the data. Gartner's CIO survey found that 71% of CIOs cite workforce adoption as the number one obstacle to AI value, ahead of cost, security, and data quality. Meanwhile, Gallup found that only 15% of employees say their company has communicated a clear AI strategy at all. Executives assume adoption will follow deployment. It never does.

A $2 million AI system with 15% adoption isn't a technology success. It's a $1.7 million write-off.

Adoption fails because it is treated as a training session at the end of the project rather than a discipline in its own right, with its own owner, budget, and number. And usage is a number. If nobody reports it to the executive team every week, then nobody is managing it.

Three reasons. Zero of them technical.

Read those three failures again. Choosing and redesigning the work. Building for production instead of applause. Getting human beings to change how they operate.

Not one is a technology problem. All three are management problems.

Researchers who have measured where the effort actually goes have put numbers to this, and they're worth memorizing: roughly 10% of the work is algorithms, 20% is data and technology, and 70% is people and processes. Which yields the shortest honest description of the Production Gap I can offer:

The model is 10% of the work, and it's the only 10% most companies actually finish.

We're about to run the experiment again, at triple the stakes

Here is why this matters more right now than it did a year ago.

According to a January survey of more than 2,300 executives, 90% of CEOs believe AI agents will produce measurable returns this year. Companies plan to roughly double their AI spend to about 1.7% of revenue. Half of CEOs believe their own tenure depends on getting AI right. Meanwhile, Gartner forecasts that more than 40% of agentic AI projects will be canceled by the end of 2027.

Both of those things are about to be true at the same time, because agents don't shrink the Production Gap. They widen it. Each of the three failures is amplified.

Bolt an agent onto an unredesigned process, and you haven't fixed the dysfunction. You've automated it at speed, sometimes without a human in the loop to pause and ask whether something looks odd. An agent doesn't just suggest, it often acts, so the production engineering around it (permissions, exceptions, auditability) goes from important to existential. And an agent your people don't trust gets switched off faster than any chatbot ever did.

Same movie, bigger budget. The sequel has higher production values and the exact same script.

What the 5% actually do

The companies on the right side of the gap aren't luckier, nor are they necessarily bigger. MIT found mid-market companies take a successful pilot to production in about 90 days, versus nine months at large enterprises. The binding constraint is organizational, and decisive organizations clear it faster.

What the 5% share is that they run three disciplines, in order, as one program.

They redesign the work first. Fewer initiatives, chosen for where the money is rather than how good the demo looks. Each one is attached to a named owner and a number that the CFO already tracks, with kill criteria written down before anyone writes code.

They build for Tuesday afternoon, not for the demo. Data readiness, integration, security, and exception handling are treated as the project, not the footnote. And they make the build-versus-partner decision on evidence rather than pride, remembering that the odds run two to one against going it alone.

They run adoption as a product. Redesigned roles, trained managers, internal champions, and usage metrics religiously reviewed on the agreed-upon cadence with the same seriousness as revenue. They measure fluency, not attendance.

And critically, they do all three together. Companies that integrate strategy, technology, and adoption are 3.6 times more likely to achieve transformative impact than companies that treat them as separate workstreams. The three disciplines aren't a menu. Each one exists to close a specific failure mode, and any one of them, skipped, is sufficient to kill the initiative.

The good news hiding in the bad numbers

If the Production Gap were a technology problem, waiting would be rational. You'd let the models mature, buy the finished product later, and skip the pain.

But it isn't a technology problem, and that changes the calculus completely. The disciplines that separate the 5% from the 95% are choosing well, building for production, and changing how people work. Those are organizational muscles. They don't arrive with the next model release. They are built through repetition, or not at all. A company that waits a year doesn't start next year at the same starting line, it starts a year behind organizations that spent that year building the muscles.

The technology has been ready for two years. The question was never the model.

The gap between the 95 and the 5 isn't measured in model quality. It's measured in management.

Ready to turn AI investment into real outcomes?

Book Time With An AI Strategist

View other posts