The Pilot Paradox, Two Years On

Rick Pozniak
August 28, 2026
8 minute read
Executives cut a ribbon on a small model bridge over a puddle while employees improvise a plank walkway across a real chasm behind them.

This is a revision of Avoiding the technology 'Pilot Paradox', first published in April 2024. The original is still up, and I would rather you could check what I actually said two years ago. What follows is what the AI pilot wave changed, including the one thing in the original I now think was wrong.

Why I am rewriting this

In April 2024 I published a piece about what I called the Pilot Paradox: the fact that a pilot, designed to be the first step toward a solution, so often turns out to be the last. I had watched too many technology projects stall somewhere between "start small" and "scale fast," and I wanted to name the pattern.

Two years on, that article holds up better than I would like. The mechanics I described are still the mechanics. What has changed is the volume, and what the volume has revealed.

When I wrote it, the AI avalanche was a passing reference, a coming pressure that might push executives toward high-visibility, low-impact projects. It arrived. Almost every organization I now speak with has run an AI pilot. Most have run several. And a large share of them are sitting exactly where the original article said they would sit: successful, well liked, and going nowhere.

That is enough new evidence to justify a rewrite. It also forced me to change my mind about one thing, and that is the part worth reading.

What I got wrong

The 2024 article listed five factors that decide whether a pilot scales: strategic alignment, comprehensive planning, change management, scaling strategy, and continuous evaluation. Change management sat third on that list, a peer of the other four.

That weighting was wrong. Change management, or more precisely the capability of the people who have to use the thing, is not one factor among five. In the AI era it is the factor. The other four are conditions you need to get right. This one is the constraint that decides the outcome.

I did not arrive at that view quickly, and I would not state it that flatly if the evidence were still ambiguous. It is not.

What the AI pilots have taught us

Three sets of numbers describe the current state, and none of them describes a technology failure.

McKinsey's State of AI survey, published August 2026 across 1,719 leaders worldwide, found that nearly two thirds of organizations have not yet begun scaling AI across the enterprise. Only 37 percent attribute any EBIT impact at all to AI, and just 6 percent qualify as high performers, meaning AI accounts for at least 5 percent of earnings. That is the largest technology investment wave in corporate history, and the scaling gap has barely moved since I first wrote about it.

MIT's Project NANDA study, published in 2025, put it more bluntly. Of the generative AI pilots it examined, 95 percent produced no measurable profit-and-loss return. That figure got quoted everywhere and misread almost everywhere, so it is worth being precise: it does not mean the technology failed. It means the pilots never changed the work they were bought to change.

Gartner forecast in June 2025 that more than 40 percent of agentic AI projects would be cancelled by the end of 2027, on escalating costs, unclear business value and inadequate risk controls. Again, not a technology failure. A business case failure.

In 2024 I wrote that 40 percent of digital and AI transformations stalled at the scaling phase, and guessed the number had climbed. It has. What I did not anticipate was the reason.

The finding that changed my mind

Buried inside that MIT study is the number that actually matters. While 95 percent of official pilots delivered nothing measurable, the researchers found that roughly 90 percent of employees were already using personal AI tools to do their jobs, at organizations where only about 40 percent had any sanctioned AI subscription.

Read those two findings side by side. The company's AI pilot failed. The company's AI adoption succeeded. They were not the same project, and only one of them had a budget.

Your real AI pilot is already running. Nobody authorized it, nobody scoped it, nobody is measuring it, and it is going considerably better than the one on the steering committee agenda.

That is not a software story. It is a story about people getting out ahead of their employer, without training, without governance, and without anyone having measured what they can actually do.

The floor problem

If capability is the constraint, the obvious question is what capability exists. That is the question our benchmarking work was built to answer. It maps workforce AI competency onto seven levels, from the Unaware at Level 1 to the AI Developer at Level 7.

As of the April 2026 refresh, roughly 80 percent of the US workforce sits at Levels 1 to 3: unaware, aware but not doing, or using AI casually in ways that have not changed their work. In Canada that figure is 82 percent, in the EU 85 percent, and globally 86 percent.

The picture is improving, and quickly. In the US, the share with no AI exposure at all fell from 39 percent in September 2025 to 28 percent in April 2026, and the fastest-growing group is Level 3, up six points over the same period. People are picking this up on their own.

They are also picking it up alone. The Future Skills Centre found that 44 percent of Canadians using AI at work had received no formal training for it. Gallup's Q1 2026 workplace survey found half of US employees now use AI at work, with 13 percent using it daily.

Now hold that against a pilot. Ten enthusiastic volunteers run a controlled test and it works beautifully, because it was always going to. Then the tool is handed to a workforce where four in five people sit at Level 3 or below, and half of those already using AI were never taught how. Nothing about the technology changed between those two moments. Everything about the user did.

That is the Pilot Paradox in its current form, and it explains why so many AI pilots that succeeded went on to produce nothing at all.

The five recommendations, revised

The 2024 advice still stands. Here is how I would restate each piece of it now.

1. Pilot a workflow, not a tool. Aligning with business goals used to mean picking the right strategic priority. In AI it means something more specific: pick a real piece of work that a real team does badly today, and pilot the whole redesigned workflow rather than the software that sits inside it. McKinsey's high performers are distinguished mainly by having redesigned workflows rather than layering AI on top of existing ones. A tool pilot proves the tool works. Only a workflow pilot proves anything worth scaling.

2. Write the scaling plan before the pilot starts. Unchanged from 2024, and still the most commonly skipped step. The second, deeper planning session that everyone intends to hold after the pilot still does not happen. If the scaling plan does not exist on day one, the pilot has no destination, and a pilot with no destination cannot fail to arrive.

3. Measure the floor before you buy anything. This is the one I would add if I could add only one. BCG puts the ratio at roughly 10 percent technology, 20 percent data and algorithms, and 70 percent people, process and organizational change, and its 2026 work found that AI leaders have AI-related skills in 13 percent of their workforce against 1 percent at laggards, and are four times more likely to run structured AI learning programs with protected time to use them. The McKinsey Rewired rule of thumb I quoted in 2024, one dollar on adoption for every dollar on technology, now looks conservative. You cannot budget any of it sensibly until you know where your people actually stand.

4. Pilot with the median employee, not the enthusiast. The original article warned that pilot groups chosen for enthusiasm hide the adoption problem. With AI this is worse, because the gap between an enthusiast and a median employee is now enormous. Put at least a few Level 1 and Level 2 people in the pilot group. They will make the pilot look worse and the business case look better, because they are the ones who will decide whether it scales.

5. Measure adoption, not satisfaction. Continuous evaluation was right, but pilot evaluations still lean on how users felt. Feelings scale poorly. Track how many people used the thing last week, on real work, without being asked. That single number predicts scaling better than any satisfaction score, and it is the number nobody wants to look at.

The key point, revised

I ended the 2024 article by saying that scaling a technology project is not just a technical challenge but a strategic and human one. Right sentence, wrong weighting. Two years of AI pilots have made the case that the human part is not a co-equal third of the problem. It is the problem. The technology has become the easy part.

The organizations getting past the pilot stage are not the ones with better models. They are the ones who raised the capability of the people they were about to hand the models to, before they handed them over.

So if you are about to run an AI pilot, the most useful thing you can do in week one is not to select a vendor. It is to find out, honestly and specifically, what your people can already do. Everything else in the plan depends on that answer, and almost nobody has it.

Move78 measures workforce AI competency against the 7 Levels, so you know where your floor sits before you invest behind it. If you want to see your own distribution, start a conversation.

Share this post

Insights on building a digital and AI-capable workforce.

Submit your email for practical research, frameworks and ideas on workforce capability.

By signing up, you agree to our Terms and Conditions.
Thank you! Your subscription has been received!
Oops! Something went wrong. Please try again.