Book / Read online / Chapter 2
Chapter 2

The Pilot Was Designed to Fail

The Blank Collar · Kristian Kabashi · about 9 min

Your pilot did not fail. It succeeded, completely, at being a pilot.

Month eleven of the AI pilot, and the steering committee is gathered for the review. The deck has forty slides. The word on most of them is "learnings." There were workshops, a vendor evaluation matrix, an internal survey with encouraging sentiment scores, and a demo that impressed the board in March. What there is not, anywhere in the forty slides, is a number that moved.

If you have sat in that meeting, this chapter is for you, and it begins with the sentence you have earned the right to say: we tried AI, and it didn't work. I believe you. You did try. The money was real, the effort was real, the disappointment is real. What I want to show you is that the trial was rigged from the first week, by nobody, which is precisely why it keeps happening to companies full of intelligent people.

And the bill was bigger than the budget line. A failed pilot spends things that never appear in the post-mortem: the board's patience, the innovation team's credibility, the organization's willingness to believe the next attempt. Every company that ran pilot theater has armored its own skeptics, which means the real transformation, when someone finally attempts it, starts in a hole the pilot dug. The cost of trying AI wrong is not zero. It is negative, and it compounds.

Your pilot did not fail. It succeeded, completely, at being a pilot.

Eliminating the suspects

When something this expensive dies, there are three suspects everyone interrogates first. Take them one at a time, the way a detective would, because watching them each produce an alibi is how you find the actual culprit.

Suspect one: budget. Sometimes it really is the money. The test: was a resource the pilot needed actually withheld, a data feed refused, an engineer never assigned? If yes, you have an ordinary funding failure. But the pattern this chapter dissects appears at every budget size, in companies that spent millions and companies that spent thousands.

Suspect two: talent. Sometimes it is the people. The test: did the team lack a capability the task demonstrably required? If yes, name the capability and point to where the work failed without it, and you have a hiring plan instead of a mystery. If no one can name it, "talent" is not an explanation. It is a place to put the file, and the pilot's design goes back on the table.

Suspect three: the technology. Sometimes the model cannot do the job, and a controlled trial on the actual work will show it missing. When that happens, write it down and stop; a real capability gap is the cheapest possible finding. What this alibi cannot survive is the tool working across the street, at your own competitor, on the same class of task.

Run the three tests. If none convicts, you are left with the possibility nobody wants to book: the pilot itself, as designed. Not sabotaged. Designed. And every design choice that killed it was made carefully, by competent people, for reasons that would survive any governance review. That is what makes this a trap rather than a blunder, and it is worth taking the design apart bolt by bolt, because you will recognize every piece.

The anatomy of a rigged trial

First choice: the pilot was sandboxed, for safety. Reasonable. Legal wanted risk contained, IT wanted systems protected, and so the AI was given a fenced area away from live customers, live data, and live money. But notice what the fence guarantees. P&L impact requires touching the P&L. A pilot sealed off from real workflows cannot produce real results, by construction, the way a caged bird cannot demonstrate migration. Anything kept safely away from the real business will die safely away from the real business.

One boundary before this indictment continues, owed to every reader in a bank, an insurer, a hospital, or a utility: some fences are not choices. They are law, and nothing in this book asks anyone to probe live patient data or move client money on a vibe. The distinction that matters is between the fence a regulator requires and the fence a committee adds on top, then attributes to the regulator. Regulated industries run real experiments constantly: shadow mode, parallel runs, historical data, consent and audit trails; the supervised probe is a discipline those industries practically invented. Theater is not the presence of a fence. It is the voluntary fence wearing a mandatory fence's uniform.

Second choice: the pilot was unowned, for consensus. Also reasonable. AI touches everything, said the deck, so every function got a seat: a steering committee, a working group, dotted lines to Legal, IT, HR, and Innovation. The result is a project that belongs to everyone, which is the corporate spelling of no one. When the pilot stalls, no single person's number suffers. When it must fight for data access or an exception to process, no single person has the authority to win that fight. A pilot nobody owns is a press release with a budget.

Third choice: the pilot was measured on activity, for accountability. The most reasonable of all, and the most lethal. Because the sandbox made money impossible and the committee made responsibility diffuse, success had to be defined as something achievable: workshops held, employees trained, use cases identified, learnings captured. The pilot was graded like a conference, and like most conferences, it passed. Activity metrics do not merely fail to measure value. They manufacture the appearance of progress precisely where none is occurring, which is why the program that produced nothing can be presented, without a single lie on the slide, as a success.

Sandboxed, unowned, measured on activity. Safety, consensus, accountability. Three virtues, assembled into a machine for producing nothing. Nobody chose the outcome. Everybody chose a component.

Assembled, the machine keeps a schedule. Month one, vendor evaluation begins. Month three, the kickoff workshop, with the good catering. Month five, a demo in the sandbox that genuinely impresses everyone, because demos are what the sandbox is good at. Month seven, the first integration request dies in a queue. Month nine, the working group's cadence slips from weekly to monthly. Month eleven, forty slides, the word "learnings." Month twelve, the committee recommends a second phase, scoped more realistically. If your calendar matched that one, understand: you did not run an experiment. You ran a ritual, and rituals always complete successfully.

Shopping, dressed as change

There is a fourth pattern, common enough to deserve its own indictment: the company that skipped the pilot theater and went straight to procurement. Licenses for everyone. A copilot in every application, a chatbot on the website, an enterprise agreement with a number large enough to feel like commitment.

Licenses are shopping. Shopping is not transformation. Buying access to intelligence changes a company about as much as a library card changes what anyone reads. Renewal reviews often show the pattern: heavy first-month use, then decay and idle seats. Who decides, who checks, and what gets done may remain unchanged. The company added intelligence to the existing machine, which produced its usual output with a new subscription line.

The licenses also lose to the shadow. Recall the employee from chapter one with the personal subscription. Nobody has timed her work against the enterprise deployment; what can be compared is the design. She configured her setup around her actual work, with instant permission to change her own process. The sanctioned tool arrived with default settings, a training video, and no authority to alter a single workflow. The company purchased intelligence and withheld the one thing intelligence needs, the right to change how work happens, then read the flat usage report as proof the technology was overhyped.

The market is beginning to say this out loud. Gartner forecasts that over 40 percent of agentic AI projects will be canceled by the end of 2027, and uses a phrase for part of what is inflating the pipeline: agent washing, vendors relabeling existing products as agentic without meaningful agent capability. Gartner's claim is narrower than the joke it invites, and I will keep it narrow: some of what companies are buying as autonomous transformation is rebranded software, and a large share of the genuine projects will not survive their own economics. The shopping is failing on both ends of the receipt.

The pilot is a self-portrait

Now step back from the wreckage and ask the only question that produces anything: why does the pilot always take this exact shape? Different industries, different vendors, different years, and the same sandbox, the same committee, the same activity metrics, as if every company copied the same doomed template. Nobody copied anything. Something generated it.

The generator is the machinery from chapter one. The old operating system builds pilots the way it builds everything: risk goes to a sandbox, authority goes to a committee, performance goes to a dashboard of activity. Ask an org chart to design an experiment and it will design a smaller org chart. Your pilot is a self-portrait of your operating system, painted at reduced scale, which is why its failure is the most informative document your company produced last year. The pilot did not fail. It told the truth: this is what the machinery does, at survivable cost and in miniature, to anything that needs speed, ownership, and permission to change the process. The same truth waits for every future initiative fed into the same machine, at full scale and full cost.

Look at the incentives. A sponsor can show the board motion without risking failure. A committee can claim relevance without owning the result. A vendor can protect a renewal; skeptics get confirmation. Each incentive is rational alone. Together, they can leave the company paying for movement without change. The pilot is not just a self-portrait of your operating system. It is the equilibrium the system rewards.

What a real probe looks like

The alternative is not a bolder pilot. It is a different species of thing, and you will get the full method in the third movement of this book. But the contrast fits in four sentences, and it belongs here so the trap has a visible exit.

A probe takes one production-relevant workflow, end to end. It is owned by one person whose name is attached to the outcome. It declares that outcome before it starts, in a unit that matters: minutes, money, error rate, risk caught, customers kept. It runs on data and controls appropriate to the risk, which in a regulated shop can mean shadow mode, historical cases, or a parallel run beside the human process, because a test can be real without being reckless. And it carries explicit permission to change the process it touches, since the process is the thing being tested. Everything the pilot's designers would veto, in other words, and that is the point. The pilot was designed to protect the machinery from the technology. The probe is designed to find out what the machinery is costing you.

Figure 1. A pilot is designed to be survivable. A probe is designed to be real: one workflow, one accountable owner, a predeclared outcome, and a legitimate route into operations.
Figure 1. A pilot is designed to be survivable. A probe is designed to be real: one workflow, one accountable owner, a predeclared outcome, and a legitimate route into operations.

Until you can tell those two intentions apart, the experiments you fund will keep serving the first one.

So take the deck from month eleven and read it again, this time as the diagnosis it always was. The learnings were real. They were just never about the technology.

One last reading before the next chapter, and it is the one this book is named for. Everything above is about your company's pilots. Point it at yourself. If you have built a working method with these tools, you are holding the same evidence at a smaller scale: it worked because you controlled the whole loop. No fence, no committee, no dashboard of activity. Now do the part the pilot never does. Surface it, through an approved channel your company will actually hear, and put your name on it, risks included, because a working method that stays hidden helps you and fixes nothing. The Blank Collar, in this chapter's terms, is the person who can tell which parts of their own week are a pilot and which are the machinery, and who takes accountability for the first instead of waiting for a program to legitimize it. Your company is running an exam. So is your calendar.

Chapter 2. The Pilot Was Designed to Fail · The Blank Collar