Substitution changes who performs a step. Redesign changes what the work requires.
Ingka Group, the largest IKEA retailer, reported that 8,500 of its call-center co-workers had been reskilled after a chatbot named Billie took over a share of the simpler customer inquiries. Sit with that number before you decide what it means. By itself it settles nothing. Reskilled toward what? A company can reskill people into a holding pattern, or into a queue for the next reduction, or into the work that finally uses what a person can do and a machine cannot. The figure is the beginning of a question, not the answer to one, and that question is the whole subject of this chapter.
Look, here is the mistake capable executives make, and I feel its pull whenever a demo works. Visible time saved invites the mind to supply a redesign that never happened. The technology moves. The workflow does not. A machine drops into one step, the person who did it is moved aside, the saved hours are counted, and the org chart around it stays exactly where it was: same handoffs, same approval, same exception path, same accountability in the same place. Something got faster. The design did not change. Because the slide shows a real saving, the company believes it bought a transformation when it bought a quicker version of what it already had.
So draw the line the title draws, then discipline it, because the title is sharper than the truth.
Substitution changes who performs a step. Redesign changes what the work requires.

Substitution is not a failure. On a clean, stable, well-defined step, swapping a machine in for a manual task can be a sound move, and pretending otherwise is its own kind of dishonesty. The failure begins one level up, when a leader treats that local swap as organization design and stops there. Substitution runs out of road because it optimizes a box while leaving every line into and out of that box untouched, and most of what makes work slow, fragile, or wrong lives in the lines, not the boxes.
Redesign starts somewhere else entirely. It starts from the outcome the work exists to produce, asks what that outcome actually requires, then rearranges everything around the answer. That is why it reaches things substitution never touches. Begin from the outcome and the handoffs change, because a handoff that existed to move paper between two people may not be needed once the routine step is machine-run, and a new handoff may be needed to route the hard cases to a human fast. The evidence changes, because a redesigned step has to prove it did its job in a form someone can check, where the old manual step was often trusted to do it without proof. Authority changes, because when a machine drafts the response, someone has to hold the decision the machine is not allowed to make, and that authority has to be named rather than assumed. Exceptions change, because the cases an experienced person used to absorb without anyone noticing now have to be caught and routed on purpose. Verification changes, because a step running at machine speed produces more output than the old spot-check was ever built for. Fallback changes, because a redesigned workflow has to say what happens when the machine is unsure or wrong, and the old workflow never needed a fallback for a human who could simply pause. And accountability changes, because all of it has to land on a role that can answer for the result. Move one of those without the others and the redesign leaks. Substitution optimizes a box. Redesign redraws the boxes and everything between them.
Return to Ingka, and take its reported facts as exactly what they are, a company describing its own operation with no outside counterfactual. Ingka said Billie was rolled out from fiscal 2021, and that from 2021 through 2023 it resolved roughly 47 percent of the inquiries it received, some 3.2 million interactions, which Ingka associated with nearly 13 million euros in savings. Those are Ingka's figures, and they describe the substitution: a bounded, repetitive class of inquiry moved to a machine.
What Ingka reported doing alongside that is the part worth studying. It said it reskilled those 8,500 co-workers toward remote interior design, digital retail sales, relationship building, and inquiries that require complex problem-solving. Read that as the company's account of moving work up toward judgment, taste, and relationships, and hold every limit on it at once. The record does not establish that every worker landed in the same role, that no job was lost, that the chatbot by itself caused the savings or any sales, or that customers were better served because of the change. It is one company's report of moving work up the value ladder rather than simply removing it, and a report is not a proof. What makes it useful is the shape it describes, because that shape is what a redesign looks like from the outside: the simple work goes down to the machine, the human work moves up, and the two moves are planned together.
For the mechanism underneath the shape, a controlled kind of evidence exists, and it complements the operating story instead of echoing it. A pre-registered field experiment, released as an NBER working paper in 2025, put 776 Procter & Gamble professionals to work on real product-innovation challenges, randomly assigned to work alone or in pairs, with or without generative AI. Task-level and controlled, it isolates what a company report never can. One disclosure belongs with it: the paper included Procter & Gamble-affiliated coauthors, the company made gifts to the relevant academic institute during 2023 to 2025, and one coauthor, Karim Lakhani, had earlier received consulting compensation from the company. Read the results with that in view.
Two findings matter. In the studied tasks, individuals working with AI matched the performance of two-person teams working without it. And the expertise boundary softened: without AI, research participants leaned toward technical proposals and commercial participants toward commercial ones, while participants using AI produced more balanced proposals regardless of their discipline, and reported more positive emotions while doing the work.
The bounded inference is where redesign gets its mechanism, and where overreach has to be stopped. Those findings describe a change in what a single contributor could reach across during that work: the expertise boundary that usually sorts people by background softened. That is the kind of change that lets a workflow be redrawn, because work that leaned on combining separate specialists can sometimes be recombined differently. It is not evidence that organizations no longer need teams, that P&G eliminated anything, or that these participants became permanent generalists. They worked across a wider boundary for the length of an experiment. Extend the finding one inch past the task and you are inventing.
If the chapter stopped there it would read as an advertisement, so here is the boundary that keeps it honest. A 2023 field experiment ran 758 BCG consultants, about 7 percent of the firm's individual contributors at the time, through tasks built to sit inside and outside the technology's competence. Across eighteen tasks chosen to be within that boundary, AI use raised speed by more than 25 percent, measured quality by more than 30 percent, and completion by more than 12 percent. On a single task deliberately designed to sit outside the boundary, consultants in the combined AI conditions were correct 19 percentage points less often than the control group. This was a company-collaborative experiment, with company-designed tasks, company-collected data, and BCG-affiliated coauthors, tied to a 2023 capability boundary that has since moved, so treat it as a field experiment run with the firm, not independent external validation, and do not stretch the negative result into a claim about all consulting, all models, or every problem-solving task.
Stretched only as far as it goes, the result changes how a redesign allocates work. The lesson is not that AI helps or that AI hurts. The same tool did both, depending on which side of a tested boundary the task fell. So redesign cannot allocate work by enthusiasm, or by how good the output looks. It allocates by tested capability, task by task, and then keeps the verification and fallback that catch the cases where the allocation was wrong, because a fluent answer to a task the machine should never have been given is still a wrong answer, and it arrives looking exactly like a right one. Persuasiveness is not correctness. A step that produces polished output nobody checks is not a redesigned step. It is an unmonitored one.
Read the three cases once through the Blank Collar framework, without pretending to grade any of them. Vision names the outcome the redesign serves, so the work is aimed before it is rearranged. Data supplies the evidence that says which step is stable enough to move and which task sits inside the boundary. Process is where the redesign physically happens, in the handoffs that get redrawn when the routine work relocates. Human Experience decides whether the people whose work moved can challenge the new design and grow into what it opens, or whether they were simply routed around. And AI performs only the work whose boundary has been tested, with everything else still owned by a person. That is not a score. It is a way of seeing that redesign moves all five together while substitution moves one and hopes.
So the single decision this chapter asks you to make, on the workflow you selected in Chapter 9, is a set of questions rather than a new framework. What outcome does this workflow exist to produce? Which step in it is repeatable and stable enough to transfer to a machine? Where does verification sit once that step moves, and who owns it? What new responsibility becomes possible for the person whose step was transferred? And what evidence will you review, after an interval you name in advance, to judge whether the redesign did what you claimed? Write those answers down before you start, name the review date when you name the change, and you have a redesign you can be held to. Skip the fourth and fifth questions and you have a substitution with a larger vocabulary.
Put those answers on one page while you make the change, not after, because a redesign remembered later is a redesign flattered later. The record is short. It names the outcome you selected, the repeatable step you transferred, and the evidence that will be available at the handoff so the next role can see what the machine did. It names the role that holds exception authority, the case the machine may not decide, and the roles that own verification and fallback when a run is wrong or the machine is unsure. It names the credible next responsibility the transferred step opened for the person who used to do it. And it names the interval after which you will look again, together with the evidence that would confirm the redesign worked and the evidence that would weaken the claim. That last pairing is what keeps you honest. A page that records only saved hours cannot tell a changed workflow from a faster old one, because saved time is produced by substitution too. A page that records the handoff, the authority, and the review lets a manager see which one they actually bought, and revise the design when the interval arrives without having promised in advance that it would succeed.
The fourth question is the one companies drop, and dropping it is not a soft HR oversight. It breaks the redesign as an operating matter. Capability, authority, and a credible next responsibility are not follow-up steps that come after the real work; they are load-bearing parts of the design, because a workflow whose people were routed around instead of moved up will not hold. The knowledge you did not transfer stays in the heads that hold it. The exceptions you did not resource keep escaping. And the people who watched the routine leave with nothing offered above it will read the message precisely and hold the next transferable piece a little closer. An organization has not finished a redesign when it removes work and builds nothing above it for the people whose work moved. It has not developed them. It has removed a rung and called the gap development.
One last observation hands the book forward. Everything here happened at the scale of a workflow: one class of inquiry, one innovation task, one bounded set of consulting problems. Redesign lives at that scale, which makes its cost and difficulty questions of continuity and capacity, not of company size. You do not transform the whole company in one swallow. You eat the elephant one workflow at a time. A small firm does not have to wait for an enterprise budget to redesign one workflow well. Whether that is an advantage or an exposure is the next chapter.