Book / Read online / Chapter 14
Chapter 14

AI Amplifies What the Company Has Built

The Blank Collar · Kristian Kabashi · about 10 min

Amplification has no preference. It takes what the workflow already carries and does more of it.

Tell the Zillow case as a clean proof of this framework first, then watch it refuse the grade. In November 2021 Zillow Group announced it would wind down Zillow Offers, its home-buying and reselling program. The company said home-price unpredictability, capacity constraints, and other operational challenges had been worsened by the housing market, the pandemic, and labor and supply-chain conditions. It recorded an inventory write-down of roughly 304 million dollars on homes bought above its later estimates of what they would sell for, and expected the wind-down to cut its workforce by about 25 percent.

The tempting reading is that a company automated a judgment it had not made inheritable, and an operation that looked sound came apart once the market moved. The public record does not support it. Nothing published establishes that anyone but its builders could operate the system, that its checks covered the work, that realized resale price was the right evaluation target, or that any single cause produced the outcome.

What the case can do is mark this chapter's limit. A system can be inspectable and still be aimed at a business assumption that does not hold, and internal verification will not catch that. Zillow does not prove that proposition. It makes it unavoidable. So the question stands: what can an audit of an AI deployment establish, and what stays beyond its reach?

What AI amplifies

Start with what amplification means, in plain operating terms. A deployment widens the consequences of the conditions underneath it when it raises any of speed, reach, consistency, or autonomy in a workflow. A direction that could fail unnoticed in one manager's head now fails across every case the system touches. A data definition that was wrong in a monthly report is now wrong at the pace the system runs. Good conditions get widened too. Amplification has no preference. It takes what the workflow already carries and does more of it, faster and in more places.

That is one way to read the AI question, and it is not the only way, because the deployment can also be the weakest reading on the page. A deployment can sit on sound direction, clean data, and a legible process and still fail on its own terms: an evaluation that checks the wrong thing, permissions that are too broad or too narrow, routing that sends the hard cases to the wrong place, a fallback that does nothing when the model is unsure, a transfer that leaves one person as the only operator, a cost or latency profile that makes the whole thing unusable. Each of those is a managerial choice in a specific deployment, and each can be broken while the others are fine. Do not treat AI as a force outside management's control. The pace of model capability in the outside world is not yours to set, and neither are the vendors, contracts, and regulations that bound what you can deploy. The design of the deployment on your workflow, within those limits, is a managerial responsibility, and many of its failure modes are choices someone made and someone can revisit.

Two states matter before any of this can be graded. If no AI deployment touches the workflow you selected, that is not a failure; record No deployment and move on. If a deployment exists but you cannot get authorized evidence to say whether the artifacts and counts it should produce exist at all, record No reading. Neither is a zero, and a confirmed missing artifact is a different thing again: that is a finding.

Correct according to what?

You cannot write an evaluation a stranger could run without a written definition of a correct outcome. That sentence sounds procedural. It is actually the whole chapter, because it separates two classes of work that companies keep confusing.

Some outcomes reconcile against an external record. Did the payment clear, did the shipment arrive, did the address match the postal file. For those, correctness is a lookup, and a deployment can be checked by anyone with access to the record. Judgment outcomes have no such record waiting. Whether a refund was fair, whether a summary was faithful, whether a case was routed to the right owner: none of these can be checked until someone writes down what a correct result would look like, what falls outside the boundary, and what evidence a reviewer would use to tell a good outcome from a plausible one. That does not mean all judgment collapses into a single rule; much of it needs a rubric, a boundary, and a review process, not a fixed rule. But until something is written, there is no evaluation, only opinion collected after the fact.

Hold on to what this measures and what it does not. This suite measures whether your company's work can be inherited by another person or system. It does not measure whether that work is correct. Those are different audits, and the adversarial turn in this chapter is that they can come apart completely. A deployment can verify every run against its written definition and pass, while the written definition itself is wrong. Complete verification certifies conformity to a standard. It says nothing about whether the standard was right.

This is why the usual dashboards mislead. Seats, sessions, active users, and messages count adoption, not completed workflow runs, and a deployment can be busy without finishing anything correctly. Return-on-investment figures produced by the vendor selling the tool, or by the sponsor who staked a reputation on it, may rest on assumptions worth reading closely; that does not make every business case invalid, but it does mean the number arrives with an interest attached. So inspect the assumptions and the workflow evidence directly, and let the completed runs, not the usage, tell you what the deployment actually did.

Audit one deployment

Take the deployment touching the workflow you selected in chapter nine. If none touches it, record No deployment and stop; you have your reading. If a deployment exists but authorized evidence cannot establish whether its artifacts or counts exist, record No reading. A confirmed missing artifact is not No reading. It is a finding.

Define the unit before you inspect it. A deployment is a workflow with an owner and a spend line, not a license or a platform. A run is one completed instance of that workflow. You are auditing runs, not seats.

Six things to find. First, the written boundary stating what the deployment may not do. Second, the written definition of the data it reads. Third, the written procedure, including the explicit exception paths, the place where the customer disputes the renewal and the system may draft or retrieve policy but may not settle the fairness call itself. That boundary between retrieval and authority is the schematic this book has carried; here it becomes a line you either find in writing or find missing. Fourth, name the two accountability functions: who notices wrong output, and who can change the deployment. Record role titles, not personal names. In a larger organization these should be distinct role holders. In a small team where one person holds both, separate the reviewing from the changing in time, against a written criterion. Record the conflict of reviewing one's own work, and bring in a qualified external or board-level check when the consequence warrants it. A known lack of independence is a finding, or a provisional finding, not No reading. Fifth, for the last complete month, record two counts: total runs, and the runs that carry a correctness artifact an authorized independent reviewer could actually inspect. Sixth, name one modification made after handover by an operator other than the person who built it.

That sixth item is the sharp one. A deployment fails the transfer test when the only person who can attest that a run was correct is the person who built it. One real change made by someone else is evidence the work transferred. Absence of such a change is weaker evidence, because it may only mean no occasion to modify the deployment has arisen yet; treat it as a question to pursue, not proof that transfer failed.

The discriminator that keeps this honest: verification coverage does not establish that the expected outcome, the tolerance, or the objective was correct. A month of fully checked runs can be a month of confidently wrong ones. Coverage tells you how much was checked. It says nothing about whether the check was pointed at the right thing.

One stop condition overrides the exercise. If evidence already shows material harm, unsafe behavior, unlawful output, or a harmful objective, contain and escalate through the authorized domain process before grading or optimizing anything. A score must never be used to make harm look handled.

The artifact is an access-controlled, one-page deployment record: the boundary, the data definition, the procedure, the two role titles, the run count, the checkable-evidence count, and the post-handover change. It holds no raw employee or customer records, no credentials, no full prompts, and no control detail that would let someone misuse the system. Call it a probe, not an AI grade.

The person-side reading is short. A Blank Collar writes the setup and the evaluation so that another person can operate the deployment and challenge its results, and then moves toward the judgment, the exception, the relationship, or the consequence the deployment cannot settle. Making the deployment inheritable is the setup. The move that grows you is what you do once it no longer needs you.

The attack I cannot fully answer

Look. This is the objection with the best chance of being right, and I have not fully answered it. This chapter and the ones before it argue that stronger organizational conditions before a deployment lead to better outcomes after it. The reverse may be doing some of the work. A company that automated a workflow years ago may have cleaner data, tighter processes, and clearer direction today partly because the automation forced those improvements. If that is true, then the conditions I am crediting as causes are partly effects, and the arrow I keep drawing points both ways.

I can give you a comparison that could weaken my own claim, which is the least a framework owes you. Inside one company, take two deployments built by the same team on the same model family, on workflows whose organizational conditions were inspected and written down before anyone opened the outcomes. Predeclare what result you expect, the rule for comparing them, and the window over which you will watch. If the workflow with the weaker pre-outcome reading repeatedly outperforms the stronger one, the claim that strong prior conditions are needed for the better result is weakened, for that company.

Be honest about what such a comparison cannot settle. Task difficulty differs between any two workflows. So does the process that generates their data, and so do the many implementation choices behind any two builds. Even with the same team on both deployments, residual management differences remain, because attention, care, and time are never spread evenly across two efforts. Two deployments are a probe, not proof, and I am not claiming to have held any of that constant.

There is no repair section

There is no repair section. I mean that narrowly and I mean it. There is no generic instruction to improve AI that stands apart from a specific workflow and its evidence, because improve the AI is a wish, not an action. The specific defects are another matter: a verification gap, permissions that are too broad or too narrow, a missing fallback, an incomplete transfer that leaves one person as the only operator, a cost or oversight problem. Each of those can be repaired in relation to the workflow it sits on and the evidence you gathered, and the next chapter can route action to them.

One deployment cannot represent a company, and the audit is easier where governance is mature, which favors the large and the old. A young or small firm with strong instincts and thin documentation may inspect worse than a slower company with a compliance department, and that is a limit of the instrument, not a verdict on the firm.

What you carry out of Movement II is a page of evidence, not a score. Not a grade of the AI, not a verdict on the company. You inspected direction, the number, the handoff, the change, and now the deployment, on one real workflow. The next chapter takes that page and asks the disciplined question the whole framework has been building toward: of the constraints you actually confirmed, which one should be addressed first, and what evidence would tell you the intervention worked, before anyone knows how it turns out.

Chapter 14. AI Amplifies What the Company Has Built · The Blank Collar