"On two occasions I have been asked, 'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?'"
Charles Babbage
A field can change meaning after a system migration while experienced staff go on silently correcting it in reporting. If that correction is never written down, the number depends on memory.
Now hand that field to a machine.
The correction layer
The machine reads the field and finds a number. What it will not find is the correction, because the correction was never written down. I treated a number's presence as proof the company shared one meaning. It did not. Experienced people were quietly translating it in use, correcting it so routinely they had stopped feeling it as a correction. I had tested whether the number was available, never whether its definition survived its translators.
Define the condition through that mechanism. Reliable data is recorded reality that can be traced to its source, interpreted without a resident translator, and reproduced: the same question, asked twice by two people, returning the same answer without an invisible correction making it so. Now say plainly what reliability does not buy. Reproducibility is necessary and insufficient for truth, because a wrong answer can be reproduced perfectly. A field miscounted the same way for years is beautifully consistent. Only the outside world catches that, which is why the exercise later in this chapter ends outside your systems.
None of this cost arrived with AI. Bad data always billed you: wrong forecasts, stock that was not where the system said, the customer invoiced twice. What changes is the geometry of consumption. Machine systems can read more of your records, faster, in more places, while the people who used to apply the corrections sit farther from the reading. Draw the line at what machines can and cannot reach. Where law and policy allow, a system can inspect logs, documents, recordings, demonstrations of the task, even physical evidence through a camera. What it cannot consult is judgment that was exercised and never recorded, and much of the correction layer is exactly that.
Pointed at your customers and your operations, AI does not clean your data. It distributes it.
The same technology, pointed at the data itself under supervision, can help repair it, and that possibility returns at the end of this chapter. Direction of aim decides which one you get.
What toxic means
Toxic is chosen against the industry's favorite word. Every company calls its data an asset, and an asset produces value sitting still. Records about to be read, joined, and acted on at machine speed do not sit still, and their defects travel.
Six signs tell you where to inspect. The number needs a human to interpret it: if a figure requires someone to explain what it excludes, the meaning lives in the person, and the person is not in the deployment. The definition drifted silently: the metric changed meaning during a migration, an acquisition, or a reorganization, nobody restated it, and one series is now secretly two. The same entity exists more than once: one customer, several records, no rule for which wins. The record sits where machines cannot reach: a PDF, a shared inbox, a laptop spreadsheet, a filing cabinet. Nobody has checked it against physical reality lately: the system says forty in stock, the shelf says twelve. And the source feeds someone's incentive: numbers that pay people get shaped by people, the sign least discussed and the one that survives every technical fix.
Resist scoring yourself against that list, because it is not a second grading system, and a count of signs is not a threshold. Every company of any age fires some of them somewhere, and a diagnostic everyone fails diagnoses no one. The signs do one job: they tell you where to point the exercise that follows, on the one number that matters, instead of auditing the whole estate.
Hold on to the quietest version of the problem, because it defeats technical audits entirely. A field can be complete, well-formatted, and technically clean while three functions mean three different things by it. No profiling tool flags it. Every checksum passes. The toxicity is in the meaning, and meaning does not show up in a schema. That is why the exercise below asks people for their definitions independently and in writing, before it asks any system: the disagreement only becomes visible once the definitions sit side by side.
One public case belongs here, held at its real size. JPMorganChase's 2024 annual report says the firm launched LLM Suite to more than 200,000 colleagues in 2024, in an environment it describes as controlled and designed to protect company and customer data, and it credits earlier technology investment for its current products and services. The scale is a fact; the credit is the company's own explanation. The inference you may want to draw, that years of data discipline are what made deployment at that scale possible, is exactly the kind of thing this chapter cares about, and the public record cannot carry it: nothing published shows whether definitions, lineage, correction capture, and external checks were reliable in any specific workflow. Treat it as a strong illustration and a weak proof, and let your own workflow supply the evidence the annual report cannot.
Put the three numbers beside each other
Watch the logic on the carried schematic first, labeled as always: schematic, not evidence. In the disputed-renewal workflow, three parties touch the same facts. Support records that a cancellation was attempted. Billing records whether a refund was approved. The system of record holds whatever the integration wrote down. Ask the only question that matters here: does cancellation attempted mean the same thing in all three places? Attempted by phone, logged after the deadline? Attempted in the app, failed at a confirmation screen? Approved as goodwill, or approved as error correction? Each function answers with total confidence from its own definition, and no invented numbers are needed to see what comes next: a decision about the refund is being made on a phrase that may be three phrases wearing one name.
Now run it for real, on the number that matters to you. The specimen is not free to choose: take the decision-critical number that the workflow you selected in chapter nine consumes or produces, the figure a real decision about that work rests on.
Name three people by role, not by convenience: the producer of the number, someone in the function the number measures, and someone in the function that consumes it to decide. Ask each, independently and in writing, for five fields: the population included, the population excluded, the time window, the source system, and any adjustment applied after extraction. That fifth field is where the correction layer becomes visible, because the adjustments are the corrections, surfacing as words for the first time.
Then have each compute the number for the same closed period. Before anyone opens the results, write down the tolerance: how far apart could the figures be before a real decision would have gone differently? Where a past decision record exists, use it, and let the variance that would have changed that decision set the tolerance. Set it before looking, because a tolerance set afterward will turn out to be exactly as wide as the disagreement.
Put the three definitions and the three results side by side, on one page, before any meeting reconciles them. That ordering is the exercise. Definitions argued in prose get harmonized by the most senior person in the room. Three numbers on a page resist harmonizing, and the gaps between them are the finding.
Last, where it is safe and lawful, check one claim against the world outside your systems: the count on the shelf, the customer your records call active, the folder that should hold the signed contract. Reproducibility inside the systems cannot catch a consistent error. The world can.
If you are small enough that one person holds two or three of these roles, do not grade yourself down for headcount. Separate the functions in time: write the definitions on different days, against a written criterion, before computing anything, and record the conflict of reviewing your own work. Where the consequence warrants real independence, borrow it: an accountant, a board member, a customer-side check. And where independence cannot be had at all, return a provisional finding or No reading, which is information, and better information than a zero earned by being small.
The artifact is one page: the number, its lineage from source to use, the three definitions, the three results, the tolerance, and the external check. Call it a probe. It shows whether one consequential number holds one meaning. It is not a formal grade of your data, and it does not need to be to change how the next meeting reads that figure.
For you, personally, it comes down to this. Your work has one honest mirror, the number it is judged by, and the temptation is never to look straight into it. Face it anyway. Know the definition that governs it, because a metric you cannot define is a sentence someone else wrote about you. If you carry corrections in your head, surface them safely, through the channel that exists for it, as findings about the data instead of confessions about yourself, and only where a good-faith disclosure is protected from being turned back on you. Where the number feeds someone's pay or bonus, that protection is not optional; a written admission that you have been quietly adjusting a compensation-linked figure belongs in a protected channel with amnesty for the disclosure, never on an open working page. The move that grows you is from keeper of the number to co-owner of a better definition, which puts your name on a contribution instead of on a dependency, and beats letting someone else write the sentence your career gets graded against.
Fix one meaning first
If the probe embarrassed you, the reflex will be a program: a data platform, an eighteen-month roadmap, a steering committee. Hold off, at least as a first move. A platform program asks the organization to trust a promise while the divergent meanings keep being consumed by every report and system already running. The narrow move pays sooner: take the number you just traced and repair it end to end. Write its definition where a machine and a new hire can both read it. Reconcile the duplicates for that entity only. Close the gap between the three computations, and record which function's adjustments were corrections and which were habits. Then take the next number. Meaning gets repaired one number at a time.
The machines have a supervised role in this, stated without the vendor gloss. Classification, reconciliation across systems, extraction from documents, and anomaly finding are tasks current systems assist with under human review, and they can shorten the mechanical part of a lineage repair. Whether that makes your cleanup faster, cheaper, or safer depends on your records, your constraints, and your review capacity, so treat each application as its own small probe with its own baseline.
The hardest part of the repair has people in it. The corrections you want to capture belong, in practice, to the people who apply them, and asking for them is an organizational change, not a documentation task. Do it as a negotiated, policy-compliant transfer, and put the offer on the table before the ask: credit attached to the improved definition, a role operating or governing the system built on it, learning that moves the person forward, responsibility for the next version. What the organization must not do is treat the handover as extraction, because the lesson gets learned instantly and generally, and the next correction stays in someone's head on purpose. And this book will not flip that into advice for the keeper. Hoarding the correction is not a strategy; the durable position is co-authoring the better definition and owning what it cannot yet settle.
And notice what the repair assumed: that once the number holds one meaning, the workflow will carry it faithfully from hand to hand. That is a large assumption. The number moves through handoffs, handoffs run on process, and whether your process is written down through explicit handoffs and exceptions or still depends on someone knowing what happens next is the next condition's question.