Skip to content
Fritz Gerald ZephirinMenu

Measure Before You Automate

4 min read

Here is a question that ends most AI success stories politely: compared to what?

The pilot went great. The team loves it. The vendor’s dashboard shows four thousand tasks completed. Compared to what? How long did the workflow take before? What was the error rate before? What did a resolution cost before? In most companies, on most deployments, nobody knows — because nobody measured before the tool arrived, and you cannot retrieve a baseline after the fact. The before is gone. Whatever story gets told now is unfalsifiable, which is precisely why everyone is comfortable telling it.

This is, I’ve come to believe, the least glamorous explanation for the era’s strangest statistic. MIT found 95 percent of enterprise AI pilots produced no measurable P&L impact — and the word doing the heavy lifting is measurable. Some of those pilots surely worked. Some surely destroyed value. Nobody can tell which, including the companies that ran them, because “success” was defined after deployment, by the people who had championed the purchase, using whatever numbers were available — which is to say, the numbers the optimizer was already bending. An unmeasured pilot doesn’t fail. It does something worse: it produces a belief unconnected to reality, which then compounds through every subsequent decision.

The fix costs two weeks. Before any tool above a trivial threshold touches a workflow, we baseline it — and the discipline is so dull it almost reads as advice from a different century, which is why it works:

Four numbers, defined in writing. Volume (how much of the work happens), cycle time (how long it takes end to end), quality (a human-scored sample — reading the actual tickets, not the satisfaction average), and cost per unit (the cost atoms — per ticket, per payroll run, per review). Two weeks of honest data. If the workflow is seasonal, note it; a baseline doesn’t need to be perfect, it needs to exist.

Success criteria pre-registered, in the purchase doc. What the tool must move, by how much, by when — written before the contract is signed, by the process owner, who is not the tool’s enthusiast. Science learned this the hard way: results announced after looking at the data are stories; predictions written down first are tests. Procurement is no different. The sentence “we’ll know it’s working when X” — dated, signed — is worth more than any demo.

The skeptic owns the readout. Whoever doubted the purchase presents the post-deployment numbers. Enthusiasts measure generously without meaning to; skeptics make the success real when it survives them. (There is no readout more convincing than one delivered by the person who voted against the purchase.)

A control where one is free. Two comparable queues, two teams, two regions: give one the tool first. The gold-standard studies that made AI’s real gains citable — the 14 percent in support, the 55 percent in the Copilot experiment — are controlled comparisons, which is exactly why those numbers survived years of scrutiny while ten thousand vendor case studies evaporated. You don’t need a journal. You need the next-best queue left alone for six weeks.

The objection is always speed: markets move; two weeks is forever; just ship it. But run the arithmetic. The baseline costs two weeks once. A false positive — a tool that feels great and does nothing — costs its license fee plus the workflow attention plus, worst of all, the next three decisions made on top of the illusion, for as long as the illusion lasts. The MIT companies didn’t lose two weeks. They lost thirty to forty billion dollars of spend they can’t even audit. Slow is smooth here, and smooth is fast.

There’s a reason this is the last lesson in the first canon, and the humblest. Everything else this site argues — judgment over generation, workflows over models, honest metrics over flattering ones — collapses into hand-waving without the discipline underneath: know your before. Measurement is not the exciting part of the AI transition. It is the part that decides whether the exciting parts actually happened.


Sources & further reading