The cheaper it gets, the more it costs
By Dino Nokic · OCT 2026 · 4 MIN READMcKinsey published a piece this week written for government finance chiefs, but the number at the center of it belongs to anyone running an operation. In a McKinsey survey of 120 enterprises, 93% had exceeded their planned AI budget. Not a few of them, slightly. Nearly half were 11 to 30% over. Only 5% came in under. And this happened while the price of running a fixed level of AI capability was falling roughly tenfold a year.
Read those two facts together and you have the whole story. The unit got cheaper. The bill got bigger. McKinsey calls it the inference-cost paradox. Any fleet owner has a shorter name for it. Diesel getting cheaper per gallon does not shrink the fuel bill if the trucks are running twice the miles and three of them are idling in the yard.
Why the bill grows
The mechanism is simple. A chatbot answers a question with one call. An agent finishes a task with a chain of them: read, check, retry, call a tool, read again. McKinsey puts the multiplier at five to 30 times the tokens of a simple query, "far more in complex cases." Then access widens from a handful of power users to the whole floor, and the simple prompts turn into multistep agents. Volume growth outruns price decline every time.
The line that should worry an operator most is the one about loops. An agent "stuck in a loop can run up charges within hours." Nobody signed a purchase order for that. There is no seat count to check. It shows up on a consumption invoice thirty days later, under a line the finance team folded into "IT." McKinsey's own survey says AI is already 30% of IT costs. The thing nobody is watching is now a third of the budget.
The zombie agent
McKinsey has a name for the worst version: the zombie agent. An always-on workflow that consumes more in compute than it returns in value, and keeps running "only because no one is measuring it." I've watched the non-AI version of this for twenty years. The report nobody reads that still gets built every Monday. The alert that fires 400 times a day until everyone mutes it. The truck in the yard with the engine running because nobody owns the yard. The difference now is the zombie costs real money per step, and it never gets tired.
Here's the contrarian part. The fix is not "spend less on AI." McKinsey is explicit that cutting usage too hard degrades quality in ways cloud savings never did. The fix is knowing what each outcome costs. Their words: track cost per resolved ticket, per summarized document, per accepted recommendation, "rather than aggregate token volume because the aggregate number tells you how much you spent, not whether it was worth it."
Cost per outcome is just cost per mile
Every operator already runs this discipline somewhere. You know cost per mile. Cost per load. Cost per delivery. Nobody in transportation and logistics manages fuel by staring at the total fuel bill. They look at what it costs to move one load one mile, and they notice the week that number moves. AI needs the same unit. Cost per inbound email answered. Cost per exception resolved. Cost per load covered. If you can't name the unit, you can't tell a working agent from a zombie, and the invoice will tell you first.
That's a baseline problem, and baselines are where AI fails. The 93% didn't blow their budgets because the models got worse. They blew them because the thing they were measuring was spend, and spend is not value. The FinOps Foundation survey McKinsey leans on says 98% of practitioners now manage AI spend, up from 31% two years ago, and 40% still can't quantify the return. Managing the bill and knowing what it bought are two different jobs. Most companies hired for the first one.
Three things to do before the invoice
McKinsey's list is written for agencies, but it translates straight to a dispatch office. Put a name on AI spend. One owner, not "IT." Tag every request to the workflow it belongs to, so when the bill moves you know which agent moved it. And cap at the individual agent, not the team, because the one that loops is the one that eats the month, and a team-level ceiling hides it until the ceiling is gone.
In our own transportation and logistics operation, the inbound system that took reaction time from three-plus days to about a minute is judged on one number: reactions, not calls. If the calls climb and the reactions don't, something is looping, and we'd see it in a week, not a quarter. That's not sophisticated. It's the same thing a dispatcher does when a truck's fuel card spikes and its miles don't. The unit is the control.
A cheaper token is not a cheaper operation. Price per step is the vendor's number. Cost per outcome is yours. If you don't know it yet, there's a zombie on the payroll already.
Sources: McKinsey, "The next phase of government AI economics" (article, September 29, 2026) · McKinsey Quarterly, "The cost of intelligence: How CIOs can manage AI demand at scale" (survey of 120 enterprises, July 20, 2026)