Uber exhausted its entire 2026 AI coding budget in four months. Microsoft withdrew a tool it had rolled out six months earlier. Neither company ran out of money. Both ran out of operating model.
There is a specific kind of corporate memory loss worth naming. Between roughly 2012 and 2020, thousands of businesses moved their server estates to the cloud on the promise of lower cost. A great many of them lifted the virtual machines across unchanged, kept them running twenty-four hours a day as they always had, and discovered eighteen months later that the bill had gone up.
The technology had changed. Nothing else had. Procurement still bought capacity in advance. Engineering still sized for peak. Finance still saw one line item arriving monthly, after the money had already been spent, with no way to attribute it to a team, a product or a customer.
The discipline that eventually fixed this — FinOps — was not a tool. It was an admission that when consumption becomes elastic and instantaneous, the operating model has to change to match. Budgets had to become allocations. Spend had to become visible to the person causing it, on the day they caused it. Somebody had to own a number that had previously belonged to nobody.
We are now watching the same film with a different technology, and it is running considerably faster.
Two companies that hit the wall in public
In May 2026, Uber’s leadership disclosed that the company had consumed its planned 2026 budget for AI coding tools within the first four months of the year. Individual engineers had been generating token bills of between $500 and $2,000 a month. The company’s response was not to abandon the tools; it was to build the missing operating model around them — a $1,500 monthly cap per engineer per tool, an internal dashboard showing real-time consumption, and a formal review process for anyone whose workload genuinely required more.
Three weeks earlier, on 14 May 2026, Microsoft began revoking internal Claude Code licences for most employees, completing the withdrawal by the end of June. It had made the tool available to thousands of staff only six months previously. Teams were steered back towards GitHub Copilot CLI — a cheaper tool that Microsoft owns. In an internal email, executive vice-president Jay Parikh asked staff to be “mindful of how we consume tokens”, and departments were given token pools with budget targets from July.
It is worth being precise about what these two events are and are not. They are not evidence that AI coding tools fail to deliver value; both companies retained capability and both continue to invest heavily. They are evidence that consumption-based AI, dropped into an organisation that budgets annually and reviews quarterly, will find the edge of the budget before anyone has agreed who owns the number.
If the most instrumented engineering organisations on earth were surprised by their own token bills, the mid-market should assume it will be surprised too.
Why this is not simply a pricing problem
The obvious rebuttal is that token prices are falling, and they are — steeply. The rebuttal does not survive contact with the data. The FinOps Foundation’s 2026 research found that 73% of enterprises exceeded their original AI cost projections. Average enterprise AI budgets have moved from roughly $1.2m a year in 2024 to around $7m in 2026. Gartner puts global AI spending in 2026 at $2.59 trillion.
This is Jevons’ paradox arriving on a quarterly invoice. When the unit cost of a capability falls, consumption of it rises faster than the price drops. Cheaper tokens do not produce a smaller bill; they produce more tokens. Agentic tools sharpen this further, because a single instruction can now trigger thousands of model calls with no human in the loop to notice.
73%
of enterprises exceeded their original AI cost projections, according to the FinOps Foundation’s 2026 research
The forecasting failure is the part that should trouble a finance director most. Mature cloud FinOps teams typically forecast within one to three per cent of actual spend. The same teams, forecasting AI consumption, have been missing by a factor of two to three. That is not a variance to absorb; it is a number that cannot be planned around. In 2025, 31% of FinOps practitioners had responsibility for AI spend. In 2026 that figure is 98%.
The mid-market version is worse, and quieter
A business turning over £15m to £300m is unlikely to have a FinOps function, a showback model, or per-team cloud budgets. It very likely has a finance team that treats technology cost as a small number of predictable subscriptions.
AI tooling arrives looking exactly like those subscriptions. It is bought per seat, it appears on a card statement, and it is approved by someone comparing it against the cost of a contractor. Underneath, it behaves like metered infrastructure: unbounded, usage-driven, and invisible until the invoice.
The failure mode is therefore not a dramatic overspend. It is a slow one. A team adopts a tool. Consumption grows because the tool is genuinely useful. Nobody watches the per-user trend because nobody owns it. Nine months later the aggregate has quietly become material and — this is the part that matters — nobody can say what it bought, because the value was never expressed as a number either.
Lift-and-shift, again

The cloud lesson was never really about servers. It was that elastic consumption requires three things an unchanged organisation does not have: attribution, ownership, and a unit economic.
Attribution means knowing which team, product or customer caused the spend, close to real time. Uber’s first move was a dashboard, not a cap — visibility precedes control.
Ownership means a named person whose budget it is. Microsoft’s departmental token pools are precisely this: converting one company-wide cost into many local ones with local consequences.
A unit economic means expressing consumption against something the business already measures — cost per ticket resolved, per claim processed, per release shipped. Without it there is no way to separate a tool that is expensive from a tool that is expensive and worth it. That distinction is the whole game, and it is the one most AI business cases never establish.
An organisation that installs AI without those three things has performed a lift-and-shift. It has changed the technology and left the operating model alone, and it will get the result the cloud migrators got, on a shorter timescale.
What to settle before the budget, not after
None of this argues for slower adoption. It argues for adopting the cost discipline at the same time as the capability, rather than eighteen months later under pressure. The two companies above are now doing in public, expensively, what could have been designed in quietly at the start.
Five questions to answer before the next AI tool is approved
- Who owns this line of spend by name, and whose budget absorbs it when it doubles?
- What will we see, and how fast, if consumption triples in a month — and who receives that alert?
- What is the unit — per ticket, per claim, per release — against which we will judge whether this is worth it?
- What is the cap, and what is the documented route for someone whose work genuinely needs to exceed it?
- If we withdrew this tool in six months, what would we have to put back, and what would that cost?
A business that can answer those five has not solved AI economics. It has done something more useful: it has made the spend legible, which means it can be defended when it is worth defending and stopped when it is not.
The organisations that look intelligent about AI cost in 2027 will not be the ones that spent least. They will be the ones that could always say what they were buying.