Skip to main content
Adomo

The Creep of Cost

Last updated: August 2026

Making AI a predictable line item: deterministic execution, automatic model routing, and budgets enforced before the spend.

Executive Summary

The first era of enterprise AI ended this spring. Token maxing, the corporate fad of consuming as many tokens as possible as a proxy for productivity, collapsed under its own invoices. Uber exhausted its 2026 AI budget in four months. Meta shut down its internal consumption leaderboard within days of launching it, and Amazon told engineers to stop using AI just to use AI. The market has already pivoted to blending premium and economy models, a shift the press now calls tokenomics.

The pivot solved the wrong half of the problem. Spend per task is lower. Predictability is not better. Gartner forecasts $207 billion in AI agent software spending for 2026, up 139 percent year over year, and 62 percent of organizations still cannot predict their monthly AI bill.

Unpredictability is architectural, not a pricing problem. Usage-based billing, agentic workflows that consume 5 to 30 times the tokens of a chat query, always-on background inference, and fragmented ownership all push the same direction, and cheaper tokens cannot correct any of them. Three things can, and they compound. Execute as much work as possible in deterministic software that consumes no tokens at all. Route what genuinely needs a model to the cheapest model that does the job well. Enforce budgets before spend occurs rather than discovering it on the invoice.

1. The Bill Nobody Chose

The layoff headlines blur two different stories, and which one is yours decides everything that follows.

When Meta cuts 8,000 jobs to offset $115 billion or more of AI capex, that is capital reallocation by a company building infrastructure it will own, at a cost it chose and can depreciate on a schedule. That story belongs to perhaps a dozen firms. Everyone else buys intelligence by the token, as an operating expense that reprices constantly and scales with adoption. This paper is about the second group, because their cost problem has no owner. Nobody chose the number. It accumulated across API contracts, AI features bundled into SaaS, cloud inference, agent platforms, and shadow spend on expense reports.

The founding claim of the era, that AI would cost less than the people it replaced, has not held. Challenger, Gray & Christmas counted 101,743 U.S. job cuts attributed to AI through June 2026, nearly double all of 2025. Separately, Orgvue found that among businesses that had already made redundancies on the back of an AI deployment, 55 percent judged those decisions wrong, and Gartner expects half the companies that attributed cuts to AI to rehire by 2027. The replacement often costs more than the person who left, because the role now includes managing and auditing the AI installed in the interim.

The productivity gains are real, and organizations that retreat will lose to those that press on. What failed is the operating model: adopt fast, measure later, let every team burn its own tokens.

2. Cheaper Tokens, Bigger Bills

The economics that killed the fad are simple. Token prices have fallen for the entire life of this market and bills have risen for the entire life of this market: inference cost for GPT-3.5-level performance dropped more than 280-fold between 2022 and 2024, and spend rose anyway, because consumption outran price. In 2026 the same pattern runs at a larger scale. Gartner forecasts $207 billion in AI agent software spending this year, up 139 percent from $86 billion in 2025, and per-developer token consumption rose 18.6 times in nine months at the height of the trend. Agentic workflows consume 5 to 30 times the tokens of a single chat query, and the FinOps Foundation now locates 80 to 90 percent of enterprise AI spend in inference. Retrieval architectures re-ship the enterprise’s own knowledge to a model on every call, and background agents run whether or not anyone asked a question.

AI agent software spending: $86 billion in 2025 to a forecast $207 billion in 2026, up 139 percent

By spring the correction had its case studies. Meta’s internal leaderboard, Claudeonomics, ranked 85,000 employees by token consumption and handed badges to the heaviest users; the top user burned 281 billion tokens in a month, the company roughly 60 trillion, and Meta shut it down within days. KAIDATA found companies already three times over their full-year AI budgets by April. Uber, where agentic-coding adoption jumped from 32 to 84 percent of 5,000 engineers in a single month, exhausted its annual AI budget by April and answered with a hard $1,500 monthly per-tool cap, a real-time usage dashboard, and executive approval required to exceed the limit.

The market corrected with it. Microsoft’s CEO warned publicly that customers of frontier models pay twice, once in tokens and once in proprietary data. Walmart and Starbucks scaled back agent plans. Vendors repriced under the strain: Cursor replaced its $20 flat plan with credits, and Salesforce, projecting a $300 million annual Anthropic bill, changed Agentforce pricing three times in eighteen months while shopping for model routing.

The emerging consensus, blend cheaper models for routine work and reserve frontier capability for hard problems, is directionally right, and the published data supports it. AI.cc’s 2026 AI API Infrastructure Report puts blended enterprise token cost, the average paid per million tokens, down 67 percent in the twelve months to April 2026, from $18.40 to $6.07. Separately, on a 200-million-token monthly workload, it prices frontier-only routing at $18.00 per million tokens against $2.80 under basic three-tier routing. Routing every workload through a single premium model, the report concludes, means overpaying by 60 to 80 percent. AI.cc sells a unified multi-model API platform, so this is a vendor’s account of the category it competes in; the independent check on it is the next paragraph.

Where the 67 percent came from matters more than its size. YipitData, tracking billing across cloud API services, finds effective token prices fell only about 6 percent in the first five months of 2026. Providers barely cut prices. Enterprises cut costs by changing which models do the work, which is to say the saving came from architecture, not from the market. AI.cc’s own analysis agrees: the driver was not that models got cheaper but that enterprises stopped over-provisioning frontier capacity for tasks that do not require it.

Effective token prices fell 6 percent while blended enterprise cost fell 67 percent: the saving came from architecture, not the market

Blending by hand is what fails. Model catalogs change monthly and prices reprice quarterly, so last quarter’s routing table is already wrong. Employees default to the most capable model they can reach, because nothing shows them the price of the choice. Every model added multiplies billing surfaces and error risk. The market got the diagnosis right and stopped halfway through the cure.

3. Unforecastable Is Worse Than Expensive

For a public company, guidance is a promise. Ramp’s corporate card data shows heavy AI spenders seeing month-over-month cost swings above 50 percent in roughly one month in four, and the FinOps Foundation finds 73 percent of organizations reporting AI costs above projection. A line item that behaves that way can quietly break a quarter, and it resists every technique FP&A was built on: it is multi-source, usage-driven rather than seat-driven, and owned by no one. Engineering provisions it, business units consume it, procurement contracts fragments of it, and finance discovers it.

Even the spender cannot connect the number to output. Uber’s COO Andrew Macdonald, mid-blowout, said it was “very hard to draw a line” between the company’s AI-assisted commits and whether it was shipping more useful features. Forrester expects 25 percent of planned 2026 AI spending to slip into 2027 as CFOs push back on projects that cannot demonstrate measurable returns.

KPMG’s Q2 2026 survey of 204 leaders at large U.S. enterprises found that two thirds have AI cost dashboards, 36 percent have direct token or usage controls, and 26 percent have full real-time visibility into what their AI systems cost to run. Dashboards describe an overrun. They do not prevent one.

The visibility gap: 66 percent of large enterprises have AI cost dashboards, 36 percent have usage controls, 26 percent have real-time visibility

Closing the gap correlates with results: KPMG finds organizations with strong cost visibility five times more likely to have achieved AI ROI.

A public company can absorb a cost that is high. It cannot absorb a cost that is unknowable.

Get the full paper

That is the diagnosis. The remaining two sections are the answer.

Section 4, What Predictability Requires, sets out the three architectural decisions in the order that matters, and prices them: a workload’s relative monthly cost falls from 100 under frontier-by-default, to 16 under tiered routing, to 5 under deterministic execution inside an enforced budget, using AI.cc’s published blended rates. It also shows which steps of a ten-step workflow reach a model at all.

Section 5, How Adomo Makes the Number Plannable, covers deterministic policy enforcement, automatic complexity-based routing, and the budget mechanism most AI spending lacks: a policy you set in advance for what happens when the budget runs out. It includes per-process unit economics from Adomo’s credit estimator, from 5 cents to triage a ticket to $4.50 for a monthly operating review deck.

Get Sections 4 and 5, free

The full paper opens right away in your browser.

No spam. Unsubscribe anytime. See our Privacy Policy.

Sources

Figures above are drawn from Gartner (AI agent software spending, rehiring forecast), Challenger, Gray & Christmas (job cut reports), the FinOps Foundation (State of FinOps 2026), KPMG (AI Quarterly Pulse Survey Q2 2026), Stanford HAI (AI Index Report 2025, cited as historical background), AI.cc (Enterprise Guide to Unified AI API Platforms in 2026, a vendor analysis of the category AI.cc competes in), YipitData (cloud LLM pricing analysis, May 2026), KAIDATA Consulting (Enterprise AI Budget Burnout in 2026), Ramp, Orgvue (2025 research, revisited by CNBC in July 2026), and Forrester. The full paper carries complete citations with figures, dates, and sample sizes.

Survey figures vary by sample and methodology; readers should consult primary sources before relying on any specific figure.

Get in touch

Reach our team. We'll respond within 24 hours.

All fields are required.

Call us

Available during business hours

+1 (866) 995-4498

Monday – Friday, 9 AM – 6 PM PST