Gartner Warns Agentic AI Inference Costs Will More Than Quintuple by 2028 Despite Cheaper Tokens
Digital Transformation

Gartner Warns Agentic AI Inference Costs Will More Than Quintuple by 2028 Despite Cheaper Tokens

Gartner's latest forecast complicates the AI business case CIOs have been building all year: token prices keep falling, but total inference spend on agentic workflows is set to rise sharply anyway.

PublishedAugust 18, 2026
Read time6 min read
Share

The forecast that undercuts the AI-is-getting-cheaper narrative

Gartner is forecasting that AI inference costs per agentic workflow will increase more than fivefold through 2028, a projection that directly contradicts the assumption most CIOs have built their AI business cases on for the past two years: that falling per-token prices would make AI progressively cheaper to run at scale. That assumption was reasonable given the trajectory of foundational model pricing, and it is also, according to Gartner, wrong for the workloads enterprises are actually scaling toward. Agentic workflows do not behave like the chatbot interactions most cost models were originally calibrated against.

Specifically, Gartner finds that routing tasks to agentic reasoning models increases provider inference costs by at least five times compared to basic chatbot interactions, before accounting for the added cost of more complex tasks that escalate well beyond that baseline. That is a substantial gap between the workload enterprises piloted, usually a simple assistant or chatbot, and the workload they are now planning to scale, multi-step agentic reasoning across business processes. Any budget built on pilot-stage inference costs is measuring the wrong thing for what production agentic deployment will actually require.

Gartner's inference paradox, explained

Gartner frames this as an inference paradox: individual token costs are declining as expected, yet overall AI expenses are rising sharply because sophisticated workflows require far more tokens than simple interactions did. Both halves of that statement are true simultaneously, and reconciling them is the part most enterprise cost models have gotten wrong. A cheaper per-token price multiplied against a much larger token volume per task can still produce a larger total bill, and agentic reasoning tasks generate exactly that larger volume, because they involve multiple reasoning steps, tool calls, and self-correction loops that a simple chatbot response never required.

The paradox matters because it breaks the intuitive mental model most finance and technology leaders have used to justify AI spending: that better model efficiency directly translates into lower AI costs over time. Gartner's data says efficiency gains instead get absorbed into more capable, more token-hungry models and workflows, keeping total spend on an upward trajectory even as the underlying unit economics improve. CIOs relying on the simpler mental model risk presenting a 2027 or 2028 budget that understates AI infrastructure costs by a meaningful margin.

Gartner attributes the fivefold increase to three converging trends. First, foundational model costs are genuinely improving, exactly as most CIOs expect. Second, that improved efficiency does not translate into savings, because it instead unlocks deployment of more powerful, and more expensive, models that would have been cost-prohibitive at previous efficiency levels. Third, complex agentic workflows consume significantly more tokens per task than the basic interactions most current usage and cost data reflects, because reasoning, tool orchestration, and multi-step task execution all generate token volume that simple chat exchanges do not.

Each of these trends is individually well understood in the industry, but Gartner's contribution is showing how they compound rather than offset one another. Falling token prices should, in isolation, reduce costs. Instead, that saved budget capacity gets reinvested into deploying more capable models for more complex workflows, which consume enough additional tokens to more than erase the per-token savings. The net effect, a more than fivefold cost increase through 2028, is what CIOs should now be modeling into multi-year AI budgets rather than treating as a worst-case scenario.

Why the business case built last year is already outdated

Most enterprise AI business cases built over the past eighteen months assumed a cost trajectory that this forecast directly contradicts. Those business cases typically extrapolated from pilot-stage chatbot or copilot deployments, where token consumption per interaction is relatively low and predictable, and assumed that scaling to more interactions simply meant multiplying that baseline. Gartner's forecast says the more relevant scaling factor is the shift from chatbot-style interactions to agentic reasoning workflows, which changes the per-task cost baseline itself by at least five times, independent of how many interactions the organization scales to.

That distinction matters enormously for any CIO whose 2026 budget approval relied on a chatbot-era cost model to justify a 2027 or 2028 agentic AI rollout. The business case was likely approved on numbers that will not hold once the organization moves from simple assistants to genuine agentic workflows handling multi-step business processes. Revisiting that business case now, before the gap surfaces as a budget overrun mid-deployment, is materially cheaper than explaining it to a board after the fact.

What Gartner is telling CIOs to do instead

Gartner analyst Will Sommer put the core recommendation directly: product leaders cannot rely on more efficient token economics to rationalize AI costs, because each successive model generation will necessitate more, and often more expensive, tokens. That framing shifts the cost management conversation away from waiting for the market to get cheaper and toward actively architecting for cost control. Gartner's specific recommendations center on building complex multimodel ecosystems rather than standardizing on a single model or vendor, and implementing sophisticated inference-tiering, routing, and orchestration strategies that match task complexity to the cheapest model capable of handling it.

Gartner also warns enterprises against defaulting to generic autonomous intelligence, the kind of broad, always-on agentic capability that sounds appealing in a vendor pitch but creates effectively unbounded inference costs once deployed at scale. The alternative Gartner describes is deliberately scoped agentic deployment, where advanced reasoning capability gets reserved for tasks that generate returns large enough to justify the cost premium, while simpler, cheaper models or workflows handle everything else. That scoping discipline is the practical difference between an AI program that scales sustainably and one that scales its bill faster than its returns.

The budget conversation to have before the next planning cycle

For CIOs heading into the next budget cycle, the immediate action is re-running AI cost projections using agentic-workflow token consumption rather than chatbot-era baselines, and presenting that revised number to finance and the board before it surfaces as a surprise variance. That conversation is uncomfortable in the short term, since it likely means telling stakeholders that the AI program they approved will cost meaningfully more than modeled. It is a substantially better conversation than having it forced after the fact by a budget overrun that erodes confidence in the technology organization's ability to forecast its own spending.

The longer-term action is building the inference-tiering and orchestration capability Gartner recommends before scaling agentic deployments further, so that cost control is architected into the system rather than bolted on after spend has already outpaced projections. Enterprises that treat this forecast as a planning input now will be positioned to scale agentic AI within a defensible budget. Enterprises that wait for the fivefold increase to show up in an actual invoice will be having this same conversation eighteen months from now, under considerably worse circumstances.

Tagged#news#digital-transformation#enterprise#cio#erp#strategy#governance#gartner#ai-inference-costs#agentic-ai-economics#it-budgeting#cost-management