AI Inference Costs Set to Rise More Than Fivefold Through 2028: Gartner

As AI products evolve from assistive features to multistep execution, product leaders face a new margin challenge with falling model prices subsidizing more complex workflows and escalating total AI costs.
AI Inference Costs Set to Rise More Than Fivefold Through 2028: Gartner
Published on
2 min read

AI inference costs per agentic workflow will increase more than fivefold through 2028 according to Gartner, Inc.,

As AI products evolve from assistive features to multistep execution, product leaders face a new margin challenge with falling model prices subsidizing more complex workflows and escalating total AI costs. As a result, inference cost management has become a top priority for product leaders.

Gartner has identified three fundamental trends driving token economics:

  1. Foundational model cost economics are rapidly improving.

  2. Improved AI efficiency is unlocking the deployment of more powerful, and more expensive, models to enable higher-value, more sophisticated AI applications.

  3. More sophisticated AI workflows use far more tokens than simple chatbot interactions, driving higher overall inference costs.

These dynamics mean that tokens are becoming more cost-efficient, but not as quickly as AI capabilities and the costs associated with those capabilities are increasing (see Figure 1). The rate of innovation is outpacing the cost curve.

This is the Inference Paradox, defined as better unit economics escalating the overall cost of AI without providing a clear pathway to commensurate and predictable value.

“The harsh economics of the Inference Paradox are exemplified by the differences between a simple chatbot and an AI agent,” said Sommer. “Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate, and question itself.”

All of these responsibilities add up. Compared to a basic chatbot interaction, routing a task to an agentic reasoning model increases provider inference costs by at least five times, and often much more as task complexity grows.

Ensuring ROI from advanced AI, like reasoning agents, demands exponentially higher returns relative to basic models, or highly optimized inference-tiering, routing and orchestration to calibrate complex tasks relative to more cost-efficient intelligence. Both of these outcomes are eminently possible but will require significant effort across complex workflows.

“Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems,” said Sommer.

𝐒𝐭𝐚𝐲 𝐢𝐧𝐟𝐨𝐫𝐦𝐞𝐝 𝐰𝐢𝐭𝐡 𝐨𝐮𝐫 𝐥𝐚𝐭𝐞𝐬𝐭 𝐮𝐩𝐝𝐚𝐭𝐞𝐬 𝐛𝐲 𝐣𝐨𝐢𝐧𝐢𝐧𝐠 𝐭𝐡𝐞 WhatsApp Channel now! 👈📲

𝑭𝒐𝒍𝒍𝒐𝒘 𝑶𝒖𝒓 𝑺𝒐𝒄𝒊𝒂𝒍 𝑴𝒆𝒅𝒊𝒂 𝑷𝒂𝒈𝒆𝐬 👉 FacebookLinkedInTwitterInstagram

logo
DIGITAL TERMINAL
digitalterminal.in