The falling price of AI inference creates an appealing assumption that cheaper models should make AI products cheaper to operate.
However, Gartner expects a different outcome.
Inference costs per agentic workflow are projected to increase more than fivefold through 2028, driven by the growing amount of computation required for AI systems that can execute multistep tasks.
Gartner calls this the “Inference Paradox.” Improvements in model economics make more capable applications financially possible, yet those applications demand more inference to operate.
The result is a new margin problem for AI products, especially those moving from assistive features toward agents that can reason and act across a workflow.
Why It Matters: Token prices are becoming a less useful proxy for the economics of an AI product. The more consequential question is how much intelligence a system consumes to complete useful work. That puts architecture, model choice and workflow design closer to the economics of AI deployment.
- Agentic Workflows Change the Unit of AI Cost: A chatbot typically processes a request and produces a response, while an agent may require repeated rounds of reasoning throughout a task. Gartner estimates that using an agentic reasoning model can increase provider inference costs by at least five times compared with a basic chatbot interaction, with more demanding workflows costing even more. This makes the economics of the full workflow more important than the price of individual tokens.
- Model Efficiency Is Helping Developers Build More Ambitious Products: Foundation model economics are improving, giving AI applications access to greater capability for less money at the unit level. Those gains can be reinvested into applications that perform more sophisticated work and consume more tokens along the way. Gartner expects innovation in AI capability to move faster than reductions in token costs, which helps explain why total inference spending can keep climbing even while individual units of computation become cheaper.
- There Is No Clear One-Model Answer to the Cost Problem: Gartner does not expect a reliable, economical model capable of serving every workload equally well. AI products may need multimodel ecosystems that assign different types of work to different models. A lightweight model might handle routine processing, while an expensive reasoning model is used where added capability has enough value to justify its inference bill.
- Routing and Orchestration Become Part of the Margin Equation: Gartner identifies inference tiering, routing and orchestration as ways to calibrate the intelligence used for a given task. These mechanisms can limit unnecessary use of expensive reasoning models and match computational effort with task requirements. The implication reaches into product design. Autonomy does not have to be applied uniformly across every feature or workflow. Gartner warns that defaulting to generic autonomous intelligence can produce costs orders of magnitude above an optimized product ecosystem.
- Higher Inference Spending Ultimately Needs Higher Economic Returns: Expensive reasoning can still make sense when the work it performs creates enough value. Gartner says advanced AI will need returns far above those of basic models, or substantial optimization of the inference supporting it. That places ROI at the center of the agentic AI equation. The relevant measure becomes the relationship between what an agent costs to complete a task and what completing that task is worth.

