Why AI Inference Is Becoming the Next Big Infrastructure Challenge

Why AI Inference Is Becoming the Next Big Infrastructure Challenge

Share

For the first phase of the AI boom, the headline problem was training: gathering thousands of chips to build ever-larger models. The next problem looks different. Once models are built, they have to answer billions of requests, all day, for every user and every automated agent. That work is called inference, and it is quickly becoming the largest, least predictable, and most expensive part of running AI. For more on the industry behind it, see our AI infrastructure and data center coverage.

For the first phase of the AI boom, the headline problem was training: gathering thousands of chips to build ever-larger models. The next problem looks different. Once models are built, they have to answer billions of requests, all day, for every user and every automated agent. That work is called inference, and it is quickly becoming the largest, least predictable, and most expensive part of running AI. For more on the industry behind it, see our AI infrastructure and data center coverage.

Training builds the model, inference pays the bills

Training is a one-time (if enormous) cost, while inference recurs every time someone uses the product. One 2026 technical walkthrough puts inference at roughly two-thirds of all AI compute this year, up from about one-third in 2023, and says the industry rule of thumb is that it eventually makes up 80 to 90 percent of a deployed model’s lifetime compute cost. Those are estimates from a single technical blog, so treat the exact shares as approximate, but the direction matches what the chipmakers themselves are saying.

That shift is reshaping hardware roadmaps. We covered how Google responded in why Google designs its own AI hardware: its newest generation splits into separate training and inference chips, and CNBC reported Google’s claim that the inference chip is 80 percent better than its predecessor. NVIDIA made a similar move, as we described in why NVIDIA is expanding beyond GPUs, pairing its Rubin platform with a Groq-based inference accelerator.

Agents multiply the demand

The biggest reason inference is straining infrastructure is that usage per task is exploding. A chatbot answers a question; an AI agent plans, calls tools, reads results, and tries again. Citing Gartner’s March 2026 analysis, one cost guide says agentic models need 5 to 30 times more tokens per task than a standard chatbot query, and that inference already represents 85 percent of enterprise AI budgets in the teams it describes.

Reasoning models add to the load. TrendForce notes that test-time scaling, where models “think” for longer, is driving more than five times as many tokens per year, with inference placing three demands on hardware: more queries per second, longer context windows, and more steps in agent loops. Nothing about this resembles the neat, scheduled workloads of training.

The cost paradox

Per-token prices are falling quickly, which sounds like good news. But a research summary on inference economics describes a paradox: unit costs fall by more than 80 percent while total spending rises, because cheaper tokens make more ambitious agent workflows affordable, and those workflows consume far more of them. The same source cites IDC’s warning that even large organizations with dedicated FinOps teams can underestimate AI infrastructure costs by as much as 30 percent.

The lesson for infrastructure planners is that efficiency gains don’t reduce total demand. They enable more of it, which is the same dynamic behind the power forecasts in why AI data centers need much more power.

The real bottleneck is memory

Inference doesn’t stress chips the way training does. The walkthrough explains that the first phase, prefill, is compute-bound and sets the time to first token, while the second, decode, is memory-bandwidth-bound: to generate each token, the GPU must stream the model’s weights and the growing context out of memory again.

That growing context is the KV cache, the working memory of a conversation. TrendForce explains that when GPU memory isn’t enough for long-context, high-volume workloads, the system must discard the cache and recompute it, which raises latency and total cost of ownership. A storage-industry analysis makes a related point about economics: cached reads are priced much lower than uncached ones, but the number of cache slots is capped because the cache sits in DRAM, which is finite. Both are vendor-aligned sources, but they describe a real engineering problem.

The response is a new memory hierarchy. Spilling the cache into CPU memory or SSDs saves expensive GPU memory but can add latency, as one infrastructure guide points out. NVIDIA’s answer includes a BlueField-4-powered inference context memory storage platform for long-running agents. This ties directly to the shortages we described in why semiconductor supply chains matter for AI: inference pulls on HBM, DRAM, and storage at the same time.

Splitting the work: new serving architectures

Because prefill and decode stress hardware differently, engineers now run them on separate machines. The walkthrough describes this as disaggregated serving, splitting compute-bound prefill and memory-bound decode into different GPU pools that scale independently, with the KV cache streamed between them. NVIDIA’s Dynamo software is built around this idea, and NVIDIA says Dynamo 1.0 boosts Blackwell inference by up to 7 times. That is a company claim, not an independent benchmark.

The Crusoe tokenomics article, a vendor blog, adds that agentic prompts tend to be prefill-heavy because tool outputs become input for the next turn, making each step more expensive than the last. Better software can help, but it cannot remove the underlying demand.

Latency and the CPU question

Speed matters differently for agents. Each step depends on the one before it, so delays compound. The same Crusoe piece notes that agentic workflows require infrastructure that handles growing context windows, stateful multi-turn execution, sustained high utilization, and compounding latency.

That helps explain a surprising trend: CPUs are back in focus. TrendForce says agentic AI creates a new CPU and memory market, pointing to NVIDIA’s Vera CPU, which NVIDIA describes as built for agentic AI. Agents run many tool calls, simulations, and checks that don’t need a GPU, so the ratio of CPUs to accelerators is shifting.

Specialized chips for inference

The industry’s hardware response is to build chips tuned to inference. Google’s eighth-generation TPU 8i triples on-chip memory according to Futurum’s analysis, and Microsoft’s Maia 200 was designed for inference economics, both of which we covered in how cloud companies are becoming AI hardware companies. NVIDIA claims its Groq-based accelerator working with Vera Rubin delivers up to 35 times higher inference throughput per megawatt than prior designs, again a company figure. And as we described in how Apple is bringing more AI processing to its devices, some inference is moving to phones and laptops to relieve the load on data centers.

What this means for infrastructure

Three consequences stand out. First, capacity planning is harder, because inference demand depends on how users and agents behave, not on a training schedule. Second, power is the ceiling: the cheaper tokens get, the more of them are consumed, so total electricity use keeps rising. Third, the winners may be decided by utilization: the operator that keeps expensive chips busy, caches effectively, and routes work to the right hardware can serve the same tokens at a much lower cost. This is part of the reasoning behind the heavy spending we examined in why Big Tech is spending billions on AI infrastructure.

The risks and unknowns

Much of the data on inference is still thin. Shares of compute, token multipliers, and speedups often come from vendors, blogs, or single analyses rather than audited sources. Forecasts depend on agents becoming reliable and widely adopted; if adoption stalls, the capacity being built could be underused. And a technical fix, such as a better caching method or a more efficient model architecture, could change the picture quickly.

The bigger picture

Inference is becoming the next big infrastructure challenge because it’s where AI meets real usage. Every new agent, longer context window, and reasoning step adds tokens, and every token has to be produced fast, cheaply, and close to the user. Training made AI possible; inference decides whether it can be delivered at scale and at a price people will pay. Follow our AI hardware coverage to track how chips, memory, and software evolve to meet it.

Scroll to Top