Why AI Models Are Getting Larger and More Expensive to Run

Why AI Models Are Getting Larger and More Expensive to Run

Share

Artificial intelligence has improved dramatically over the last few years.

Modern AI models can write software, analyze documents, understand images, generate video, reason through complicated problems and increasingly operate computer applications.

But there is another side to this progress.

More capable AI models generally require more computing resources.

Training frontier models can require enormous clusters of accelerators running for extended periods. Running those models for users also requires substantial computing power, especially when models use long context windows, multimodal inputs, reasoning and tool-based workflows.

This creates a fundamental trade-off:

More capability
      ↓
More computation
      ↓
More hardware
      ↓
More electricity
      ↓
Higher operating cost

The AI industry is therefore facing a difficult question:

How much additional computing power is worth paying for another improvement in AI capability?

The answer is becoming increasingly important as AI moves from simple chatbots toward reasoning models, coding systems and autonomous agents.


AI Models Are Not Getting Larger for Only One Reason

AI Models Are Not Getting Larger for Only One Reason

It is common to describe AI progress as a race toward larger models.

That explanation is incomplete.

Modern AI systems can become more computationally demanding because of several different factors:

  • larger model architectures
  • longer context windows
  • more training data
  • additional reasoning computation
  • multimodal inputs
  • larger outputs
  • tool usage
  • agentic workflows
  • higher-quality training
  • repeated inference

Sometimes the model itself gets larger.

In other cases, the model may have a similar number of parameters but require substantially more computation to produce a high-quality answer.

This distinction is important.

AI compute demand

Model size
     +
Training compute
     +
Inference compute
     +
Context
     +
Reasoning
     +
Tools / agents
     ↓
Total infrastructure demand

So when people say AI is becoming more expensive, they are talking about a much larger problem than model parameters alone.


What Does “Larger AI Model” Actually Mean?

A model’s size is often described using parameters.

Parameters are learned numerical values that allow a neural network to transform input information into an output.

A simplified neural network might look like:

Input
  ↓
Layer 1
  ↓
Layer 2
  ↓
Layer 3
  ↓
Output

Each layer can contain many learned parameters.

Modern large language models contain enormous numbers of them.

A simplified comparison is:

Small model
     ↓
Millions / billions of parameters

Large model
     ↓
Billions / hundreds of billions of parameters

Frontier-scale system
     ↓
Potentially much larger overall architecture

However, parameter count alone does not tell you how expensive a model is to run.

Two models with similar parameter counts can have very different computational requirements.

Architecture matters.


Why More Parameters Can Increase Compute Requirements

During inference, the model needs to process the user’s input through its neural-network layers.

A simplified view is:

Prompt
  ↓
Tokenization
  ↓
Neural network
  ↓
Multiple layers
  ↓
Next-token prediction
  ↓
Next token
  ↓
Repeat

If the network contains more parameters, more computation and memory movement may be required.

The model also has to generate output one token at a time in many autoregressive systems.

For example:

Input
 ↓
Token 1
 ↓
Token 2
 ↓
Token 3
 ↓
Token 4
 ↓
...

A longer response therefore requires additional inference work.

This is one reason the cost of AI is closely connected to the number of tokens processed and generated.

Our article on AI inference and how AI models generate responses goes deeper into what happens during this process.


Training Is Extremely Expensive

Training Is Extremely Expensive

Inference receives a lot of attention because millions of users interact with AI systems.

But training frontier models can require enormous amounts of compute before a model is released.

The basic process looks like:

Huge dataset
     ↓
Training cluster
     ↓
Millions / billions of calculations
     ↓
Model update
     ↓
More training
     ↓
More model updates
     ↓
Finished model

Training requires accelerators to perform mathematical operations repeatedly.

Large clusters can contain thousands or even much larger numbers of processors.

This creates several costs:

  • accelerator hardware
  • electricity
  • networking
  • storage
  • cooling
  • data-center capacity
  • engineering
  • model evaluation
  • failed or experimental training runs

And training does not necessarily happen only once.

A company may train multiple versions of a model, run experiments, perform post-training and conduct reinforcement learning before releasing a production system.


Scaling Has Historically Driven AI Progress

One major reason models became larger is that scaling has repeatedly produced better results.

Researchers discovered that increasing combinations of:

  • model size
  • training data
  • compute

could improve model performance across many tasks.

The basic idea can be simplified as:

More data
   +
More parameters
   +
More compute
   ↓
More capable model

This does not mean that simply making a model larger guarantees better results.

Architecture, training methods and data quality matter enormously.

But scaling has remained one of the most important drivers of progress.


The New Problem: Reasoning Requires More Compute

Modern AI development has introduced another source of computation.

Some models are designed to spend additional computation reasoning through difficult problems before producing the final answer.

A traditional simplified workflow might look like:

Prompt
 ↓
Generate answer

A reasoning-oriented workflow can look more like:

Prompt
 ↓
Analyze problem
 ↓
Break into steps
 ↓
Evaluate possibilities
 ↓
Use tools
 ↓
Check result
 ↓
Generate answer

The model may therefore perform substantially more computation for a single user request.

This is particularly relevant to:

  • mathematics
  • programming
  • research
  • planning
  • complex analysis
  • multi-step problem solving

The result is a new trade-off:

Better reasoning can require more inference compute.


AI Models Are Not Just Getting Bigger — They Are Thinking More

This is one of the most important changes in modern AI.

Earlier discussions often focused on parameter count.

Today, another question matters:

How much computation does the model use while answering a particular request?

Consider two simplified systems.

Fast response

Question
   ↓
Model
   ↓
Answer

Extended reasoning

Question
   ↓
Reasoning
   ↓
Intermediate computation
   ↓
Tool use
   ↓
Verification
   ↓
More reasoning
   ↓
Answer

The second system may produce a better result on difficult tasks, but it can require substantially more computation.

That means AI companies increasingly have to optimize not only the model but also how much computation is spent on each request.


Context Windows Are Also Increasing

Another major source of compute and memory demand is the context window.

The context window determines how much information an AI model can process as part of a request.

Early language models worked with relatively small amounts of text.

Modern systems can process extremely large contexts containing:

  • documents
  • source code
  • conversations
  • datasets
  • images
  • tool results
  • search results

Our article on AI model context windows explains how context length affects AI systems.

A simplified example:

Short context

User request
     ↓
Small amount of information
     ↓
Model

Compared with:

Long context

User request
     ↓
Large document
     ↓
Source code
     ↓
Previous conversation
     ↓
Tool results
     ↓
Model

The second system has much more information to process.

Long-context AI therefore creates additional demands on memory and computation.


The KV Cache Makes Long Conversations More Complicated

Large language models often use something called a KV cache during inference.

KV stands for:

  • Key
  • Value

During generation, the model can store information from previously processed tokens rather than recalculating everything from scratch for every new token.

A simplified view looks like:

Previous tokens
      ↓
Key / Value cache
      ↓
New token generation

This improves efficiency.

But the cache itself consumes memory.

As context becomes longer, the amount of information that needs to be stored can increase substantially.

For large-scale AI serving, that creates another hardware requirement:

Long context
    ↓
Larger KV cache
    ↓
More memory
    ↓
Higher memory bandwidth requirements

This is one reason AI inference increasingly depends on high-bandwidth memory and carefully designed accelerator architectures.


Memory Can Become a Bigger Problem Than Raw Compute

AI workloads are not only about mathematical operations.

They are also about moving enormous amounts of data.

A processor may be capable of performing calculations extremely quickly, but it still needs to access:

  • model weights
  • activations
  • KV cache
  • input tokens
  • output data

The simplified architecture looks like:

                 AI accelerator
                      │
              ┌───────┴───────┐
              │   Compute     │
              └───────┬───────┘
                      │
                 Memory system
                      │
                     HBM
                      │
                Model + cache

This is why modern AI chips rely heavily on high-bandwidth memory.

It is also one reason AI hardware is evolving so quickly.

Our article on how AI chips work explains the relationship between compute, memory and AI workloads.


Multimodal AI Adds More Computation

Text is only one type of information modern AI systems process.

Many models now work with combinations of:

  • text
  • images
  • audio
  • video
  • documents

This is known as multimodal AI.

Consider a simple text request:

Text
 ↓
Language model
 ↓
Answer

Now consider a video analysis request:

Video
 ↓
Frames
 ↓
Visual processing
 ↓
Audio
 ↓
Speech processing
 ↓
Text representation
 ↓
Reasoning
 ↓
Answer

The second workflow can involve dramatically more data.

Video is particularly expensive because it contains many frames over time.

As AI becomes more multimodal, infrastructure has to process more information per request.


AI Agents Can Multiply Inference Demand

AI agents introduce another major source of cost.

A normal chatbot interaction might involve one primary model response.

An agent may call a model repeatedly.

For example:

User
 ↓
Agent reasoning
 ↓
Search
 ↓
Model call
 ↓
Read result
 ↓
Model call
 ↓
Write code
 ↓
Model call
 ↓
Run code
 ↓
Model call
 ↓
Check result
 ↓
Final response

A single user request could therefore trigger many inference operations.

Our article on AI agents and how they work explores this transition from simple responses toward multi-step AI workflows.

This creates an important economic effect.

Even if the cost of an individual model call falls, an application can still become expensive if it performs many calls per task.


Why AI Data Centers Are Becoming So Large

The growing computational requirements of AI are directly affecting physical infrastructure.

A traditional data center might contain a mixture of:

  • CPUs
  • storage
  • networking
  • virtualization infrastructure
  • enterprise applications

An AI data center can be much more accelerator-heavy.

Traditional data center

CPU
CPU
Storage
Network
CPU
Storage


AI data center

GPU / AI accelerator
GPU / AI accelerator
GPU / AI accelerator
GPU / AI accelerator
High-speed network
High-bandwidth memory
Advanced cooling

Our article on how AI data centers are different from traditional data centers explains this difference in more detail.

As model requirements increase, companies need more accelerators and more supporting infrastructure.


Power Consumption Is Becoming a Major Constraint

Every AI calculation requires electricity.

At small scale, the cost may not seem significant.

At global scale, it becomes a major infrastructure issue.

The chain looks like:

More AI computation
       ↓
More accelerators
       ↓
More electricity
       ↓
More cooling
       ↓
Larger power infrastructure
       ↓
Larger data centers

This is one reason AI companies are investing in custom accelerators, more efficient chips and better data-center designs.

Our article on why AI is increasing data-center power demand explores this issue in greater detail.

The challenge is not simply producing more electricity.

The electricity has to reach the accelerator efficiently, and the resulting heat has to be removed.


Why Cooling Becomes More Difficult

AI accelerators can consume substantial amounts of power.

Most of that electrical energy eventually becomes heat.

The basic relationship is:

Electrical power
      ↓
AI computation
      ↓
Heat
      ↓
Cooling system

As accelerator density increases, air cooling can become less practical for some high-density systems.

This is one reason liquid cooling is becoming increasingly important for advanced AI infrastructure.

The result is that a more powerful AI model can indirectly require investment in:

  • cooling systems
  • pumps
  • heat exchangers
  • facility plumbing
  • rack design
  • power distribution

AI model development is therefore affecting physical data-center architecture.


More AI Compute Means More Networking

Large AI models are often distributed across many accelerators.

Instead of one processor handling everything:

One accelerator
     ↓
Model

a large system may look like:

Accelerator ─── Accelerator
     │               │
     ├── Network ────┤
     │               │
Accelerator ─── Accelerator
     │               │
     └── Network ────┘

These processors have to exchange data extremely quickly.

As models become larger and clusters grow, high-speed networking becomes increasingly important.

This is why modern AI infrastructure involves not only GPUs or custom AI chips but also:

  • high-bandwidth interconnects
  • switches
  • optical networking
  • RDMA
  • specialized fabrics

The accelerator is only one component of the system.


Why AI Companies Are Building Their Own Chips

The rising cost of AI models is one reason companies are designing custom processors.

A company running enormous AI workloads may benefit from hardware optimized for:

  • inference
  • memory access
  • model architecture
  • specific numerical formats
  • networking
  • power efficiency

The basic idea is:

General-purpose accelerator
          ↓
Good at many workloads


Custom AI accelerator
          ↓
Optimized for known AI workloads

Companies such as Google, Microsoft, Amazon, Meta and OpenAI are pursuing custom silicon strategies for precisely this broader reason.


AI Models Are Also Getting More Expensive Because They Are Used More

There is another important factor that is easy to overlook.

Even if a model’s architecture remained unchanged, the cost of operating AI could still increase if usage grows rapidly.

Imagine:

1 million requests/day

Then:

10 million requests/day

Then:

100 million requests/day

The model did not necessarily become larger.

But the infrastructure requirement increased dramatically.

This creates two different forms of scaling:

Model scaling
     ↓
More computation per request


Usage scaling
     ↓
More requests per second

The AI industry is dealing with both simultaneously.


Training Cost and Inference Cost Are Different

It is useful to separate the two.

Training

Training is a large upfront computational expense.

Massive dataset
      ↓
Large cluster
      ↓
Long training process
      ↓
Finished model

Inference

Inference is the continuing operational expense.

Users
 ↓
Requests
 ↓
Model inference
 ↓
Responses
 ↓
More requests
 ↓
More inference

A company can spend enormous amounts training a model and then spend even more over time operating it for millions of users.

This changes the economics of AI products.

The question is no longer only:

How much did it cost to train the model?

It is also:

How much does every useful AI task cost to execute?


The Cost of a Token Matters

Large language models operate on tokens rather than ordinary words.

A token may represent:

  • part of a word
  • a complete short word
  • punctuation
  • a number
  • another piece of text

The exact tokenization depends on the model.

If an application processes enormous amounts of text, token consumption can become a useful way to measure computational workload.

A simplified cost model is:

Input tokens
      +
Output tokens
      +
Reasoning / processing
      ↓
Compute consumption
      ↓
Infrastructure cost

This is why AI providers often publish pricing based on tokens.

However, token price is only one part of the real cost.

Infrastructure utilization, memory, networking and hardware efficiency also matter.


Bigger Does Not Always Mean Better

It is important not to assume that larger models are automatically superior.

A smaller model can be preferable when the task is simple.

For example:

Simple classification
       ↓
Small / efficient model
       ↓
Fast + inexpensive

While:

Complex research
       ↓
Larger reasoning model
       ↓
More computation
       ↓
Potentially better performance

The ideal model depends on the workload.

This is why AI companies increasingly offer families of models rather than one model for everything.


Mixture-of-Experts Changes the Equation

Another important technique is Mixture-of-Experts, commonly called MoE.

Instead of activating every part of a very large model for every token, an MoE architecture can route different inputs through different subsets of the network.

A simplified representation:

Input
  ↓
Router
  ↓
┌───────┬───────┬───────┬───────┐
Expert A Expert B Expert C Expert D
└───────┴───────┴───────┴───────┘
          ↓
       Selected
       experts
          ↓
        Output

This can allow a model to have a very large total parameter count while activating only part of the network for a particular token.

That creates an important distinction:

Total parameters ≠ parameters used for every computation.

MoE architectures can therefore change the relationship between model size and inference cost.

They are one example of how AI researchers are trying to increase capability without making every computation proportionally more expensive.


Quantization Can Reduce AI Costs

Another important technique is quantization.

AI models often use numerical representations such as:

  • FP32
  • FP16
  • BF16
  • INT8
  • lower-precision formats

Using lower precision can reduce memory usage and potentially improve inference efficiency, although the impact depends on the model and hardware.

A simplified example:

Higher precision
      ↓
More memory
More bandwidth
More compute


Lower precision
      ↓
Less memory
Less data movement
Potentially faster inference

The challenge is maintaining model quality.

Reducing precision too aggressively can affect accuracy or numerical stability.

This creates another engineering trade-off:

Model quality
      ↕
Compute efficiency

Modern AI infrastructure increasingly tries to find the right balance.


Why Inference Optimization Matters So Much

Suppose an AI company has a model used by millions of people.

Even a small efficiency improvement can have a large impact.

Imagine reducing the compute required for every request by a small percentage.

At enormous scale:

Small optimization
       ×
Millions / billions of requests
       ↓
Large infrastructure savings

This is why companies invest heavily in:

  • optimized kernels
  • quantization
  • caching
  • speculative decoding
  • model distillation
  • custom accelerators
  • better scheduling
  • efficient serving systems

The goal is not always to make the model smaller.

Sometimes the goal is to do the same amount of useful work with less computation.


The AI Industry Is Facing a Capability-Cost Trade-Off

The central challenge can be represented like this:

Higher capability
       ↑
       │
       │
       │
       └──────────────→ Higher compute
                         and cost

AI companies want to move the capability curve upward without allowing costs to rise at the same rate.

That means future AI progress will depend heavily on efficiency.

The industry needs:

  • better model architectures
  • better training methods
  • more efficient inference
  • better chips
  • better memory systems
  • better networking
  • better cooling
  • better software

The model alone is not enough.


What Happens If Models Keep Getting More Expensive?

There are several possible consequences.

More specialized hardware

AI companies will continue designing chips around specific workloads.

More efficient models

Researchers will look for ways to achieve similar performance using less compute.

Model routing

Applications may automatically send simple requests to smaller models and difficult requests to larger ones.

User request
     ↓
Complexity detection
     ↓
┌───────────────┐
│               │
Simple        Complex
 ↓               ↓
Small model    Large model

More local AI

Some workloads may move from cloud servers to phones, PCs and edge devices when smaller models become capable enough.

More efficient data centers

Power delivery, networking and cooling will become increasingly important parts of AI infrastructure.


The Future May Not Be About One Giant Model

One possible direction for AI is a collection of specialized models working together.

For example:

                 AI System
                    │
       ┌────────────┼────────────┐
       ↓            ↓            ↓
   Small model   Reasoning     Vision model
       │           model           │
       ↓            ↓              ↓
   Simple tasks  Complex tasks   Images
                    │
                    ↓
                 AI Agent

The advantage is that the system does not have to use the most expensive model for every request.

This could become increasingly important as AI applications scale.


Why AI Models Are Getting Larger and More Expensive to Run

The reason is not simply that companies are chasing bigger parameter counts.

AI systems are becoming more capable in several dimensions at once.

They are handling:

  • more data
  • longer contexts
  • more modalities
  • more reasoning
  • more tools
  • longer workflows
  • more users
  • more complex tasks

All of those factors can increase computational requirements.

The bigger picture looks like this:

More capable models
        ↓
More computation per task
        ↓
More inference demand
        ↓
More accelerators
        ↓
More memory + networking
        ↓
More electricity
        ↓
More cooling
        ↓
Larger AI infrastructure

That is why the AI race is increasingly becoming an infrastructure race as well.

The next generation of AI will not be determined only by who can build the most capable model.

It will also depend on who can train, serve and scale that intelligence efficiently.

The companies that solve that problem can potentially make advanced AI available to more users at lower cost.

And that may ultimately matter just as much as making the models more intelligent.

Scroll to Top