What Happens Inside an AI Data Center

What Happens Inside an AI Data Center

Share

From the outside, an AI data center looks like any other big, windowless warehouse. Inside, almost everything about it is built around one goal: keep a huge number of very expensive processors busy, fed, powered, and cool. The easiest way to understand it is to follow what happens when you type a question into a chatbot.

A request comes in

Your prompt travels over the internet to a front door made of ordinary-looking servers: load balancers, security checks, and routing software. Their job is to decide which group of accelerators should handle the request. Nothing about this part is exotic, but it matters, because the expensive hardware behind it only makes money when it is working on something.

Your prompt travels over the internet to a front door made of ordinary-looking servers: load balancers, security checks, and routing software. Their job is to decide which group of accelerators should handle the request. Nothing about this part is exotic, but it matters, because the expensive hardware behind it only makes money when it is working on something.

The prompt is then handed to a pool of GPUs. Serving a modern model usually means splitting the job in two. First the model reads your whole prompt in one burst of heavy computation. Then it generates the answer one token at a time, which leans less on raw compute and much more on how fast the chip can pull the model’s weights and its working memory out of memory. Many operators now run those two stages on separate groups of machines, so each can be sized for what it is actually limited by.

Inside the rack

The machines doing this work are not stand-alone servers. In NVIDIA’s current reference design, a single rack holds 72 GPUs and 36 CPUs spread across 18 compute trays, tied together by NVLink switch trays so the whole rack behaves like one big accelerator. That matters because a large model rarely fits on one chip. It is cut into pieces that live on different GPUs, and those GPUs have to trade results constantly while your answer is being produced.

Between racks, a separate network, either InfiniBand or Ethernet, links the rooms of equipment together. We looked at why this layer has become so important in why AI data centers need faster networking. If a link is slow or congested, the GPUs on either side simply wait, and waiting is the most expensive thing a data center can do.

Training is a different life

While some racks answer user requests, others may be in the middle of training a new model, which looks quite different. Thousands of GPUs work on a single job for weeks, reading enormous datasets from fast storage and periodically saving checkpoints, snapshots of the model’s progress, so that a failure doesn’t wipe out days of work. With so many components running flat out, something is always going wrong somewhere. Much of the software in the building exists to notice a failed chip or link, move the job around it, and carry on.

The part nobody sees: heat and electricity

All of this turns electricity into heat, and the amount is large. Racks like these draw far more power than air can carry away, so operators run liquid through cold plates attached directly to the chips. Oracle’s engineers have said that air cooling is no longer viable at this scale, so heat is pulled out with coolant loops, pumps, and heat exchangers that look more like an industrial plant than an office server room.

The electrical side is similar. Power arrives from the grid through transformers and switchgear, is converted and distributed down to each rack, and is backed up by batteries and generators in case the supply blips. The people who work in these buildings spend much of their time watching dashboards for temperature, power draw, and failing hardware, and swapping parts before a small fault becomes a stalled training run.

Why it all matters

Put together, an AI data center is less like a warehouse of computers and more like a single machine whose parts happen to be spread across a campus. The chips, the network, the cooling, and the power all have to keep pace with each other, and whichever one lags sets the speed for everything else. That’s why the biggest bottleneck right now is often not the processors themselves but the electrical equipment around them, a problem we explored in why data center power is becoming a technology bottleneck.

Scroll to Top