A few years ago, a cluster of a few thousand GPUs counted as a supercomputer. Today, AI companies are assembling machines with hundreds of thousands of accelerators in a single site, drawing power on the scale of a small city. The numbers are staggering, but the logic behind them is straightforward: in modern AI, the size of the machine you can build largely decides what models you can train. For more on the infrastructure behind it, see our AI infrastructure and data center coverage.

How big the clusters have become
The leading projects now measure themselves in gigawatts. Sherwood News’s roundup lists OpenAI’s Stargate partnership with Oracle and SoftBank as a $500 billion plan totaling 5.5 gigawatts and about 2 million GPUs, alongside a separate NVIDIA and OpenAI partnership covering 10 gigawatts and 4 to 5 million GPUs. One Stargate report says the flagship Abilene, Texas campus is set to reach 1.2 gigawatts and over 450,000 GB200 GPUs, though the same source dates from earlier in 2026, so check current progress.
xAI’s Colossus in Memphis is the other headline project. A detailed March 2026 analysis says the site began in 2024 when xAI filled a former appliance factory with 100,000 H100 GPUs in 122 days, and by the end of March the combined cluster had about 1 gigawatt of compute across roughly 770,000 GPUs, with Musk projecting around 2 gigawatts. Other reports put the GPU count lower, for example 555,000, so treat the totals as approximate. Project Rainier, which we covered in how Amazon is building its own AI silicon, is a comparable effort built on custom chips, with nearly 500,000 Trainium2 processors for Anthropic.
Reason 1: Scale drives capability
The central reason is the way modern models improve. Researchers have long observed that model performance tends to rise predictably with more compute, more data, and more parameters, a pattern known as scaling laws. Whether or not those gains continue forever, AI labs have so far found that bigger training runs usually produce more capable models. A lab that cannot assemble enough compute risks falling behind rivals that can.
That explains why spending has become so concentrated. As we described in why Big Tech is spending billions on AI infrastructure, companies treat compute as a competitive necessity, and the labs that don’t own hyperscale infrastructure make multi-year, multi-gigawatt commitments instead. On NVIDIA’s latest earnings call, the company said OpenAI’s existing and planned commitments represent roughly 12 gigawatts of NVIDIA compute through 2030.
Reason 2: One big job beats many small ones
Training a frontier model isn’t like running thousands of independent tasks. The GPUs must work on the same problem at the same time, exchanging results constantly. That favors one tightly connected cluster over many scattered ones. The xAI analysis notes that newer rack systems raise the coherent scale-up domain from 8 to 72 GPUs in a single generational step, and Google’s latest training pods link 9,600 chips in one system.
This is why the network matters as much as the chips, a point we explored in why NVIDIA has become more than a GPU company. Distance adds delay, and delay wastes expensive GPU time, so labs try to pack as many accelerators as possible into the same campus.
Reason 3: Speed to market
Bigger clusters finish training runs faster. If a model takes months on a small cluster, a larger one can cut that sharply, and in a market where rivals release new models every few months, time matters. xAI’s build is a vivid example: the analysis says its first Colossus 2 cluster of about 110,000 GB200 GPUs came online in 91 days, followed by a second of similar size in 64 days. The speed of construction is itself part of the strategy.
Extra capacity also lets labs run several experiments in parallel, test new training methods, and spend more on post-training steps that improve reasoning and safety.
Reason 4: Training and serving compete for the same hardware
Clusters don’t only train models. The same hardware is needed to serve them to users, and demand there is growing quickly, as we described in why AI inference is becoming the next big infrastructure challenge. Agents and reasoning models use many more tokens per task, so a lab that wins customers needs capacity for both building new models and answering requests. Building big clusters is partly a way to avoid choosing between the two.
Reason 5: Capacity is scarce
Many buyers want the same chips, and supply is limited. As we explained in why semiconductor supply chains matter for AI, advanced chips depend on a small number of manufacturing, packaging, and memory suppliers. NVIDIA says it expects supply to remain a bottleneck through fiscal 2028. A company that can secure large allocations early gets a head start, which is why long-term deals and early reservations have become common.
What it takes to build one
Power. The hardest part is electricity. xAI’s approach illustrates the extremes: an industry tracker counts 1,498 MW of gas turbine capacity at its Memphis-area campus, and we covered the wider power picture in why AI data centers need much more power.
Cooling and design. At these densities, racks need liquid cooling and purpose-built facilities, as we explained in why AI is changing the design of data centers.
Financing. The costs are huge. NVIDIA is involved on this front too, and one summary says it committed credit support for AI-cloud infrastructure including up to $105 billion tied to an Ohio OpenAI campus. We covered NVIDIA’s expanding role as a financier in why NVIDIA has become more than a GPU company.
Reliability. With hundreds of thousands of components, failures are routine, and software must keep a run going when individual chips or links drop out. Commentary on the largest clusters points to synchronization and fault tolerance as a growing engineering challenge, though much of it comes from secondary sources.
The risks
The approach has clear downsides. Clusters are expensive and can become outdated quickly as new chips arrive. They depend on grids, permits, and local acceptance, and projects have drawn controversy over emissions and utility costs. There’s also an open question about returns: if scaling gains slow, or if smaller, more efficient models close the gap, the largest clusters could turn out to be less decisive than their builders hope. And the financing structures, which tie suppliers, labs, and lenders together, mean trouble in one place can spread.
Data quality is another caution for anyone writing about this field. Reports on cluster sizes vary widely, and many popular summaries rely on social posts or unverified claims, so cite company statements or filings where possible.
The bigger picture
AI companies are building massive GPU clusters because scale has been the most reliable path to better models, because large training runs need tightly connected hardware, and because compute is scarce enough that those who secure it first gain an advantage. The same race links chips, memory, networking, power, and finance, which is why it touches every part of the story we’ve covered. Follow our AI hardware coverage to see whether the biggest clusters keep paying off.
SiliconeUpdate.com is a technology news platform that publishes updates and informational content related to silicon technology, software, artificial intelligence, and emerging technologies.
All articles published on this platform are attributed to SiliconeUpdate.com instead of individual authors. Content is presented in a neutral, informational format without personal opinions.
—
Content Publishing
SiliconeUpdate.com publishes news and updates based on publicly available information, official announcements, and industry developments. The focus is on clarity, relevance, and timely reporting.
—
Editorial Control
All editorial decisions, updates, and content management are handled at the platform level. No individual human or AI identity is presented as the author of articles.
—
Contact
For editorial communication or general queries, contact:
Email: neemasharma@gmail.com