When people talk about AI infrastructure, they usually talk about chips. But a cluster of thousands of GPUs is only as fast as the network that connects them. If data can’t move between accelerators quickly and predictably, the expensive chips spend their time waiting. That is why networking has become one of the fastest-growing and most contested parts of the AI build-out.

Training turns a network into part of the computer
A traditional data center runs many independent jobs, so a slow link affects only one of them. AI training is different. Thousands of GPUs work on a single model at the same time and must repeatedly exchange results, often after every step. The slowest link in the group sets the pace for everyone. In effect, the network stops being a connection between computers and becomes part of one big computer.
That’s the logic behind the rack-scale designs we described in how GPU clusters are built for AI workloads. Inside an NVL72 rack, 72 GPUs communicate over a dedicated NVLink fabric, and a separate network ties racks together. Oracle’s engineers describe it as the scale-up network within a rack and an InfiniBand or RoCE network as the scale-out layer across racks. Each layer needs to be fast, and the more GPUs in a job, the more traffic crosses the slower outer layer.
Bandwidth: bigger chips need bigger pipes
Each generation of accelerator is more powerful, so it needs more bandwidth to avoid starving. One 2026 networking report says the 800G generation, using PAM4 signaling at 100G per lane, is now mainstream for backend AI fabrics. A separate analysis expects 800G to remain the primary data rate for large-scale deployment around 2026, with 1.6T demand tied to the maturity of 200G-per-lane technology and to architectural choices inside big clusters. It also notes that 800G-plus optical transceivers are projected to exceed 60 percent of shipments in 2026.
The switches are keeping pace. Broadcom’s Tomahawk 6 delivers 102.4 Tbps, twice the throughput of its previous generation, and Broadcom says it targets both scale-up and scale-out AI networks. Reports differ on exactly when it reached volume, with some saying sampling began in late 2025 and others that production shipments were announced in March 2026, so check Broadcom’s latest status before quoting a date. NVIDIA is also moving, with Spectrum-6 and Quantum-X Photonics switches planned for 2026.
Latency and the long tail
Bandwidth is only part of it. AI traffic is bursty and synchronized: many GPUs send data at once, then all wait for the slowest message. That makes tail latency, the delay of the unluckiest packets, more important than average speed. One networking article argues that as models scale, traditional Ethernet struggles with micro-packet inefficiencies and jitter, and describes Broadcom’s Tomahawk Ultra as aiming for 250 nanoseconds per hop with streamlined headers. That’s a vendor-aligned description, but it shows where the engineering effort is going.
Congestion control, load balancing, and smart routing matter as much as raw speed. This is part of why NVIDIA’s networking revenue has grown so fast, as we covered in why NVIDIA is expanding beyond GPUs: the company sells the fabric, the network cards, and the software that keeps traffic flowing evenly.
InfiniBand versus Ethernet
The biggest strategic fight is over which technology wins. InfiniBand, historically NVIDIA’s strength, offers low latency and tight integration. Ethernet is open, widely supplied, and cheaper to scale. One market analysis says that thanks to RoCEv2 and the Ultra Ethernet Consortium, Ethernet has begun to outsell InfiniBand in 800G deployments in massive clusters, though that comes from a switch vendor with its own interest in Ethernet.
The standards effort is real. The Ultra Ethernet Consortium’s 1.0.2 specification describes AI and HPC deployments of roughly 80,000 to 256,000 Ethernet ports with target port speeds starting at 800G. NVIDIA competes on both sides, selling InfiniBand and its Spectrum-X Ethernet platform, whose revenue Converge Digest says grew 2.6 times year over year. Cloud providers pick their own paths: Microsoft’s Maia 200 uses standard Ethernet in a two-tier scale-up design for clusters of up to 6,144 accelerators, and Google built its own Virgo network for the TPU 8t. We covered those strategies in how cloud companies are becoming AI hardware companies.
The optics problem: speed costs power
Pushing faster links creates a physical problem. At 800G and especially 1.6T per port, even the few centimeters of circuit board between the switch chip and a pluggable optical module start to hurt signal quality. Traditional pluggable optics also draw a lot of power. One infrastructure analysis argues that power becomes the ceiling as links move to 1.6T, pushing designers toward co-packaged optics, which place the optical engines next to the switch silicon to cut watts per bit.
The two main switch vendors are racing here. Broadcom’s Tomahawk 6 “Davisson” co-packages 16 optical engines of 6.4 Tb/s each around the chip, and Broadcom claims roughly 70 percent lower optical power for the design. NVIDIA describes production deployments of Quantum-X Photonics InfiniBand switches with AI cloud operators in mid-2026 according to one transceiver guide, though vendor timelines tend to slip. Treat these power savings as company claims until independent measurements appear.
This links directly to the power constraint we covered in why AI data centers need much more power: every watt spent moving bits is a watt not spent computing.
The market is responding
Demand for optics is rising quickly. One analyst-sourced piece expects shipments of 400G and 800G datacom optical modules to climb from about $9 billion in 2024 toward $16 billion by 2026, though that is a forecast from a secondary source. Another says IDC data shows 800G switch revenue surging 220 percent quarter over quarter in the first half of 2025. The direction is clear even if exact figures vary: networking is growing with the clusters themselves, a theme we described in why AI companies are building massive GPU clusters.
Why inference needs it too
Networking isn’t only a training story. Inference spreads work across many GPUs, moves large KV caches between machines, and splits stages across separate pools, as we explained in why AI inference is becoming the next big infrastructure challenge. Disaggregated serving only works if the network can shuttle cached data quickly enough to avoid slowing responses.
The risks and trade-offs
Faster networks cost more, draw more power, and add complexity. Standards are still maturing, so buyers face choices between open Ethernet, proprietary InfiniBand, and custom fabrics, each with different performance and lock-in. Optical components add supply-chain exposure, in the same way chips and memory do, as we described in why semiconductor supply chains matter for AI. And because many speed and power figures come from vendors, independent testing will matter as 1.6T and co-packaged optics roll out.
The bigger picture
AI data centers need faster networking because training and serving large models turns thousands of separate chips into one machine, and that machine is only as quick as its slowest connection. The race to 800G, 1.6T, and co-packaged optics is a race to keep expensive accelerators busy within a limited power budget. The companies that control the network, from NVIDIA and Broadcom to the cloud providers’ own designs, increasingly control how efficiently AI can scale. Follow our AI hardware coverage as networking keeps pace with the chips.
SiliconeUpdate.com is a technology news platform that publishes updates and informational content related to silicon technology, software, artificial intelligence, and emerging technologies.
All articles published on this platform are attributed to SiliconeUpdate.com instead of individual authors. Content is presented in a neutral, informational format without personal opinions.
—
Content Publishing
SiliconeUpdate.com publishes news and updates based on publicly available information, official announcements, and industry developments. The focus is on clarity, relevance, and timely reporting.
—
Editorial Control
All editorial decisions, updates, and content management are handled at the platform level. No individual human or AI identity is presented as the author of articles.
—
Contact
For editorial communication or general queries, contact:
Email: neemasharma@gmail.com