
Data center networking is undergoing its most fundamental transformation in two decades. The shift is not incremental—it is architectural. AI workloads, particularly large language model (LLM) training and inference, generate traffic patterns that traditional data center networks were never designed to handle. The result is a wholesale redesign of how data moves between compute nodes, how switches are interconnected, and how networks are managed and scaled.
The numbers tell part of the story. Modern AI training involves thousands of GPUs exchanging data simultaneously, often across racks and even across facilities. East-west traffic—server-to-server communication within the data center—has risen from roughly 20% of total traffic to as much as 80% in AI clusters. A single training session for a trillion-parameter model can generate terabytes of inter-node communication as model parameters, gradients, and intermediate activations are shuttled between GPUs.
But the numbers tell only part of the story. The deeper change is qualitative: AI workloads are continuous, tightly synchronized, and massively parallel, in contrast to the bursty, independent request-response patterns of traditional cloud applications. This shift has pushed network design from a background concern to a primary determinant of AI system performance and cost.
1. The Shift from North-South to East-West Dominance
Traditional data center networking was built around a north-south model: user traffic entering and exiting the data center, with relatively modest internal communication requirements. The rise of microservices and distributed applications shifted the balance toward east-west traffic—server-to-server communication within the data center. AI has pushed this pattern into overdrive.
In AI training clusters, internal traffic is no longer just frequent—it is constant and highly sensitive to delay. East-west traffic has become the dominant workload. More than 76% of data center traffic now flows east to west, moving between GPUs, endpoints, APIs, datasets, and internal services.
The practical consequences are severe. Oversubscription ratios that were once comfortably designed at 3:1 or 5:1 are now collapsing toward 1:1 in AI clusters, where every GPU must communicate with every other GPU simultaneously. Bandwidth demands have risen sharply: 400Gbps Ethernet is becoming baseline, 800Gbps deployments are accelerating, and 1.6Tbps is already on the horizon.
2. Elephant Flows and Synchronized Traffic Bursts
AI workloads generate a distinctive traffic profile that differs fundamentally from traditional cloud applications. Instead of many small, independent flows, AI training produces "elephant flows"—massive, continuous data streams that saturate links and operate in lockstep synchronization.
These elephant flows are characterized by:
Low entropy: A small number of very large flows, rather than many small flows, dominates the traffic mix.
High synchronization: Traffic bursts occur at predictable intervals, typically at the end of each training iteration when GPUs exchange gradients and model updates.
Sustained high throughput: Observed transfer rates between storage systems and GPU clusters range from 100 Gbps to 1 Tbps.
Fewer, larger endpoints: Neocloud traffic consolidates around a relatively small set of persistent endpoints, with each flow carrying larger volumes of data.
The synchronized nature of these flows creates a particularly challenging environment for traditional load balancing. Conventional Equal-Cost Multi-Path (ECMP) routing, which distributes traffic based on flow-level hashing, struggles with a small number of large elephant flows. Hash collisions can cause multiple elephant flows to compete for the same link while parallel paths sit idle. The slowest flow in a collective communication operation dictates the efficiency of the entire training task.
3. The Rise of Rail-Optimized Topologies
Traditional data center networks use CLOS leaf-spine topologies, which provide universal any-to-any connectivity through a hierarchical architecture. In a CLOS network, each server connects to a leaf switch, which connects to multiple spine switches, creating a structured mesh where inter-leaf communication traverses a three-hop path: leaf-spine-leaf.
CLOS networks work well for general-purpose cloud traffic, but they are inefficient for AI workloads. LLM training has highly predictable communication patterns: GPUs of the same local rank (0–7) across different nodes communicate most frequently with each other. In a CLOS topology, this traffic must traverse the spine layer, consuming bandwidth and adding latency that could be avoided.
Rail-optimized topologies address this inefficiency by grouping same-rank GPUs into dedicated "rails" with single-hop connectivity. In a rail-optimized design, GPU 0 of every server connects to the same leaf switch, GPU 1 to the same leaf switch, and so on. This alignment with the communication pattern dramatically improves efficiency for collective operations like all-reduce.
| Topology | Connectivity Model | Typical Path Length | Best For | Trade-off |
|---|---|---|---|---|
| CLOS (Fat-Tree) | Universal any-to-any | 3 hops (leaf-spine-leaf) | Mixed workloads, general cloud | Higher cost, unnecessary hops for AI |
| Rail-Optimized | Same-rank grouped into rails | 1 hop within rail | Dense transformer training | Requires perfect cabling |
| Rail-Only | No spine layer | 1 hop within rail | MoE models with reduced inter-rail traffic | Limited cross-rail flexibility |
| Rail-Optimized + Spine | Rail grouping with limited spine | 1 hop within rail, 3 hops cross-rail | Large clusters needing cross-rail | Higher cost than rail-only |
Rail-optimized architectures achieve 95% scaling efficiency for dense transformer models, while CLOS networks offer better flexibility for mixed workloads. Rail-only designs, which eliminate the spine layer entirely, achieve the same training performance while reducing network cost by 38% to 77% and network power consumption by 37% to 75%.
The rise of Mixture-of-Experts (MoE) models has further altered the network equation. MoE architectures deliver increased token performance while reducing collective network communication traffic to levels comparable to much smaller models, making rail-only designs more viable for these workloads.
4. RoCE vs. InfiniBand: The Protocol Battle for AI
One of the most consequential decisions in AI data center design is the choice between InfiniBand and Ethernet with RDMA over Converged Ethernet (RoCE). Each has distinct technical characteristics, and the "correct" choice depends on workload characteristics, scale, and economic constraints.
InfiniBand emerged from supercomputing requirements where microseconds determine success or failure. Its architecture assumes lossless transmission through credit-based flow control, where senders only transmit when receivers guarantee buffer availability. This eliminates packet drops but requires tight coupling between endpoints. InfiniBand delivers consistent sub-microsecond latency but struggles with dynamic workloads that deviate from expected patterns.
Ethernet evolved from local area networks where simplicity and interoperability mattered more than absolute performance. Its architecture assumes lossy transmission with best-effort delivery, relying on higher-layer protocols for reliability. Modern data center Ethernet adds features like Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) to approach InfiniBand's lossless behavior, while RoCE enables RDMA capabilities over standard Ethernet infrastructure.
The performance gap between the two has narrowed significantly. RoCE can achieve near-linear scaling performance comparable to InfiniBand when properly configured. However, InfiniBand still delivers 15% better performance for training workloads, at 2.3 times the cost. For Meta, InfiniBand's performance advantage could not justify the higher total cost of ownership across its 600,000 GPU fleet, leading the company to standardize on Ethernet. OpenAI, by contrast, credited InfiniBand's superior congestion control for enabling GPT-4 training to complete 40% faster than initial Ethernet-based attempts.
The industry trend is toward a hybrid model: InfiniBand for training, Ethernet for inference. As Ethernet-based solutions like NVIDIA Spectrum-X close the performance gap, Ethernet is positioned to overtake InfiniBand across high-performance deployments, matching performance while maintaining cost and operational advantages.
5. Congestion Control and Load Balancing Innovations
AI workloads impose stringent demands on congestion control and load balancing. Traditional algorithms assume equal-cost paths, but newer topologies for HPC and AI create unequal-cost paths that existing flow-based schemes fail to handle. Random packet spraying is also ill-suited to the synchronized elephant flows of AI training.
Several innovations have emerged to address these challenges:
Latency-Aware Packet Spraying (LAPS): Manages both packet send rate and distribution based on real-time one-way path delay, achieving joint load balancing and congestion control regardless of topology.
Scalable Global Load Balancing (SGLB): A distributed, global congestion-aware load balancing system that uses global congestion information to distribute traffic across all available paths, accelerating All-to-All collective communication by up to 60%.
Ultra Ethernet Consortium (UEC) Programmable Congestion Management: Standardizing congestion signaling (CSIG) that allows packets to carry high-fidelity congestion information.
These innovations reflect a broader shift in network design philosophy. The assumption of equal-cost paths no longer holds for newer data center network topologies catering to HPC and AI workloads. The network must adapt dynamically to the actual communication patterns of AI workloads, rather than assuming generic traffic behavior.
6. The Optical Interconnect Revolution
Electrical interconnects are reaching their practical limits, accelerating a shift toward optical technologies. Co-packaged optics (CPO), which bring photonic interfaces closer to compute units, are reducing both latency and power consumption. The transition to 800G and 1.6T optics is well underway: while 800G modules remain the mainstream workhorse, 1.6T modules have already entered mass production.
Optical interconnect has moved from a "nice-to-have" to a "must-have" for AI data centers. The reasons are threefold:
Bandwidth density: AI clusters are designed from the outset to deploy sufficient 800G/1.6T optical modules, with upgrades from 800G to 1.6T requiring only higher-speed optical module swaps.
Power efficiency: Networking alone accounts for nearly 10% of total compute power consumption in AI data centers, and this fraction continues to rise. Optical circuit switching can reduce spine-layer power consumption by nearly 99% compared to electrical packet switching.
Reliability at scale: Large-scale AI clusters comprising hundreds of thousands of transceivers experience frequent component failures, where disruptions in systems of 500,000 XPUs can lead to losses exceeding $3 million per day.
Co-packaged optics addresses these challenges by integrating photonic components directly within processor packages, dramatically reducing electrical path lengths and enabling unprecedented bandwidth densities while lowering energy consumption per transmitted bit. Optical circuit switches (OCS) are also emerging as a promising approach, enabling scalable, reconfigurable, and energy-efficient infrastructure for AI clusters.
7. Telemetry and Observability in AI Networks
AI networks require a level of observability that traditional polling-based monitoring cannot provide. Continuous streaming telemetry detects transient issues such as packet loss, congestion, and microsecond-level latency spikes that traditional polling-based monitoring misses. In-band Network Telemetry (INT) has become a cornerstone of modern high-precision congestion control protocols, collecting temporal attributes of the network to enable real-time path planning and adaptive routing.
The stakes are high. In AI clusters costing millions of dollars and operating continuously, every second of network-induced GPU idle time is directly measurable in financial terms. A single straggler flow, silent packet loss event, or node failure can stall an entire training job across thousands of GPUs. Telemetry provides the visibility needed to detect and resolve these issues before they cascade into system-wide slowdowns.
8. The Ultra Ethernet Consortium and Standardization
The Ultra Ethernet Consortium (UEC) represents the industry's collective response to the networking demands of AI and HPC workloads. Formed by industry leaders including Microsoft, Oracle, and others, the UEC aims to build an Ethernet-based open, interoperable, high-performance full-communications stack architecture.
The UEC released its 1.0 specification in June 2025, culminating thousands of hours of work toward enhancing Ethernet for AI and HPC needs. The specification targets the scale-out network and includes several key innovations:
Programmable Congestion Management (PCM): Standardizing congestion signaling that allows packets to carry high-fidelity congestion information.
Link Layer Retry (LLR): Improving reliability over lossy Ethernet links.
Credit-Based Flow Control (CBFC): Approaching InfiniBand's lossless behavior while maintaining Ethernet interoperability.
In-Network Collectives (INC): Standardizing collective communication primitives for Ethernet networks.
The UEC also aims to redesign the transport layer to avoid classic connection-state bottlenecks by using a highly scalable connectionless transport with ephemeral packet-delivery contexts rather than heavy connection-oriented dependencies. Much of this innovation sits at the endpoints and NICs, with gains coming from simplified RDMA behavior, multi-pathing, higher utilization, and lower tail latency.
9. Power and Thermal Constraints
Power density in AI data centers has reached levels that push air cooling beyond its limits. Power densities of 50 to 150 kilowatts per rack are accelerating the adoption of direct-to-chip liquid cooling and immersion systems. Networking hardware is undergoing a similar transition, as electrical interconnects reach their practical limits.
Network subsystems, once viewed as minor factors in energy consumption, are now major contributors to data center power consumption. AI data centers are projected to consume 68 GW of power by 2027—nearly doubling the global data center energy usage from 2022 and approaching California's total power capacity of 86 GW in that year.
Optical technologies offer a path to reduce this burden. Optical circuit switching can reduce spine-layer power consumption by nearly 99% compared to electrical packet switching, with 8-year lifecycle costs reduced by 76%. Eliminating most energy-intensive Optical-Electrical-Optical (O-E-O) conversions offers up to 60% energy savings. These savings are not merely operational—they directly affect the economic viability of AI infrastructure at scale.
10. What This Means for Network Design
The transformation of data center networking for AI workloads has several practical implications for network architects and operators.
| Design Dimension | Traditional Cloud | AI-Optimized |
|---|---|---|
| Traffic Dominance | North-south | East-west (up to 80%) |
| Oversubscription | 3:1 to 5:1 | 1:1 (non-blocking) |
| Topology | CLOS leaf-spine | Rail-optimized, rail-only |
| Protocol | TCP/IP | RDMA (RoCE or InfiniBand) |
| Congestion Control | ECMP, DCQCN | LAPS, SGLB, global congestion-aware |
| Telemetry | Polling-based | Streaming INT, microsecond visibility |
| Optics | 100G–400G pluggables | 800G–1.6T, CPO, OCS |
| Cooling | Air | Liquid, immersion |
Several principles emerge from this transformation:
Bandwidth is no longer a background concern. Network capacity directly determines GPU utilization and training cost. A network that cannot sustain the required bandwidth leaves expensive GPUs idle.
Topology must match workload patterns. Rail-optimized designs outperform CLOS for AI workloads because they align with the actual communication patterns of distributed training.
Protocol choice has economic consequences. The InfiniBand-versus-Ethernet decision affects not just performance but capital and operational costs over the life of the cluster.
Observability is essential. Traditional monitoring is insufficient for AI networks. Streaming telemetry and in-band network telemetry provide the visibility needed to detect and resolve transient issues.
Optics are the future. As electrical interconnects reach their limits, optical technologies—CPO, OCS, silicon photonics—become essential for scaling bandwidth while managing power consumption.
11.Conclusion
AI workloads are not simply adding more traffic to data center networks—they are changing the fundamental requirements of network design. East-west traffic has become dominant, elephant flows have replaced bursty request-response patterns, and synchronized collective communication has made predictable, near-lossless performance a necessity rather than a luxury.
The network industry is responding with a wave of innovation: rail-optimized topologies that align with AI communication patterns, congestion control algorithms that handle elephant flows and unequal-cost paths, optical interconnects that break through electrical bandwidth limits, and standardization efforts like the Ultra Ethernet Consortium that aim to bring InfiniBand-class performance to the Ethernet ecosystem.
For network architects and operators, the message is clear: the assumptions that guided data center network design for the past two decades no longer hold. AI workloads require a fundamentally different approach—one that treats the network not as a background utility but as a primary determinant of system performance. The organizations that recognize this shift and adapt their infrastructure accordingly will be best positioned to capture the full value of AI compute at scale.
TEL:+86 132 6656 7067




















































>
>
>
>
>
>
>
>