The Next Bottleneck in AI Infrastructure Is No Longer the GPU

Over the past decade, discussions around AI infrastructure have largely centered on scaling compute capacity. Organizations have raced to deploy increasingly powerful GPUs as AI models have grown in size, complexity, and computational requirements. The emergence of massive AI clusters powered by thousands of GPUs has only reinforced this focus.

However, a fundamental shift is underway.

As AI deployments scale from hundreds to thousands—and increasingly tens of thousands—of GPUs, the primary challenge is no longer computation. It is communication.

In modern AI infrastructure, the network is rapidly becoming just as important as the compute layer itself.

The Rise of East-West Traffic


For decades, data center networks were primarily designed around a relatively simple model: users accessed applications hosted inside a data center. Most traffic therefore flowed between end users and servers.

As cloud computing gained traction, applications became increasingly distributed, with web servers, application servers, databases, storage systems, and analytics platforms communicating continuously with one another. The growth of Artificial Intelligence has accelerated this trend dramatically.

Unlike traditional enterprise applications, AI workloads are inherently collaborative. Training a large language model or running large-scale inference often requires thousands of GPUs working together as a single distributed system. These GPUs constantly exchange model parameters, gradients, intermediate results, and training data. As AI clusters grow in size, the volume of data moving between compute nodes can become enormous.

This has fundamentally changed the nature of data center traffic. Instead of data moving primarily between users and applications, a significant portion now moves between servers, accelerators, and storage platforms inside the data center itself.

Network architects refer to this as East-West traffic.

In many large AI environments, overall application performance depends not only on available compute capacity, but also on how efficiently thousands of GPUs can communicate with one another. As AI infrastructure scales, the amount of data exchanged across the network grows dramatically, placing unprecedented demands on switching fabrics and interconnect technologies. This is why networking is emerging as a critical factor in determining overall AI infrastructure performance.

Why More GPUs Alone Are Not Enough


As AI systems grow in scale, the network must transport enormous volumes of data between thousands of GPUs while minimizing latency, congestion, and packet loss.

This requires:
  • Extremely high bandwidth
  • Low and predictable latency
  • Efficient load balancing
  • Rapid recovery from failures
  • Non-blocking or low-oversubscription architectures
As a result, application performance is increasingly determined by the efficiency of the network fabric that connects compute resources.

The Emergence of Clos and Leaf-Spine Architectures


To meet the demanding networking requirements of large-scale AI workloads, modern data centers have increasingly adopted Clos-based architectures, most commonly implemented as leaf-spine fabrics.

A typical leaf-spine fabric consists of two layers:

Spine Layer
  • High-capacity backbone switches that interconnect the fabric
Leaf Layer
  • Access switches that connect servers, GPUs, and storage systems
In this architecture, every leaf switch connects to every spine switch, creating multiple equal-cost paths between any two endpoints. As a result, traffic can be distributed efficiently across the network, helping to avoid bottlenecks and improve resilience.

Leaf-spine architectures offer several advantages that make them particularly well suited for AI environments:

Predictable Latency

Traffic typically follows a consistent path through the fabric: Leaf → Spine → Leaf


Figure 1: Simplified Leaf-Spine Architecture for AI Data Center Networking

This helps maintain predictable network performance regardless of where workloads are located within the data center.

Massive East-West Bandwidth

Multiple parallel paths enable the network to support the high volume of server-to-server and distributed GPU communication characteristic of large-scale AI workloads.

High Availability

If a switch or link fails, traffic can be automatically redirected through alternate paths, minimizing the impact on applications.

Scalable Growth

Additional compute resources can be introduced by adding more leaf switches and servers, allowing the fabric to grow alongside AI infrastructure requirements.

These characteristics have made Clos and leaf-spine architectures the preferred foundation for modern cloud and AI data centers.

The Hidden Challenge: Connecting AI Islands


While much of the discussion around AI networking focuses on what happens inside a data center, the importance of interconnecting geographically distributed AI infrastructure is growing rapidly.

Modern AI infrastructure is increasingly distributed across multiple facilities rather than being confined to a single location. Organizations are deploying primary and secondary AI data centers, geo-distributed training clusters, disaster recovery sites, and multi-region cloud environments to improve resilience, compliance, and scalability.

As a result, AI workloads are no longer limited to communication within a single data center fabric. Increasingly, large volumes of data must move between geographically separated data centers.

This marks an important shift in network design. What was once a challenge of connecting GPUs within a facility is now becoming a challenge of connecting entire AI clusters across cities, regions, and even countries.

Unlike communication within a data center, these connections may span from tens of kilometers to hundreds or even thousands of kilometers.

At this scale, high-capacity transport infrastructure becomes a critical part of the AI stack.

This is where Data Center Interconnect (DCI) networks play a pivotal role. Leveraging technologies such as coherent optics, DWDM, and OTN, DCI networks enable operators to transport massive amounts of traffic between data centers while maintaining performance, reliability, and scalability.


Figure 2: End-to-End Networking Architecture Connecting Multiple AI Data Centers

The Convergence of Compute, Networking and Optical Infrastructure


Historically, compute, switching, routing, and transport networks were often planned and operated as largely independent domains. A data center team focused on servers and switches, while WAN and transport teams focused on routing and optical infrastructure.

The AI era is rapidly blurring these boundaries.

A modern AI workload may begin on a GPU cluster in one data center, traverse a leaf-spine fabric, cross an inter-city transport network, and access compute or storage resources in another location. The performance of such workloads increasingly depends on the efficiency of the entire infrastructure stack rather than any single component.

As a result, networking can no longer be viewed as a collection of isolated layers. Every component contributes to overall performance, resiliency, and resource utilization. A bottleneck in any part of the infrastructure can impact application response times, training duration, service availability, and the effective utilization of expensive AI compute resources.

This shift is driving significant investment across the entire networking stack, including:
  • High-capacity Ethernet fabrics
  • AI-optimized data center architectures
  • 400G and 800G networking
  • Multi-terabit routing platforms
  • Coherent optical transport systems
  • Advanced data center interconnect solutions
AI infrastructure performance will increasingly depend on how effectively compute, networking, and transport technologies operate together as a unified system.

Looking Ahead


As AI clusters scale across data centers, regions, and even countries, networking infrastructure will play an increasingly decisive role in determining overall system performance. The ability to move data efficiently—within data centers and between them—will be critical to unlocking the full potential of AI.

This is driving a new era of innovation across the networking stack, spanning data center switching, routing, transport, and Data Center Interconnect (DCI) technologies.

At Tejas Networks, we see this convergence of AI, data center networking, and transport infrastructure as a defining trend for the industry. From high-performance data center switching fabrics to Data Center Interconnect (DCI) networks powered by IP/MPLS routing and optical transport technologies, the ability to build scalable and resilient connectivity infrastructure will be central to supporting next-generation AI workloads. Our growing portfolio across switching, routing, and optical transport positions us to address the evolving networking requirements of modern AI data centers and distributed AI infrastructure.

The future of AI is not just about processing data faster—it is about moving data smarter.

More Resources

5G Advanced – what it holds for the world?

6G – Trying to look through a crystal ball

Revolutionizing Fishermen’s Safety with Tejas Networks’ Satcom based Vessel Tracking Solution

Scroll to Top