Ted Kang
← All writing

The Invisible Constraint of AI Infrastructure: Why Light Is Replacing Copper

AI data center optical network diagram

1. Introduction

AI infrastructure is usually discussed through GPUs, power, and capital expenditure. That framing misses a quieter bottleneck inside the cluster: the network fabric linking accelerators together.

As model training scales, copper interconnects run into physical limits on reach, signal integrity, and power efficiency. Optical links are moving from the edge of the data center toward the center of the AI stack.

2. Copper vs Optics

Copper has historically dominated short-reach connections because it is familiar, cheap, and simple to deploy. The problem is that bandwidth density and distance requirements are now rising faster than copper can comfortably absorb.

Optics cost more upfront, but they increasingly solve the right problem: moving more data over longer distances with lower loss and better scaling characteristics.

Copper versus optics comparison infographic
Attribute Copper Optics
Reach Strong at very short distances Scales better over longer runs
Bandwidth density Increasingly constrained Better fit for high-speed scaling
Power per bit Rises as speeds increase Can be more efficient at scale
Cluster implications Good for legacy and short-reach links Critical for larger AI fabrics

3. The AI Data Center Network

An AI data center is not just a room full of GPUs. It is a tightly coupled system of servers, switches, transceivers, and software orchestration that has to move enormous volumes of data with very low latency.

The network now determines whether expensive compute is fully utilized or left waiting on communication overhead.

AI data center optical networking architecture

4. The Optical Supply Chain

The optical stack spans lasers, DSPs, modulators, transceivers, fiber infrastructure, and switch integration. Each layer has different economics and different bottlenecks.

That matters for investors because value will not accrue evenly. Some layers benefit from volume growth, others from technical scarcity.

Optical supply chain infographic
Layer Role Why it matters
Laser and photonic components Create and shape optical signals Foundational performance layer
DSP and signal processing Convert and manage high-speed data transmission Critical for efficiency and reliability
Transceivers and modules Package optics into deployable hardware Main interface with switching gear
Fiber and interconnect infrastructure Carry signals through the cluster Enables scale-out architecture

5. Co-Packaged Optics

Co-packaged optics aims to bring optical engines closer to the switch silicon itself. The goal is straightforward: reduce power, improve signal integrity, and avoid the growing penalties of driving electrical traces at extreme speeds.

The concept is technically attractive, but the path to mass adoption depends on packaging complexity, thermal management, and ecosystem readiness.

Co-packaged optics architecture diagram

6. AI Cluster Communication

Large training clusters behave like communication machines as much as compute machines. Gradient exchange, parameter synchronization, and model sharding all create network-heavy workloads.

At smaller scale, inefficiency is tolerable. At frontier scale, communication becomes a first-order variable in both training time and total cost.

AI cluster communication infographic
Cluster metric Implication Networking requirement
More accelerators per training run Higher east-west traffic Denser interconnect fabric
Larger parameter counts More synchronization overhead Lower-latency communication
Longer cluster reach Copper becomes less practical Greater optical penetration

7. Networking Bandwidth Evolution

The progression from lower-speed networking to 800G and beyond is not just a spec sheet story. It changes rack design, power budgets, thermal envelopes, and system architecture.

The faster the bandwidth cycle turns, the more pressure it places on every component connected to the network roadmap.

Networking bandwidth evolution chart
Era Typical bandwidth step Architectural pressure
Legacy cloud scaling 25G to 100G Incremental rack efficiency
Modern AI buildout 200G to 400G Cluster-level throughput
Frontier AI era 800G and beyond Power, reach, and package redesign

8. Why This Matters

Optics is becoming one of the hidden denominators of AI infrastructure. If compute demand continues to rise, the network can no longer be treated as a commodity afterthought.

That makes optical infrastructure strategically important not only for operators, but also for investors trying to map where the next layer of value accrues in the stack.

9. Follow Along

I’ll continue writing about the less obvious constraints in AI infrastructure, especially where hardware, energy, and capital markets intersect.

If you follow this theme, the optical stack deserves to be tracked alongside chips, power, and data center development.