AI Networking: Where Optics Earns Its Place

How copper, pluggable optics and co-packaged designs affect useful AI compute, and where technical gains become commercial value.

Revised Sep 16, 2026 · An editorial update to the original Mar 16, 2026 note.

Illuminated fiber strands connecting a server to a glass optical engine.
Explore in 3D

See the network behind the link.

Explore the rack-to-switch fabric and its connections, then inspect the server network interface. This is a conceptual topology; the optical comparisons and assumptions follow in the article.

The central argument

The economics of an AI cluster depend on useful work delivered. Optical networking earns its return when it removes a real communication constraint at an acceptable cost, power draw and level of operational complexity.

The accelerator gets the attention. The network helps determine whether a collection of accelerators can work effectively together. Buying more chips does not guarantee a proportionate increase in output if those chips spend more time waiting for one another.

That creates an opportunity for optics, but “light is replacing copper” is too broad a description. Copper, conventional optical modules and more integrated optical designs can coexist. The useful question is which connection needs to change, what problem that change solves, and who gets paid for solving it.

Start with useful compute

Large training jobs divide work among accelerators and exchange intermediate results or model updates. Some exchanges can overlap with computation; others sit on the critical path and hold up the next step. A communication improvement matters most when it shortens that critical path. Faster networking cannot fix every cause of low productivity: software, memory access, storage and the availability of customer work also matter.

It helps to distinguish two network roles. Scale-up connects accelerators within a tightly coupled compute domain. Scale-out connects those domains or servers into a larger cluster. NVIDIA’s rack-system documentation describes NVLink within its rack-scale compute domain and InfiniBand or Ethernet between racks. That is a concrete architecture, not a universal rule that every scale-up network must stop at one rack. NVIDIA DGX rack-system guide.

These roles are separate from the physical medium. Ethernet and InfiniBand describe networking technologies; copper and fiber describe ways of carrying signals. Co-packaged optics describes where optical conversion sits. Treating all three as competing choices confuses different layers of the system.

Hypothetical example: a job takes 100 seconds, including 30 seconds of communication that cannot overlap with computation. Halving that communication time reduces the job to 85 seconds: a 15% reduction in completion time, or about 18% more jobs per hour if everything else stays fixed. It does not double output. The same logic explains why a spectacular link benchmark can produce a modest application improvement.

I would measure time to finish a representative job, or requests served at the required response time and quality. Technical efficiency also differs from paid utilization. An operator can run a technically efficient cluster and still lack enough paying work to earn an attractive return.

BandwidthHow much data can move in a given time.
LatencyHow long an exchange takes before work can continue.
ReliabilityHow consistently the job can progress without disruption.
Explore the mechanism

The same hardware budget can produce different output

Modeled computation · 75 hoursWaiting · 25 hours
Illustrative allocation of 100 accelerator-hours. Waiting is assumed to displace computation one-for-one, with no overlap and no other bottleneck. Real systems can overlap communication and computation; this is not a benchmark or a definition of measured GPU utilization.

Where copper becomes difficult

Copper remains attractive for short connections because a passive cable can avoid the cost and power of optical conversion. As signaling rates rise or paths lengthen, attenuation and interference make the electrical signal harder to recover. Active electrical cables extend the useful range by adding electronics, which brings its own cost and power consumption. There is no single distance at which every copper design stops working.

An Arista and Broadcom deployment guide compares passive copper, active electrical cables and several optical options for particular AI network configurations. The ranges differ by product and link design. I read that as evidence for choosing the medium connection by connection, rather than applying one reach limit to an entire data center. Arista–Broadcom AI networking deployment guide.

Optics converts an electrical signal into light, carries it over fiber and converts it back at the other end. Its advantage is the ability to carry high data rates over useful distances with different loss and density tradeoffs. It does not eliminate electronics, laser power, connector maintenance or the need to route traffic efficiently.

Moving a switch from the top of a rack to a separate network rack can change the economics simply by lengthening the connection. Density, cable routing and repair access therefore belong in the discussion alongside bandwidth. A design optimized for one compact rack need not be the right design across a hall.

Why move optics closer to the chip?

A conventional pluggable optical module sits at the equipment faceplate. The signal travels electrically between the main switch chip and that module. The module typically uses digital signal processing to recover and condition high-speed signals before or after optical transmission.

Linear pluggable optics, or LPO, keeps the removable module but removes its digital signal processor. It relies more on the host chip’s signal-processing capabilities and the quality of the full electrical and optical path. The LPO industry group explicitly identifies capable host chips and well-designed transmission lines as requirements. That makes compatibility and full-link validation central to deployment. LPO MSA technical questions.

Co-packaged optics, or CPO, brings optical engines onto the same package assembly as the main chip, shortening the electrical path. NVIDIA’s 2025 technical discussion explains why this can reduce the power spent overcoming electrical loss. A switch CPO design does not, by itself, mean optical engines are integrated into every accelerator package. NVIDIA on co-packaged optics.

Closer integration changes manufacturing and maintenance. Optical engines, fiber attachment and thermal design must work together, and the failed component may be harder to access than a front-panel module. Some designs keep lasers externally replaceable. The OIF framework discusses soldered and socketed engines, external lasers and their different reliability and repair tradeoffs. CPO is a family of implementations, not one fixed service model. OIF co-packaging framework, sections 7.2–7.8.

Any efficiency claim needs a denominator. A 50% saving in a component that accounts for a hypothetical 4% of facility electricity would directly save about 2% of facility electricity if everything else were unchanged. Cooling effects or additional productive capacity require separate analysis. A module-level saving is not a facility-level saving.

A closer look

Shorten the electrical journey

Two conceptual switch boards: longer copper traces reach optical modules at the edge on the left, while shorter traces reach optical engines near the switch on the right.
Left · Edge opticsA longer electrical path before conversion to light.
Right · Optics near siliconA shorter electrical path before conversion.
Conceptual proximity comparison, not a package layout. Co-packaged optics integrates optical engines with the switch package; it does not remove every electrical connection. Drawn distances and component counts do not imply measured savings. View full size ↗

Follow the value through the supply chain

The ecosystem includes switch chips, optical engines, lasers, fiber, connectors, packaging and testing. More optical connections can expand demand across these categories without giving each supplier the same pricing power. Integration can even remove a component or transfer its function to another part of the system.

I would follow three pieces of evidence: a qualified position in a customer’s production design, repeatable manufacturing economics, and continued relevance in the next generation. A qualification win establishes technical acceptance; it does not automatically establish a large order, attractive pricing or a durable margin.

The commercial test is the return after development spending, capacity expansion, warranty obligations and price reductions. Customer concentration deserves particular attention when a supplier funds a specialized production line for one architecture.

The thesis weakens if copper, LPO or improved conventional modules meet customer needs at a better total cost, or if CPO production and service complexity outweigh its advantages. It strengthens when customers show that optical integration improves completed work per dollar and suppliers convert adoption into cash. That is the bridge from an engineering trend to a business thesis.

Networking earns its place by improving the return on the whole cluster.

Sources and review. Reviewed September 13, 2026. Primary sources are linked alongside the relevant claims. Company examples and forecasts are identified by date; technical descriptions do not establish future investment returns. Numerical examples are hypothetical.

Keep exploring

Follow the next idea.

Back to Research ↗Explore the charts ↗