← Back to AI & Technology
Sidy's Intelligence Brief — AI & Technology

Co-Packaged Optics: AI Scaling Is Becoming a Data-Movement Problem

2026-09-2217 min read

As AI systems add more accelerators, the difficulty is no longer only performing more arithmetic. The system must also move far more data between processors, memory and switches. At very high link speeds, electrical distance becomes costly in loss, power and signal integrity. Co-packaged optics changes the architecture by moving optical engines much closer to the switching or compute silicon, shortening the hardest electrical path. That can improve bandwidth density and energy efficiency, but it also moves packaging, thermals, laser delivery, reliability, serviceability and interoperability into the critical path.

AI infrastructureCo-packaged opticsData movementSilicon photonicsAdvanced packaging

The Brief in One Sentence

Every scaling wave moves the bottleneck: when compute multiplies faster than affordable data movement, the network stops being plumbing around the computer and becomes part of the computer itself.

Why It Matters Now

AI infrastructure is scaling by connecting more accelerators into larger systems. That increases not only compute capacity but the quantity of data that must move across links during training and inference. Peer-reviewed 2026 work on co-packaged optics describes resistive loss, capacitive loading and frequency-dependent distortion in electrical interconnects as increasingly important constraints on bandwidth, latency and energy efficiency.

The consequence is architectural. A switch or accelerator can become more capable while the links around it consume more power and occupy more front-panel, package and cooling budget. The relevant question therefore shifts from how fast is the chip? to how much useful computation can the whole system sustain after the cost of communication is included?

Explain It Simply

Imagine a factory that keeps buying faster machines but leaves the same narrow corridors between them. Eventually the machines spend too much time waiting for parts to arrive. The factory has more machinery but not proportionally more useful output.

AI systems can hit a similar problem. GPUs and accelerators are powerful, but they constantly exchange data. Electrical signals become harder and more power-hungry to push over very fast links as distance and data rate rise.

Co-packaged optics shortens that difficult electrical trip. Instead of sending the fastest electrical signals all the way to a separate optical module at the edge of a board, the system converts them to light much closer to the switch or compute silicon, then carries the data optically.

Evidence Map

  • Observed / technical limit: a 2026 Nature Electronics review identifies resistive loss, capacitive loading and frequency-dependent distortion as growing constraints on electrical interconnect bandwidth, latency and energy efficiency.
  • Observed / power comparison: a 2026 Journal of Lightwave Technology analysis estimates current CPO transceiver energy-efficiency improvements of 23% versus LPO and 67% versus DSP pluggables in its analyzed generations; the authors explicitly warn that results depend on system boundary and technology generation.
  • Observed / standardization: the OCI MSA released a 200G optical-compute-interconnect specification in March 2026 for low-power, high-density AI backend interconnects, with implementation options that place optics at different distances from the ASIC.
  • Observed / industrialization signal: NVIDIA says its Spectrum-X Ethernet Photonics CPO platform entered production with Vera Rubin in 2026. Broadcom says it is shipping production CPO systems. These remain vendor claims, not an independent census of deployment.
  • Observed / remaining constraints: current literature highlights thermal management, heterogeneous integration, manufacturing, reliability/serviceability and standardization as first-order issues.
  • Inference: the architectural boundary between networking and computing is moving inward because data movement itself is becoming a limiting resource.
  • Uncertain: public evidence does not yet establish one universal CPO cost, reliability or power advantage across every scale-up and scale-out topology.

What CPO Actually Changes

The central architectural move is simple: shorten the highest-speed electrical path before converting the signal to light.

In conventional pluggable architectures, switch or compute silicon drives high-speed electrical signals across a package, board and connector path to a removable optical module. As link speed rises, that electrical path requires more equalization, retiming or DSP work to preserve signal integrity.

Co-packaged optics places optical engines beside, on, or within the same package environment as the ASIC. The optical fiber then carries the signal for the longer reach. The important change is not that electricity disappears. It is that the electrical-to-optical boundary moves closer to the silicon.

Why Moving Bits Can Dominate

A large AI system is useful only if its accelerators can exchange parameters, activations, gradients, tokens and synchronization traffic quickly enough. Adding compute therefore increases pressure on the communication fabric.

This creates a migration-of-bottleneck effect. A generation may begin constrained by compute. New accelerators reduce that constraint. The system then hits memory bandwidth, interconnect bandwidth, network power, topology or cooling. Progress in one layer exposes the next weakest layer.

CPO matters because it attacks one of those next layers: high-speed communication energy and density. It does not remove the other bottlenecks.

Sidy’s Synthesis — The Bottleneck Migration Rule

Compute → Data Movement → Conversion Boundary → New Constraints

When compute capacity rises, ask what must move in order for that compute to stay busy. If the answer requires more bandwidth than the existing electrical path can provide economically, the system has to move the conversion boundary.

CPO therefore illustrates a broader technology rule: solving one bottleneck does not remove constraint; it relocates constraint. Electrical reach gives way to photonic packaging. Pluggable serviceability gives way to tighter integration. The architecture improves one dimension and makes another dimension more important.

Power Numbers Need a Boundary

Power claims around optical interconnects are easy to misuse. One comparison may count only the optical module. Another may include the host SerDes, switch ASIC contribution or retiming. A newer pluggable generation can also narrow the difference.

The 2026 JLT analysis is useful precisely because it separates these boundaries. Its reported improvements should therefore be read as evidence that architecture matters, not as a universal percentage to copy into every data-center business case.

The correct unit is increasingly system energy per useful bit moved, not the headline wattage of one component.

Integration Creates New Failure Modes

Moving optics closer to hot, expensive silicon creates benefits and new coupling. Thermal gradients can affect photonic components. Fiber attach and optical alignment enter the package-manufacturing problem. A failed removable pluggable can be swapped; a deeply integrated optical engine changes the service model.

External laser architectures can separate some heat and replacement risk, but they add their own distribution, coupling and redundancy questions. Manufacturing yield also matters because tighter heterogeneous integration can make one defective subsystem more economically consequential.

That is why the useful question is not whether CPO is elegant. It is whether the complete package can be manufactured, cooled, tested, serviced and operated reliably at scale.

Standards Matter Because Packaging Can Lock the Ecosystem

Tighter integration can improve performance while increasing dependence on one supplier’s package, optical engine, connector, management interface or manufacturing flow. That is why the 2026 OCI effort matters: it attempts to define an open optical-compute interface and a multi-vendor ecosystem.

A released specification does not prove mature interoperability. Compliance, test tooling, production volume and multi-vendor field experience still have to follow. But standardization changes the strategic option set by making it possible, in principle, to separate architecture from a single proprietary implementation.

What Most People Miss

The most important shift is not from copper to fiber in the abstract. Data centers already use enormous amounts of optical fiber. The shift is where the optical link begins.

Moving that boundary by centimeters can change equalization needs, power, bandwidth density, package design, cooling, field replacement and supplier structure. A seemingly small physical relocation can therefore alter the architecture of the entire system.

Critical View

CPO should not be treated as inevitable everywhere. Advanced pluggable and linear-pluggable optics continue to improve and preserve important serviceability advantages. Different reaches, switch radices, topologies and cost structures can favor different solutions.

Vendor roadmaps toward million-GPU systems are ambitions, not independent proof of economical deployment at that scale. Production announcements also do not reveal installed share, field failure rates or total lifecycle cost.

The correct conclusion is narrower: as bandwidth density and energy per bit become harder constraints, bringing optics closer to silicon becomes a more credible architectural response. Whether CPO wins a particular system still depends on the complete economics and reliability envelope.

What to Monitor Next

  • Independent watts-per-bit and switch-level power measurements across CPO, LPO and DSP-pluggable generations.
  • Field failure rates, optical-engine replacement strategy and laser redundancy.
  • Thermal behavior and manufacturing yield in high-volume packages.
  • OCI and other interoperability/compliance progress beyond specification publication.
  • Actual deployed CPO volume rather than announced platform availability.
  • Whether CPO expands from switching into compute-package scale-up links.
  • Lifecycle cost, including service, downtime, spares and package replacement.

Remember This

Faster compute creates more communication. When moving data becomes too expensive electrically, architecture changes where light begins. CPO is important because it changes that boundary — and with it, the location of the next bottleneck.

Primary sources

Facts, figures and quotations should be traceable to the sources below. Sidy's synthesis is labeled as synthesis and does not replace sourced facts.

  1. https://www.nature.com/articles/s41928-026-01681-6
  2. https://opg.optica.org/jlt/abstract.cfm?uri=jlt-44-16-7158
  3. https://www.nature.com/articles/s44310-025-00105-1
  4. https://oci-msa.org/assets/files/200G-OCI-Optical-Phy-Specification-v1.0.pdf
  5. https://oci-msa.org/
  6. https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory
  7. https://www.broadcom.com/info/optics/cpo
  8. https://nvidianews.nvidia.com/news/nvidia-and-coherent-announce-strategic-partnership-to-develop-optics-technology-to-scale-next-generation-data-center-architecture