400G Optical Modules in Modern Networks
Dec 17, 2025|
400G optical module represents both a triumph of engineering pragmatism and a source of constant operational headaches. At its core, it does something straightforward: push 400 billion bits per second through glass using light. The implementation sprawls across multiple form factors, modulation schemes, wavelength configurations, and vendor interpretations of what "compatible" actually means. PAM4 modulation brought the industry to this speed threshold by encoding two bits per symbol instead of one, effectively doubling throughput without doubling baud rate-but that decision carries consequences that ripple through every layer of the deployment stack, from the DSP silicon burning 12 watts inside the module to the FEC engines on the host platform scrambling to correct the elevated bit errors that PAM4 inherently produces.

QSFP-DD and OSFP emerged from the standards process like two siblings who couldn't agree on anything except that they both wanted 400G. The industry needed eight electrical lanes at 50Gbps each, and two different consortiums decided to solve that problem in two different ways.
QSFP-DD won the compatibility argument. It fits existing QSFP28 cages if you squint hard enough and don't mind the second row of pins. Backward compatibility matters when you have tens of thousands of deployed ports and a CFO who asks pointed questions about stranded assets.
OSFP won the thermal argument. The slightly larger housing and integrated heatsink mean you can actually dissipate the 15-20 watts these modules draw without cooking adjacent ports. I've seen line cards where the corner QSFP-DD ports consistently run 8℃ hotter than the middle ones because the airflow design assumed 100G power envelopes.
Neither really won. Most hyperscalers went QSFP-DD for inventory simplicity. Most telecom deployments went OSFP because their coherent modules needed the thermal headroom. Everyone else picked whatever their switch vendor shipped and moved on.
The QSFP112 variant deserves mention because it confuses everyone. Four lanes at 100G each-same 400G aggregate, fewer lanes, newer SerDes. It matters for NIC connectivity where you want server-to-TOR links without the DSP gearbox complexity. It matters less than vendors claim elsewhere.
PAM4 Changed Everything (And Broke A Few Things)
Here's what nobody adequately explains when they're selling you on 400G: PAM4 signaling trades noise immunity for bandwidth efficiency, and that tradeoff isn't free.
NRZ encoding used two signal levels. High or low. One or zero. Your receiver just needed to distinguish between those two states, and the eye diagram gave you comfortable margins. PAM4 uses four levels-00, 01, 10, 11-which means your receiver now has to distinguish between three threshold crossings with one-third the voltage separation. The theoretical 9.54dB SNR penalty isn't theoretical at all. It shows up in your pre-FEC BER counters every single day.
The DSP inside a 400G module does heroic work compensating for this. Feed-forward equalization, decision feedback equalization, clock and data recovery-all running at 53.125 GBaud per lane. When it works, it's invisible. When it doesn't work, you get bursts of correctable errors punctuated by occasional uncorrectable ones, and good luck figuring out whether the problem is your module, your fiber, your host, or the cosmic background radiation.

I spent two weeks last year chasing an intermittent error condition on a DR4 link that turned out to be a DSP firmware bug that only manifested when the ambient temperature exceeded 31℃. The vendor acknowledged the issue three months after we opened the case. The firmware update that fixed it also broke interoperability with one of our older switch platforms.
The FEC situation compounds this. KP4 FEC-RS(544,514) for the standards wonks-can correct up to 15 symbol errors per codeword, which sounds generous until you realize how often you need it. Running 400G without FEC isn't just inadvisable; it's impossible for most use cases. The coding gain buys you roughly 7dB of margin, which PAM4 promptly consumes.
The reach specifications tell only part of the story.
400G-SR8 uses 850nm VCSELs across eight parallel fibers, targeting 100 meters over OM4. It's cheap. It's multimode. It requires an MPO-16 connector with eight TX and eight RX fibers. Within a rack or between adjacent racks, this works fine. The moment someone asks about running it "just a little further," remind them that modal dispersion at 850nm doesn't negotiate.
400G-DR4 operates at 1310nm over four parallel single-mode fibers, rated for 500 meters. The MPO-12 connector uses the outer eight fibers and leaves four unused-a fact that confuses cable installers roughly once per deployment. DR4 has become the workhorse for leaf-spine connectivity in single-mode plants because 500 meters covers most datacenter geometries with room to spare.
400G-FR4 uses CWDM4 wavelengths (1271, 1291, 1311, 1331nm) multiplexed onto a single fiber pair via duplex LC. Two kilometers reach. This is where 400G starts feeling economical for campus interconnects, because you're not pulling eight-fiber MPO trunks between buildings.
400G-LR4 stretches the same CWDM4 approach to 10 kilometers with higher launch power and better receivers. The price jump from FR4 to LR4 still surprises procurement departments who haven't updated their mental model from 100G-LR4 pricing.
400G-ZR deserves its own section because it represents a fundamentally different technology dressed in the same form factor.
Everything I've described so far uses direct-detect optics. Light goes in, photodiode converts it, DSP cleans it up. Coherent optics encode information in both amplitude and phase across two polarizations simultaneously, then use a local oscillator and sophisticated digital signal processing to recover everything at the receiver. The result: 400Gbps over 120+ kilometers of unamplified fiber in a pluggable module.
The OIF 400ZR standard specifies 16QAM modulation at 60 GBaud with dual polarization. The concatenated FEC (soft-decision inner Hamming, hard-decision outer staircase) provides about 10.8dB of net coding gain. The whole thing draws 15-20 watts and generates heat that would make a QSFP-DD module weep.
I've seen ZR modules installed in switches that weren't designed for that thermal load. The switch chassis reported normal temperatures because its intake sensors measured cool air. The module reported 73℃ because it was sandwiched between two other ZR modules with inadequate airflow. The link worked-barely-with elevated FEC corrections that nobody noticed until the pre-FEC BER trended past the threshold and packets started dropping.
ZR+ and MZR variants push reach further at the cost of interoperability. Vendor-specific enhancements to launch power, receiver sensitivity, and FEC algorithms can extend links past 400km, but you're buying a solution rather than a commodity.

The Third-Party Question
I've had this conversation approximately six hundred times.
"Can we use third-party 400G optics?"
Technically yes. The MSA specifications exist precisely to enable multi-vendor interoperability. A compliant QSFP-DD from manufacturer X should behave identically to one from manufacturer Y. The IEEE standards define the optical and electrical parameters. CMIS (Common Management Interface Specification) standardizes how the host talks to the module.
Practically, it depends.
Cisco's authentication mechanisms have evolved from the blunt "error-disable the port" approach of older platforms to more sophisticated vendor verification that logs warnings but doesn't necessarily disable functionality. The service unsupported-transceiver command remains the escape hatch. Arista tends to be more permissive but will decline to support issues that might stem from third-party modules. Juniper's stance varies by platform and software version in ways that require consulting their compatibility matrices.
I run third-party optics in lab environments without hesitation. For production paths carrying revenue traffic at 2 AM when something fails? I want to be able to call TAC and have them actually help instead of immediately deflecting to "replace with supported transceivers."
The cost math changes this calculation for hyperscalers who buy modules by the tens of thousands and employ optics engineers who can characterize and qualify suppliers independently. It's different math for enterprises buying hundreds of modules through distribution channels with limited technical resources.
A 400G QSFP-DD module draws somewhere between 10 and 15 watts depending on variant and vendor. A 400G coherent ZR module draws 15-20 watts. An 800G QSFP-DD800 module-already deployed in AI clusters-draws 18-25 watts.
Put 64 of these in a 2RU switch and you have 640 watts just from optics before accounting for the switch ASIC, memory, fans, and power supplies. The thermal design problem has moved from "adequate" to "critical" in a single generation.
I watched a thermal imaging camera sweep a fully-loaded 400G spine switch during a qualification test. The hottest modules weren't the ones you'd expect. Corner positions, downwind of the ASIC exhaust, ran hotter than face-plate center modules that got fresh air. The standard DDM temperature readings showed a 17℃ spread across ports that were supposedly identical.
The module specifications promise operation from 0℃ to 70℃, but the performance curves don't look the same at 70℃ as they do at 25℃. Laser threshold current increases. Slope efficiency decreases. Wavelength drifts-and for CWDM4 and DWDM systems, wavelength drift means crosstalk with adjacent channels.
Air-cooled systems are approaching their limits. Liquid cooling for switches remains exotic but increasingly necessary for AI/ML clusters where GPUs and optics compete for the same thermal budget.

The IEEE standards define compliance points. They don't guarantee your specific link will work.
TDECQ (Transmitter and Dispersion Eye Closure Quaternary) is the PAM4 equivalent of OMA (Optical Modulation Amplitude) but more complicated. It attempts to characterize transmitter quality in a way that predicts receiver performance. The measurement requires reference receivers and mathematical transforms that vary between test equipment vendors in ways that cause endless standards committee debates.
Pre-FEC BER testing matters more than it ever did. The "fingerprint" of your bit errors-random versus bursty, uniformly distributed versus concentrated in specific PAM4 symbols-determines whether your FEC can actually correct them. True random errors play nicely with Reed-Solomon codes. Burst errors from clock recovery issues or DSP misbehavior can overwhelm the FEC even when the raw BER looks acceptable.
I've learned to demand pre-FEC statistics from every 400G link, not just post-FEC. A link showing 0.00 post-FEC BER while running pre-FEC BER at 2×10⁻⁴ looks great until you realize there's almost no margin left. Add a slightly dirty connector or an aging laser, and that link will tip over the FEC cliff without warning.
At 400G the contamination problem becomes acute. The modulated eye has less margin. Particles that would have been invisible at lower speeds now attenuate enough to matter.
Single-mode fiber cores are 9 micrometers across. An MTP/MPO-12 connector carries eight active fiber paths (four TX, four RX) plus four unused. Every mating cycle risks contamination. Every contaminated end-face risks insertion loss that eats into your link budget.
The cleaning discipline required is non-negotiable but rarely followed consistently. One-click cleaners, dry wipes with static concerns, wet cleaning with isopropyl alcohol that must be wiped dry immediately rather than allowed to evaporate-every method has adherents and critics. What everyone agrees on: inspect with a fiber scope before connecting, and if it's dirty, clean it and inspect again.
I watched a deployment team burn an entire afternoon troubleshooting an intermittent 400G-DR4 link. Multiple module swaps. Configuration reviews. Finally broke out the inspection scope and found construction debris on the bulkhead adapter that nobody had thought to check. Twenty seconds with a cleaning tool fixed what four hours of troubleshooting couldn't.

If you're deploying new datacenter fabric today, 400G is the baseline for spine layer and increasingly for leaf-spine uplinks. The cost per bit has dropped to where 4×100G breakout from a 400G module is often cheaper than individual 100G modules. DR4 for anything over 30 meters inside a building. FR4 for campus interconnects. LR4 or ZR if you're reaching between sites.
If you're an enterprise contemplating your first 400G deployment, the switching platforms have matured, the module supply chain has stabilized, and the pricing no longer requires executive sign-off on each purchase order. Start with a leaf-spine refresh, prove out your cabling infrastructure can handle the tighter contamination tolerance, and understand that your management tools need to start collecting FEC statistics before you actually need them.
If you're a hyperscaler reading this, you're already past 400G for GPU clusters and wondering how 1.6T will actually deploy. Good luck with the thermal problems; I'll read your papers in two years.
The modules themselves have become remarkably reliable. The problems live everywhere else: contaminated connectors, misconfigured FEC modes, thermal designs that assumed yesterday's power envelopes, and support organizations still learning how to troubleshoot PAM4 signal integrity issues. The unglamorous fundamentals-clean your connectors, measure your temperatures, understand your FEC budget-matter more than the spec sheet debates ever will.


