← ANALYSIS
AI · WAFER-SCALET1T2Nov 20, 2025· 8 min read

Scaling AI Hardware: The 15kW Wafer-Scale Challenge

How Cerebras's wafer-scale engine and the physics of I²R losses are driving the shift to 48V distribution in hyperscale AI infrastructure

BS
Beamed Silicon
Semiconductor Intelligence

The Cerebras CS-2, a wafer-scale machine learning accelerator, consumes approximately 15 kilowatts of power — the equivalent of 15 electric kettles running simultaneously, delivered to a chip the size of a dinner plate. No conventional computing architecture has ever operated at this power density in a single silicon package. Delivering thousands of amperes at sub-volt operating voltages across the chip exposes the fundamental physics of resistive power delivery: as current scales, losses scale with the square of the current, making traditional low-voltage distribution physically unworkable at wafer scale.

The Physics of I²R Loss at Extreme Power Levels

The resistive power loss in any conductor is given by P_loss = I² × R. This quadratic relationship is the central challenge of high-current power delivery. Doubling the current quadruples the loss. At the amperage levels required for a 15kW, 1V processor — approximately 15,000 amperes — even a conductor with resistance of one milliohm would dissipate 225 watts. The wiring from the power supply to the processor package introduces real and non-negligible resistance; at these current levels, the loss is catastrophic without intervention.

The traditional server power architecture — a centralised 12V power supply distributing to point-of-load converters near each processor — was designed for single-chip systems drawing hundreds of amperes, not thousands. A 12V bus supplying 15kW must carry 1,250 amperes at the source, requiring cable and busbar cross-sections that are physically impractical in a rack-mounted chassis. Even at the chip level, the interconnects between the package substrate and the silicon die represent a resistance that, at 15,000 amperes, creates voltage drops making uniform power delivery across a wafer-scale die nearly impossible.

The 48V Transition: Reducing Current in the Distribution Network

The solution adopted by hyperscale data centre operators is the shift to 48V distribution infrastructure. By distributing power at 48V rather than 12V, the current required to deliver the same wattage is reduced by a factor of four. A 15kW load at 48V draws only 312 amperes from the distribution bus, compared to 1,250 amperes at 12V. Since loss scales with I², the 48V bus dissipates 1/16th the resistive loss of an equivalent 12V system — a fundamental improvement that enables practical high-power AI cluster deployment.

The Open Compute Project formalised the 48V architecture for data centres with its Open Rack V3 standard, which specifies 48V bus bars running the full length of the server rack and point-of-load converters — typically high-efficiency GaN-based down-converters — mounted directly on the server board or on the processor package itself. This 'distributed conversion' architecture concentrates the highest-current conversion as close to the load as physically possible, minimising the length of the highest-current conductors and dramatically reducing distribution losses.

Thermal Management at Wafer Scale

Power delivery at 15kW is inseparable from the challenge of removing 15kW of heat from a single silicon package. Air cooling at this power density is not viable — the airflow rates required would generate unacceptable acoustics and pressure drops. Cerebras uses a direct liquid cooling system in which coolant flows over a cold plate in direct contact with the back of the wafer, designed to remove the full 15kW continuously while maintaining the die junction temperature within the specified operating range.

The thermal design challenge at wafer scale differs fundamentally from conventional chip cooling. A standard GPU die of 800mm² has a relatively uniform power distribution that can be managed with a single cold plate. A wafer-scale die spanning the full 300mm wafer — approximately 70,000mm² — has a power density distribution that varies dramatically across its surface, with high-utilisation regions generating local hot spots that the cooling system must resolve without overcooling the rest of the die. This requires sophisticated thermal modelling and coolant flow optimisation that did not exist in prior generations of computing hardware.

SOURCES & FURTHER READING

Published by Beamed Silicon Intelligence. Analysis reflects publicly available information as of publication date. Nothing herein constitutes investment advice.