S
Shrey Shah
Guest
The fusion reactors, the liquid-cooled servers and the 600-kilowatt server racks are all humming in race to keep the AI boom running.
Everyone here knows by now that AI's real choke point is power, not chips. But less acknowledged is the absurdly aggressive tech-stack now being constructed to overcome the power shortage on every vector-and the bleak, predictable pattern underlying it all: as hardware gets more efficient, the industry actually demands more power. A true Jevons paradox being unleashed on a gigawatt scale.
Nvidia's Vera Rubin platform, now flowing to customers, touts as much as 10x more inference performance per watt than Blackwell, with disaggregated inference (breaking inference into prefill and decode states, running them on separate hardware components; a concept Nvidia calls Dynamo) able to boost output by a full 35x tokens per watt over GPUs-only designs. These are real and significant efficiency wins.
Yet the power consumed by a state-of-the-art server rack is unprecedented:
!AI rack power consumption by generation
From approximately 40 kW for a Hopper rack in 2023, we're already at around 210 kW for a Vera Rubin NVL72 in 2024, and the projected 600kW "Rubin Ultra" (Kyber) rack due in 2027, with 1 MW+ units supposedly already in development, represents a 15x increase in power draw on hardware with a mere 10x leap in watt efficiency over just four years. Rather than decreasing power consumption, the increased efficiency allows engineers to pack more compute into each rack, and hyperscalers are eagerly deploying this capability to serve ever-larger models and workloads. These racks can draw the equivalent power of between 40 and 170 American homes.
The resulting imperative: cooling is no longer an optimization, but a critical infrastructure requirement. Fanless GPUs are now standard in Nvidia's newest boxes because the fans cannot dissipate the massive amounts of heat generated. Everything is moving to warm-water direct-to-chip liquid cooling, even for previously air-cooled 8-GPU servers.
Dell'Oro predicts the liquid cooling market for data centers will exceed $7 billion by 2029.
Any infrastructure engineers haven't yet factor ed coolant distribution units (CDUs) into their capacity planning already will soon.
We discussed Big Tech's interest in acquiring nuclear power plants before, focusing on Microsoft taking over Three Mile Island, Amazon's deal for Susquehanna, Google's acquisition of the Kairos SMR, and Meta's significant nuclear energy portfolio. However, the story has evolved beyond press releases. These initiatives are now translating into tangible construction projects, complete with regulatory applications and steel in the ground. The final delivery of power on schedule is debatable, especially considering that some SMR designs, like Kairos's molten-salt reactor, have no precedent for commercial operation anywhere globally, but the money and permits are no longer theoretical exercises.
This scenario would have sounded like science fiction just three years ago. There are now binding power-purchase agreements in place between multiple fusion companies and hyperscalers - not research grants, but real contracts with delivery terms:
* Helion Energy has executed the first-ever fusion power purchase agreement, committing to delivering a minimum of 50 megawatts to a Microsoft data center in central Washington State by 2028. In June 2026, Helion became the first fusion company to receive operating licenses from state regulatory bodies, a Radioactive Materials License and a Radioactive Air Emissions License from Washington. It has since raised $465 million in a Series G round at a valuation of $15.5 billion.
* Commonwealth Fusion Systems is erecting Arc, a 400-megawatt commercial fusion power plant, on land leased from Dominion Energy in Virginia, not far from where the world's highest concentration of AI data centers now exists. Google has a power purchase agreement in place for 200 MW from Arc and has invested in the company twice.
* Google has also been providing funding to TAE Technologies for more than a decade, with a goal of commercial grid-scale fusion by the early 2030s.
* Eni, the Italian multinational energy company, has signed an off-take agreement valued at over $1 billion with Commonwealth Fusion.
While no power from these plants will likely come online before 2028, and fusion technology has famously missed its own deployment targets for 70 years, many experts in the field are openly skeptical that fusion can actually become economically viable in time to support this current AI cycle. However, the fact that fusion companies are now entering into enforceable commercial contracts with penalties for delayed delivery, as opposed to simple research partnerships, represents a genuine shift for an industry stuck in "20 years away." The driving force for these large-scale investments by hyperscalers is the relentless demand for AI-powered workloads, not climate policy initiatives.
Overlaying the efficiency trends for chips with the rollout of nuclear and fusion power generates an unavoidable conclusion: nobody who's serious is placing their bet solely on efficiency gains. Every hyperscaler pursuing 10x faster chips is simultaneously securing gigawatt-scale energy contracts, aware that the efficiency gains will be devoured by increasing scale, not enjoyed as cost savings. This isn't an engineering failure; it's what happens when the output (compute) grows in value at a rate that outpaces its production cost reduction.
The incentive is to use ever more compute, not less.
This manifests as ever-larger models and an ever-greater demand for agents, not token efficiency improvements leading to reduced energy consumption.
For infrastructure developers, the planning assumption for the next five years is unlikely to be "power will get easier to manage." Instead, it's likely to be: "that rack I put in two years ago is now one-fifth the power draw of the ones I'm planning today, so factor those increases into electrical and cooling budgets, and don't assume the next generation of chips will free up resources; instead, prepare to absorb them."
Everyone here knows by now that AI's real choke point is power, not chips. But less acknowledged is the absurdly aggressive tech-stack now being constructed to overcome the power shortage on every vector-and the bleak, predictable pattern underlying it all: as hardware gets more efficient, the industry actually demands more power. A true Jevons paradox being unleashed on a gigawatt scale.
Here's a tour of what's happening under the hood:
The Chips: 10x more efficient, 15x more power-hungry
Nvidia's Vera Rubin platform, now flowing to customers, touts as much as 10x more inference performance per watt than Blackwell, with disaggregated inference (breaking inference into prefill and decode states, running them on separate hardware components; a concept Nvidia calls Dynamo) able to boost output by a full 35x tokens per watt over GPUs-only designs. These are real and significant efficiency wins.
Yet the power consumed by a state-of-the-art server rack is unprecedented:
!AI rack power consumption by generation
From approximately 40 kW for a Hopper rack in 2023, we're already at around 210 kW for a Vera Rubin NVL72 in 2024, and the projected 600kW "Rubin Ultra" (Kyber) rack due in 2027, with 1 MW+ units supposedly already in development, represents a 15x increase in power draw on hardware with a mere 10x leap in watt efficiency over just four years. Rather than decreasing power consumption, the increased efficiency allows engineers to pack more compute into each rack, and hyperscalers are eagerly deploying this capability to serve ever-larger models and workloads. These racks can draw the equivalent power of between 40 and 170 American homes.
The resulting imperative: cooling is no longer an optimization, but a critical infrastructure requirement. Fanless GPUs are now standard in Nvidia's newest boxes because the fans cannot dissipate the massive amounts of heat generated. Everything is moving to warm-water direct-to-chip liquid cooling, even for previously air-cooled 8-GPU servers.
Dell'Oro predicts the liquid cooling market for data centers will exceed $7 billion by 2029.
Any infrastructure engineers haven't yet factor ed coolant distribution units (CDUs) into their capacity planning already will soon.
The Nuclear Bet: moving from hypothetical to construction sites
We discussed Big Tech's interest in acquiring nuclear power plants before, focusing on Microsoft taking over Three Mile Island, Amazon's deal for Susquehanna, Google's acquisition of the Kairos SMR, and Meta's significant nuclear energy portfolio. However, the story has evolved beyond press releases. These initiatives are now translating into tangible construction projects, complete with regulatory applications and steel in the ground. The final delivery of power on schedule is debatable, especially considering that some SMR designs, like Kairos's molten-salt reactor, have no precedent for commercial operation anywhere globally, but the money and permits are no longer theoretical exercises.
The Fusion Bet: the biggest long-shot, which is closer than you think
This scenario would have sounded like science fiction just three years ago. There are now binding power-purchase agreements in place between multiple fusion companies and hyperscalers - not research grants, but real contracts with delivery terms:
* Helion Energy has executed the first-ever fusion power purchase agreement, committing to delivering a minimum of 50 megawatts to a Microsoft data center in central Washington State by 2028. In June 2026, Helion became the first fusion company to receive operating licenses from state regulatory bodies, a Radioactive Materials License and a Radioactive Air Emissions License from Washington. It has since raised $465 million in a Series G round at a valuation of $15.5 billion.
* Commonwealth Fusion Systems is erecting Arc, a 400-megawatt commercial fusion power plant, on land leased from Dominion Energy in Virginia, not far from where the world's highest concentration of AI data centers now exists. Google has a power purchase agreement in place for 200 MW from Arc and has invested in the company twice.
* Google has also been providing funding to TAE Technologies for more than a decade, with a goal of commercial grid-scale fusion by the early 2030s.
* Eni, the Italian multinational energy company, has signed an off-take agreement valued at over $1 billion with Commonwealth Fusion.
While no power from these plants will likely come online before 2028, and fusion technology has famously missed its own deployment targets for 70 years, many experts in the field are openly skeptical that fusion can actually become economically viable in time to support this current AI cycle. However, the fact that fusion companies are now entering into enforceable commercial contracts with penalties for delayed delivery, as opposed to simple research partnerships, represents a genuine shift for an industry stuck in "20 years away." The driving force for these large-scale investments by hyperscalers is the relentless demand for AI-powered workloads, not climate policy initiatives.
The Awkward Truth
Overlaying the efficiency trends for chips with the rollout of nuclear and fusion power generates an unavoidable conclusion: nobody who's serious is placing their bet solely on efficiency gains. Every hyperscaler pursuing 10x faster chips is simultaneously securing gigawatt-scale energy contracts, aware that the efficiency gains will be devoured by increasing scale, not enjoyed as cost savings. This isn't an engineering failure; it's what happens when the output (compute) grows in value at a rate that outpaces its production cost reduction.
The incentive is to use ever more compute, not less.
This manifests as ever-larger models and an ever-greater demand for agents, not token efficiency improvements leading to reduced energy consumption.
For infrastructure developers, the planning assumption for the next five years is unlikely to be "power will get easier to manage." Instead, it's likely to be: "that rack I put in two years ago is now one-fifth the power draw of the ones I'm planning today, so factor those increases into electrical and cooling budgets, and don't assume the next generation of chips will free up resources; instead, prepare to absorb them."