AI transparency: articles and reports are produced with the help of artificial intelligence and checked through a process that does not constitute specialist validation. They may contain errors: verify relevant information with independent sources. Read the full disclaimer.
Merlintrader Special Report · AI Infrastructure · September 16, 2026

AI’s New Bottleneck Is Power: What Happens When the Constraint Moves From the GPU to the Megawatt?

How NVIDIA, Google and the AI Energy Management Alliance are turning data centers into grid assets — and what the shift from GPU scarcity to megawatt scarcity means for NVIDIA, CoreWeave, Nebius and the next phase of the AI infrastructure boom.

$NVDA $CRWV $NBIS $AMD AI Factories Power & Grid Long-form research
Research note: company-reported performance figures are identified as such and should not be treated as universal benchmarks. The Lambda and Emerald AI results below are reported by NVIDIA; they are not independently audited benchmarks. This report separates demonstrated results, company projections and Merlintrader analysis.

The AI infrastructure race is moving beyond the GPU

The scarce resource increasingly deciding how much AI capacity can actually be deployed is electricity. GPUs, networking and models still matter, but the new economic question is how efficiently operators can turn secured megawatts into useful, billable compute. NVIDIA’s DSX platform and the newly launched AI Energy Management Alliance are attempts to make power itself programmable — inside the data center and, potentially, at the grid interface.

+24%
Higher token throughput
Lambda validation using NVIDIA DSX MaxLPS under the same aggregate power budget.
+23%
Better performance per watt
Reported in the same Lambda HGX B200 validation.
4 → 3 MW
Automatic load reduction
Emerald AI / Silicon Valley Power demonstration while preserving priority workloads.
96 MW
Dedicated commercial DSX Flex deployment
Planned NVIDIA AI Factory Research Center deployment in Manassas, Virginia.

Merlintrader thesis

The infrastructure winner may not simply be the company that owns the most GPUs. It may be the company that can secure power fastest, activate it reliably, keep accelerators utilized and generate the most economically useful AI output from every energized megawatt.

What matters next

Watch active vs. contracted power, time-to-power, utilization, tokens per megawatt, financing cost, flexible-load agreements and whether grid regulators create real economic incentives for controllable AI data centers.

01

The AI Infrastructure Story Has Changed

For most of the generative-AI boom, investors could understand the infrastructure story through a relatively simple hierarchy. At the top were the frontier model developers. Beneath them were the hyperscalers and specialized AI clouds. Underneath those platforms sat accelerators, high-bandwidth memory, networking equipment and data-center hardware. The dominant investment question was usually some variation of the same theme: who can secure the most advanced GPUs, and how fast can those GPUs be deployed?

That framework is no longer enough.

By September 2026, the bottleneck is moving outside the chip. In many regions, the problem is not whether an AI cloud company can order another rack of accelerators. The problem is whether it can find a site with sufficient electrical capacity, secure an interconnection agreement, procure transformers and switchgear, construct or contract suitable generation, remove the heat generated by extremely dense racks, and bring all of that infrastructure online quickly enough for the GPUs to begin generating revenue before the next hardware generation arrives.

This change sounds subtle, but it alters the economics of the entire AI trade. A GPU that cannot be powered is inventory. A data-center building without energized capacity is real estate. A signed customer contract that cannot be served because power delivery is late is a backlog item rather than recognized revenue. In other words, electricity increasingly determines the conversion rate between capital expenditure and revenue-producing compute.

The International Energy Agency has estimated that global electricity consumption from data centers could roughly double by 2030, with AI-focused data centers growing considerably faster than the overall category. In the United States, Lawrence Berkeley National Laboratory’s 2025 update estimates that data centers could account for roughly 11.8% of national electricity consumption by 2030 in its central estimate, with a wide scenario range around that figure. Electric Power Research Institute scenarios likewise envision several hundred terawatt-hours of U.S. data-center consumption by the end of the decade.

These forecasts should not be interpreted as precise endpoints. AI efficiency can improve quickly. Demand can accelerate or disappoint. New power supply can arrive. Data centers can move geographically. Models can become less computationally expensive per task even while aggregate usage explodes. But the direction is clear enough: AI is becoming one of the most important new sources of electricity-load growth in advanced economies, and the physical infrastructure required to support that growth is not expanding at software speed.

The central AI infrastructure question is no longer simply “How many GPUs can you buy?” It is increasingly “How many megawatts can you secure, connect, cool, finance and convert into useful tokens?”
02

Why Electricity Is Becoming the Real AI Currency

There is an important difference between electricity consumption and electrical capacity. An AI data center does not merely need enough energy over the course of a year. It needs enough instantaneous power capacity to operate extremely dense clusters reliably, plus enough redundancy and cooling infrastructure to handle variations in load.

That distinction matters because AI workloads can be unusually dynamic. Large training jobs, reinforcement-learning workloads and inference systems do not necessarily draw power in the same pattern. A facility designed around static assumptions can leave substantial electrical capacity stranded because operators must reserve headroom for worst-case peaks. If every node is assumed to hit peak consumption simultaneously, operators may intentionally underfill a site even though typical aggregate consumption is lower.

This is the problem NVIDIA is now trying to attack with software and system architecture.

Rather than treating the power budget as an immutable ceiling allocated statically across racks, NVIDIA’s DSX approach treats power as something that can be measured and dynamically assigned. Different workloads receive different power allocations. Lower-priority work can temporarily yield. Training and inference clusters can be managed differently. The objective is to keep more accelerators productively occupied without violating the site’s maximum electrical envelope.

Economically, that is a powerful concept. A company that can get 20% more productive compute from the same energized facility may not need 20% more land, 20% more transmission capacity or 20% more interconnection capacity to produce the same amount of AI output. If the incremental software and control-system cost is low relative to the avoided infrastructure spend, the improvement flows directly into capital efficiency.

Key concept: this is why NVIDIA increasingly talks about tokens per megawatt. Peak accelerator performance remains important, but it is no longer sufficient as the primary infrastructure metric. What ultimately matters to a cloud operator is how much billable or strategically useful work can be produced by the scarce resources already secured.
03

NVIDIA DSX: Moving the Optimization Boundary From the GPU to the Factory

NVIDIA’s DSX platform is significant because it expands the company’s optimization boundary. NVIDIA historically optimized the GPU. It then expanded into scale-up interconnect through NVLink, scale-out networking through InfiniBand and Spectrum-X Ethernet, CPU-GPU systems, DPUs, inference software and full rack-scale architectures. DSX moves one level higher again: the object being optimized is increasingly the entire AI factory.

The DSX family includes several layers. DSX MaxLPS focuses on extracting more productive compute from a fixed site power budget. DSX Flex focuses on adjusting facility consumption in response to external grid conditions. DSX Sim models infrastructure before deployment. DSX OS addresses lifecycle operations and resiliency. Reference designs extend the architecture into power distribution, cooling, networking, storage and facility design.

The strategic pattern should be familiar to anyone who has followed NVIDIA over the past decade. The company repeatedly identifies a bottleneck surrounding the GPU and then turns that bottleneck into part of its own platform.

CUDAProgramming environment becomes part of the moat.
NVLinkGPU-to-GPU communication becomes part of the platform.
Spectrum-X / InfiniBandNetworking becomes strategic infrastructure.
Dynamo / BlueFieldInference and infrastructure processing become integrated.
DSXPower and facility operations move inside the optimization stack.

That does not mean NVIDIA owns the grid. It does not manufacture every transformer, turbine, battery or cooling tower. But it does mean the company wants the operating logic of the data center to understand the behavior of NVIDIA compute deeply enough to optimize the full system around it.

Lambda provides the first useful real-world number

The most interesting evidence presented during AI Infra Summit week came from Lambda’s validation of DSX MaxLPS on an HGX B200 cluster.

According to NVIDIA and Lambda, the test used a five-rack, 19-node cluster and workloads designed to produce reproducible high utilization. The baseline was 16 nodes operating at full power. Using dynamic power management, Lambda operated 19 nodes inside the same aggregate facility power budget. Cluster-wide token throughput increased by approximately 24%, from roughly four million tokens per second to roughly five million, while performance per watt improved by about 23%.

What was demonstrated: 19 nodes operated inside the same aggregate power envelope previously used by 16. The reported outcome was approximately +24% total token throughput and +23% performance per watt.

These numbers should be interpreted carefully. This is not proof that every data center can simply install software and obtain 24% more revenue. Workload mix matters. Hardware configuration matters. Customer-service constraints matter. Power-management policies matter. Some environments will have more stranded headroom than others. An inference-heavy facility with strict latency commitments may have less flexibility than a facility running a large pool of interruptible research jobs.

But the result is still important because it validates the basic thesis: static power allocation can leave economically valuable capacity unused, and software aware of real workload behavior can recover some of that capacity.

For an AI cloud operator, the impact can be significant. Assume two otherwise identical data centers each have 100 MW of usable IT power. If one architecture can consistently convert a materially larger percentage of that envelope into productive GPU work, the more efficient facility effectively owns more compute capacity without having secured another power connection. When power interconnections can take years, that difference is strategically meaningful.

04

DSX Flex: Can the Data Center Become Part of the Grid?

MaxLPS optimizes power inside the fence. DSX Flex addresses the second problem: the relationship between the data center and the grid outside the fence.

This is where the story becomes more consequential.

Traditional large electricity users are often modeled as relatively inflexible load. A factory, office building or conventional data center requests an interconnection and the utility plans around the assumption that the load will need to be served whenever it appears. If the local transmission and generation system cannot support that demand under required reliability conditions, new infrastructure may have to be built before the customer can receive its full connection.

AI data centers are unusual because portions of their workload may be flexible in time. Some inference must run immediately. Some customer workloads have strict service-level agreements. But other work — model training, batch jobs, checkpoints, synthetic-data generation, certain research jobs and internal development workloads — can sometimes be slowed or shifted without destroying economic value.

That creates the possibility of treating the AI factory as a controllable grid resource rather than a passive load.

NVIDIA highlighted a commercial demonstration involving Emerald AI and Silicon Valley Power. This used Emerald AI Conductor at NVIDIA’s Eos facility; it was not a DSX Flex installation. The planned 96-MW Manassas deployment is a separate project. Emerald AI’s Conductor software received utility signals and automatically reduced the data center’s draw by changing the workload hierarchy. In the initial test, power reportedly fell from approximately 4 MW to 3 MW while high-priority work continued. NVIDIA says Silicon Valley Power subsequently sent more than 200 demand signals and that the system responded successfully.

The important part is not the single-megawatt reduction. The important part is the control model.

If utilities can trust a 100-MW, 500-MW or eventually gigawatt-scale AI campus to shed a predictable fraction of load within a defined response time, interconnection planners may be able to treat that facility differently from a fully inflexible customer. Flexible interconnection agreements could allow projects to connect earlier or at greater initial capacity while transmission upgrades are still being completed.

This does not create electricity from nothing. During prolonged shortages, the underlying system still needs enough generation. Transmission lines still need to be built. Transformers still need to be manufactured. But flexibility can improve utilization of infrastructure that would otherwise be sized around relatively rare peak conditions.

In power systems, that distinction is enormous. The grid is engineered for reliability under stressed conditions, not merely for average demand. If AI workloads can become dispatchable enough to reduce those peak requirements, some existing infrastructure can support more nominal data-center capacity.

05

The AI Energy Management Alliance: From Technology Demo to Industry Standard

On September 16, Emerald AI, Google and NVIDIA launched the AI Energy Management Alliance, or AEMA. The founding group is joined by a broader set of launch partners spanning AI, semiconductors, utilities, generation, storage and grid software. Named participants include Anthropic, Analog Devices, AES, National Grid, Constellation, NRG, RWE, Fluence and others.

The significance of AEMA is not that a new industry association has been created. Technology industries create associations constantly. The significance is the specific problem the organization is trying to standardize: how grid operators should evaluate a data center that promises to behave as a flexible load.

A utility cannot rely on marketing language. “Our AI data center is flexible” is meaningless unless the operator can answer several technical questions: how many megawatts can the facility curtail, how quickly can it respond, how long can the reduction last, how often can it be called, whether it can remain online through a voltage or frequency disturbance, what telemetry is available, whether the response is automated, what workloads are protected and what happens if the facility fails to deliver the promised reduction.

AEMA says it intends to focus on measurable performance rather than prescribe one technology. That is the right conceptual approach. One operator might create flexibility through workload scheduling. Another could combine compute orchestration with batteries. Another might use onsite fuel cells or gas generation. Others may use hybrid architectures combining grid supply, generation and storage.

The grid should care less about the brand of technology and more about the electrical behavior of the load.

If AEMA contributes to standardized interconnection frameworks, the economic impact could be larger than the initial technology announcements. The value would come from making flexibility financeable and contractible. Developers, utilities and lenders could model a known set of obligations rather than negotiate every large AI project from scratch.

06

Why the Grid Cannot Simply “Build More Power” Fast Enough

The obvious response to a shortage of data-center power is to build more generation and transmission. That is indeed part of the solution. But the timelines do not align cleanly.

An AI data center can be planned and built comparatively quickly. GPUs can move from product announcement to large-scale deployment within a few years. Cloud demand can change in months. Electricity infrastructure operates on a very different clock.

Large transmission projects may require many years of planning, permitting, land acquisition and construction. New substations require specialized equipment. High-voltage transformers have long manufacturing lead times. Gas turbines have experienced tightening supply. Nuclear projects are measured in many years. Even renewable projects that can be built quickly still require grid connections, storage or complementary dispatchable generation to support reliable round-the-clock demand.

In June 2026, the U.S. Federal Energy Regulatory Commission launched proceedings aimed specifically at the rules governing connection of very large new loads such as data centers. The fact that federal regulators are focusing explicitly on “speed to power” illustrates how infrastructure constraints have moved from a niche utility issue into the center of U.S. industrial policy.

The IEA estimates that roughly one-fifth of planned data-center projects could face delay risks if grid constraints are not addressed. That number should not be treated as a forecast of cancellations, but it captures the underlying mismatch: the AI project pipeline is growing faster than conventional power infrastructure can be delivered.

Why flexibility matters: a flexible data center may be able to operate before a transmission upgrade is complete. Onsite fuel cells can bridge grid delays. Batteries can handle fast transients. Software can shift non-urgent compute. New generation can be added over time.

The emerging AI energy architecture will therefore probably be hybrid rather than singular.

07

Vera Rubin Makes the Power Problem More Urgent — Not Less

Chip efficiency improves every generation, but that does not necessarily reduce total electricity demand. The reason is straightforward: better economics stimulate more usage.

If a new accelerator cuts the cost per token in half, customers do not necessarily spend half as much. They may train larger models, run more inference, deploy more agents, process more video and create entirely new workloads that were previously uneconomic.

This is the infrastructure version of the Jevons paradox: greater efficiency can increase total consumption by making the underlying service cheaper and more useful.

NVIDIA’s Vera Rubin architecture is a good example. Rubin systems are designed to improve performance and efficiency dramatically relative to previous generations. But the rack-level systems are also extraordinarily dense. The industry is moving from individual accelerators toward integrated racks and multi-rack fabrics that behave increasingly like single giant computers.

That makes power delivery, cooling and facility design inseparable from compute architecture.

Important distinction: NVIDIA says DSX MaxLPS combined with future Vera Rubin deployment planning could, in suitable environments, permit up to roughly 40% more GPU capacity within the same megawatt budget. This is a company projection rather than a broadly independently demonstrated result and should be treated accordingly.

NVIDIA is also moving toward 800-volt DC power distribution in its reference architecture. Higher-voltage DC distribution can reduce the conversion stages and conductor requirements involved in moving enormous amounts of power inside very dense facilities.

A few percentage points may sound small. At gigawatt scale they are not. A 3% improvement on a 1-GW facility represents 30 MW. Efficiency improvements that appear incremental at rack scale become infrastructure assets at campus scale.

08

CoreWeave: Why Power Is Already a Financial Metric

CoreWeave provides perhaps the clearest public-market example of why investors need to treat power capacity as a financial variable rather than an engineering footnote.

At the end of the second quarter of 2026, CoreWeave reported approximately 1.5 GW of active power and approximately 3.7 GW of contracted power. It also reported roughly $104 billion of revenue backlog at June 30, before more than $25 billion of additional customer commitments added early in the third quarter.

The backlog represents demand. Power represents the physical ability to serve that demand. Between those two numbers sits execution.

Every additional megawatt that moves from “contracted” to “active” can support installed equipment and revenue-producing workloads. Every delay in that transition postpones the conversion of customer commitments into service revenue. This is why power construction timelines, electrical equipment, financing and site development deserve almost as much attention as GPU supply.

CoreWeave has also been aggressive in adopting new NVIDIA architectures. The company reported the industry’s first bring-up and validation of Vera Rubin NVL72 and on September 16 announced a multi-rack Rubin deployment connecting hundreds of Rubin GPUs in a scale-out cluster. That technological leadership is valuable, but Rubin availability alone is not enough. The equipment needs energized facilities.

CoreWeave’s current infrastructure footprint makes the DSX thesis especially relevant. At multi-gigawatt scale, recovering even a modest percentage of stranded power headroom can translate into a large amount of incremental usable compute capacity.

Consider a purely illustrative example. If an operator has 1.5 GW of active capacity and system-wide optimization could make just 5% more of that envelope economically productive, that is equivalent to 75 MW of incremental usable capacity. At 10%, the figure becomes 150 MW. These are not CoreWeave guidance figures and should not be interpreted as expected results. They simply show why software-level power optimization can become financially material at cloud scale.

CoreWeave’s advantage — and risk

CoreWeave’s growth model effectively monetizes speed. Customers pay for access to scarce advanced compute. The faster CoreWeave can secure power, install new generations and convert them into usable clusters, the faster it can recognize revenue against demand.

But this model also requires enormous capital. Power infrastructure, buildings, GPUs, networking and financing all arrive before much of the associated revenue. In Q2 2026, CoreWeave reported $2.575 billion in revenue and $1.510 billion in adjusted EBITDA, but also $640 million of net interest expense and a GAAP net loss of $626 million.

The trade-off: extraordinary demand can coexist with extraordinary capital intensity. More productive tokens per megawatt can improve the equation, but financing and execution remain fundamental equity risks.
09

Nebius: Power Strategy Becomes Part of the Product

Nebius provides another useful case because the company has been explicit that compute-capacity growth has historically been limited by how quickly infrastructure could be brought online.

The company is scaling toward multi-gigawatt capacity and has pursued several approaches to the power problem. One of the most notable is its partnership with Bloom Energy.

In May 2026, Nebius and Bloom announced plans for approximately 328 MW of fuel-cell capacity at a U.S. deployment. The key rationale was time to power. Behind-the-meter generation can reduce reliance on transmission upgrades and allow compute capacity to become operational more quickly than waiting exclusively for conventional grid expansion.

This is strategically important because it illustrates how AI clouds are beginning to design their own energy stacks.

The conventional data-center model starts with available grid capacity and builds a facility around it. The emerging AI-factory model may increasingly begin with customer demand and then assemble the fastest viable portfolio of grid electricity, fuel cells, storage, onsite generation and demand flexibility required to serve it.

Nebius has also emphasized efficiency. Its 2025 sustainability reporting stated an average portfolio PUE of approximately 1.25. PUE measures total facility power relative to IT equipment power. Lower numbers indicate less electricity is being consumed by cooling and other overhead.

PUE is useful, but it is no longer enough.

A data center can have excellent PUE and still underutilize its GPU power envelope. This is the conceptual innovation behind the “tokens per megawatt” framework: facility efficiency must eventually be combined with compute productivity.

10

NVIDIA, CoreWeave and Nebius Are Not the Same Trade

CompanyPrimary exposureWhy power mattersPotential advantageMain risk
$NVDAAccelerators, networking, systems and AI-factory softwarePower shortages can constrain the amount of NVIDIA hardware customers can deploy.DSX can extend NVIDIA’s platform into power optimization and factory control.Efficiency gains could reduce hardware needed per unit of AI output if demand does not expand fast enough.
$CRWVAI-native cloud infrastructureRevenue growth depends on bringing enormous amounts of contracted capacity online.Large active-power footprint makes efficiency gains economically meaningful.Very high capital requirements, financing costs and execution complexity.
$NBISAI cloud and infrastructureCapacity availability has been a practical limit on growth.Hybrid energy strategy, owned infrastructure and efficiency engineering.Rapid expansion requires capital and flawless infrastructure execution.
HyperscalersIntegrated cloud + AI platformsGigawatt-scale demand competes with other grid customers and infrastructure timelines.Balance-sheet strength and ability to finance generation directly.Regulatory, community and return-on-capital pressure.
Utilities / PowerGeneration, transmission and distributionAI creates enormous new load but can destabilize planning if poorly integrated.Long-duration demand growth and new infrastructure investment.Ratepayer protection, stranded-asset risk and uncertain data-center forecasts.
11

The Hidden Winner: Utilization

AI infrastructure economics are often described through hardware acquisition prices, but utilization can be just as important.

A $40,000 accelerator running productive customer workloads 90% of the time is economically different from an identical accelerator sitting idle because power cannot be delivered, a network job failed, storage is bottlenecked or a cluster must reserve excessive headroom for peak power events.

NVIDIA’s full-stack strategy increasingly aims to eliminate those idle periods. Dynamo addresses inference orchestration. Spectrum-X addresses networking. BlueField addresses infrastructure processing. Storage architectures reduce data bottlenecks. DSX addresses electrical and facility constraints.

AI cloud providers are not primarily in the business of owning GPUs. They are in the business of selling useful compute generated by GPUs.

The distinction is similar to the difference between an airline owning aircraft and an airline maximizing revenue-producing flight hours. Asset count matters, but utilization determines economics.

For AI infrastructure, the ultimate metric may therefore evolve from installed GPU count to productive token throughput per dollar of invested capital.

12

Flexible Load Could Change the Economics of Interconnection

The most important long-term consequence of AEMA may have little to do with NVIDIA software itself.

If regulators and utilities formally recognize flexible AI loads as lower-risk interconnection customers, the data-center development process could change.

Imagine two proposed 500-MW facilities. Facility A requires 500 MW at all times and cannot curtail. Facility B guarantees that it can automatically reduce consumption by 100 MW within one minute during predefined grid conditions, maintain that reduction for several hours and provide real-time telemetry proving compliance.

Those facilities impose different risks on the grid.

If interconnection rules treat them identically, Facility B receives no economic reward for flexibility. If tariffs recognize the difference, Facility B might connect earlier, require fewer immediate network upgrades or receive a different cost allocation.

That is the market design problem AEMA is trying to influence.

It is also why utility participation matters. The technology companies cannot simply declare themselves flexible and demand faster connections. Grid operators must trust the response under extreme conditions. The standards therefore need to address response speed, duration, telemetry, ride-through capability and penalties for non-performance.

If those rules become standardized, flexible data centers could eventually participate in electricity markets similarly to other demand-response assets.

Long-term implication: compute scheduling could become an energy-market activity.
13

AI Workloads Are Better Suited to Flexibility Than Many Industrial Loads

A steel mill cannot pause a blast furnace because electricity prices rise for fifteen minutes. A chemical process may not be safely interruptible. Many manufacturing facilities require predictable continuous operation.

Some AI workloads are different.

A large training job can checkpoint. A batch inference workload can be delayed. Synthetic-data generation may be shifted. A research experiment can wait thirty minutes. Low-priority internal workloads can yield power while customer-facing inference stays online.

This does not mean all compute is flexible. Inference serving a real-time consumer product may be extremely latency sensitive. Financial applications, healthcare systems and defense workloads can carry strict availability requirements. But the presence of both urgent and non-urgent work within the same campus creates a portfolio that can potentially be managed.

Software controls the workload, so software can potentially control the load. Electricity demand becomes programmable.

That property could become one of the most valuable differences between AI data centers and previous generations of industrial electricity demand.

14

The Energy Stack Is Becoming Part of the AI Competitive Moat

During the first phase of the AI boom, competitive advantage came from access to GPUs. During the second phase, it expanded to networking and cluster architecture. The third phase is incorporating energy.

SiliconAdvanced accelerators, CPUs, memory and networking.
SystemsRack architecture, liquid cooling, storage and power conversion.
SitesLand, fiber, permits and construction capacity.
EnergyGrid connection, generation, storage and fuel.
SoftwareOrchestration that converts the physical stack into billable output.

Companies that control more of these layers can potentially deploy faster and operate more efficiently. But vertical integration also increases complexity and capital requirements.

This is why the AI infrastructure winners may not necessarily be the companies with the lowest theoretical chip cost. They may be the companies that execute best across the entire chain.

15

What Does This Mean for AMD and Other Accelerator Competitors?

The power bottleneck is not an NVIDIA-only issue. AMD, custom accelerators and emerging architectures face the same physical constraints.

In fact, a world focused on tokens per megawatt could strengthen competition.

If customers evaluate AI systems through total output per watt and total cost per token rather than raw benchmark performance, alternative accelerators have more ways to compete. A chip that is slower in one benchmark but materially cheaper or more efficient for a specific inference workload can still win economically.

That is why AI Infra Summit discussions increasingly focus on memory bandwidth, networking, workload specialization and rack-level economics rather than only peak FLOPS.

AMD’s open ROCm strategy, Qualcomm’s efforts to enter AI data-center inference, AWS Trainium and Google TPUs all become relevant within this framework.

The real battle is increasingly between complete systems rather than individual chips.

NVIDIA’s advantage is that it already owns unusually large portions of the system. The risk is that the broader and more valuable the AI infrastructure market becomes, the more incentive hyperscalers and competitors have to design around NVIDIA’s margins.

16

Power Efficiency Can Be Bullish and Bearish for NVIDIA at the Same Time

There is an apparent contradiction in the DSX story.

If NVIDIA can help customers produce 24% more tokens from the same power budget, those customers may need fewer GPUs to produce a fixed amount of AI output. Why is that bullish for NVIDIA?

The answer depends on demand elasticity.

If demand for AI output is fixed, efficiency is deflationary for hardware demand. A customer that needs one billion tokens per day can satisfy the requirement with fewer resources.

But if lower cost per token creates new use cases and expands consumption, efficiency can accelerate overall demand. History in computing strongly supports the possibility that cheaper compute creates more computing.

The same dynamic occurred with storage, bandwidth and CPU performance. Dramatic efficiency improvements did not cause the world to stop buying computers. They enabled entirely new applications.

Agents that reason continuously, multimodal video generation, robotics, autonomous research systems and personalized inference could consume orders of magnitude more compute than current chat applications. If those markets emerge, improvements in token economics may expand the total addressable market faster than they reduce hardware intensity.

Core risk: if AI usage does not expand fast enough to absorb continuous efficiency improvements and massive new capacity, the same technologies that improve token economics could eventually contribute to overcapacity. Power scarcity delays that risk today, but it does not eliminate it.
17

Financing Is the Second Bottleneck

Electricity is not the only scarce resource. Capital is becoming another.

The IEA noted in its 2026 analysis that data-center investment has become too large to rely solely on technology-company balance sheets. The sector increasingly depends on debt markets, infrastructure funds, joint ventures, project financing and long-term customer commitments.

This matters because energy infrastructure extends the capital cycle.

A company that once needed to finance GPUs and leases may now finance land, substations, transmission equipment, cooling infrastructure, batteries and onsite generation. Every additional layer increases upfront investment.

CoreWeave’s financial statements illustrate the trade-off. The company can report enormous adjusted EBITDA while still carrying very large interest expense because so much capital must be raised before capacity becomes revenue producing.

For investors, that means the AI infrastructure boom cannot be evaluated purely through revenue growth.

Questions that matter: capex per incremental MW, time from contracted to active power, revenue per active MW, cost of financing, contract duration, residual hardware value and the economics of repurposing older accelerators for inference.
18

Utilities May Become an Underappreciated Part of the AI Trade

If AI infrastructure becomes one of the largest sources of electricity-load growth in the United States, utilities and energy infrastructure deserve far more attention from technology investors.

Data-center developers increasingly compete for sites based on power availability. Regions with abundant generation, transmission capacity, favorable permitting and constructive utilities can attract billions of dollars of investment.

But not all load growth is automatically attractive for utilities.

A utility may need to invest heavily in generation and transmission to serve a customer whose technology evolves quickly. If a speculative data-center project never reaches full utilization, existing ratepayers should not be left paying for stranded infrastructure. These concerns explain why ratepayer protection and cost allocation are central to FERC’s current work and to AEMA’s principles.

Long-term contracts, minimum demand commitments and customer-funded infrastructure will therefore matter.

The ideal AI customer from a utility perspective is not merely large. It is predictable, creditworthy, flexible and willing to pay the costs its connection imposes.

That is exactly the type of customer profile AEMA wants flexible AI factories to become.

19

Natural Gas, Nuclear, Renewables, Fuel Cells — There Will Not Be One Winner

The scale of the projected demand makes a single-source solution unrealistic.

Renewables will play a large role because solar and wind projects can often be developed relatively quickly and at competitive energy costs. Batteries can shift some generation and handle short-duration needs.

Natural gas remains attractive because it is dispatchable and existing pipeline infrastructure can support new generation in many regions, although turbine availability and permitting can become constraints.

Nuclear is increasingly attractive to hyperscalers because it offers dense, reliable, low-carbon baseload power, but new nuclear capacity generally operates on longer development timelines.

Fuel cells can provide distributed generation directly behind the meter, which is why Nebius’s Bloom Energy partnership is strategically interesting.

Geothermal may contribute in specific regions. Hydro remains valuable where geography allows it. Future small modular reactors could eventually serve campuses directly.

Likely outcome: an energy portfolio optimized for time-to-power, reliability, cost and emissions rather than one dominant generation technology.
20

Data Centers Are Becoming Energy Companies — Without Wanting to Be

Cloud companies historically wanted electricity to behave like an invisible utility input. Buy electricity from the grid, focus on computing, and leave energy infrastructure to specialized operators.

AI scale is breaking that separation.

Google, Microsoft, Amazon, Meta and specialized clouds increasingly sign long-term power agreements, invest in generation projects, secure nuclear output, explore geothermal, build batteries and negotiate directly with utilities.

Nebius is deploying fuel cells. CoreWeave is accumulating gigawatts of contracted power. NVIDIA is designing software that speaks to utility demand-response systems. Google is helping form an industry alliance around flexible load.

The technology sector is being pulled downstream into electricity because it cannot wait for the conventional grid-development cycle.

This is not necessarily permanent. If grid capacity catches up, energy may become less differentiating. But throughout the current infrastructure buildout, energy execution can be a competitive moat.

21

The Metric Investors Should Start Tracking: Revenue per Active Megawatt

Public disclosures around AI clouds are still inconsistent. Companies report GPU counts, active power, contracted power, backlog, capex and revenue using different definitions.

A useful future metric could be revenue per active megawatt.

This would not be perfect. Different facilities have different hardware generations. Training workloads have different pricing from inference. Utilization changes over time. Some companies own buildings while others lease them. Power reported as active may not equal IT load under comparable definitions.

But conceptually the metric connects physical infrastructure to financial output.

If two cloud companies each operate 1 GW but one generates substantially more revenue and gross profit from that power, the difference may reflect better hardware, utilization, pricing, software, customer mix or operational execution.

A second useful metric would be the conversion rate from contracted power to active power.

Securing 5 GW sounds impressive. If only 500 MW is operational, most of the economic value remains in the future. Investors should track the speed and cost of conversion.

22

AI Infra Summit 2026: The Industry Is Already Talking Differently

The agenda at AI Infra Summit 2026 provides a useful snapshot of how the infrastructure conversation has changed.

Sessions no longer focus only on accelerator performance. They discuss compute as capital, power constraints, token economics, memory bottlenecks, data movement, custom silicon, networking and infrastructure strategy.

NVIDIA’s Ian Buck centered his keynote on AI-factory efficiency. Other sessions explicitly addressed the power challenge in AI data centers. Amazon executives discussed optimization across the stack from silicon to services. Intel presented a systems-level AI strategy. Qualcomm demonstrated its attempt to enter data-center inference. AMD emphasized infrastructure openness.

This is the mature phase of a technology buildout.

During the early phase, demand is obvious and bottlenecks are simple: there are not enough GPUs. During the mature infrastructure phase, performance emerges from the interaction of dozens of systems. The winners optimize the whole chain.

23

What Could Break the Thesis?

AI demand could disappoint

The most obvious risk is that enterprise AI revenue does not grow fast enough to justify the enormous infrastructure investment currently planned. If model providers struggle to monetize usage or customers become more cost conscious, data-center utilization could fall.

Efficiency could outrun demand

Model optimization, quantization, sparsity, mixture-of-experts architectures and better inference software continuously reduce the amount of compute required per unit of output. If demand does not expand faster than efficiency, hardware requirements could undershoot current forecasts.

Power projects may arrive after technology changes

Transmission projects and campuses are long-duration assets. GPU generations are short-duration assets. A site designed around today’s infrastructure assumptions may need redesign as rack power density increases.

Community opposition could intensify

Data centers compete with residential and industrial users for electricity, water, land and transmission capacity. Projects that appear to raise local electricity bills or provide too few permanent jobs relative to infrastructure cost may face political resistance.

Financing conditions matter

High interest rates make infrastructure more expensive. AI clouds that rely heavily on debt are sensitive to credit markets. Even strong demand can produce poor equity returns if too much value accrues to lenders, equipment suppliers or customers.

Flexible load may be less flexible than advertised

A cloud operator can only curtail workloads that customers allow it to curtail. As AI moves from research into mission-critical production applications, the fraction of truly interruptible compute may fall.

24

Why AEMA Still Matters Even If DSX Is Not the Final Standard

One should separate the technology thesis from the company thesis.

DSX may become widely adopted. It may coexist with competing orchestration systems. Hyperscalers may build similar capabilities internally. Open standards may emerge. Different utilities may implement flexibility in different ways.

None of those outcomes invalidates the larger trend.

The important development is that AI compute is beginning to interact programmatically with the electricity system.

Once utilities recognize flexible data centers as a distinct class of load, the concept can survive regardless of which vendor supplies the control software.

NVIDIA’s strategic success would be making DSX one of the preferred control layers. But from an industry perspective, the bigger change is the creation of a market in which compute scheduling and grid scheduling become linked.

25

NVIDIA’s Moat Is Expanding Into Physical Infrastructure

NVIDIA is frequently described as a semiconductor company. That description becomes less accurate each year.

The company increasingly provides a vertically integrated computing architecture: accelerators, CPUs, networking, interconnect, DPUs, rack designs, libraries, inference software, orchestration and now facility-level optimization.

The more layers NVIDIA optimizes, the harder it becomes to compare competing products on chip specifications alone.

A competitor may offer an accelerator with attractive theoretical performance. But customers must compare the total system: networking, software compatibility, cluster reliability, memory architecture, power behavior, rack density, cooling and operational tooling.

DSX extends the moat one step further by arguing that the correct unit of comparison is not even the rack. It is the factory.
26

The Most Important Investment Read-Throughs

NVIDIA · $NVDA

The positive interpretation is that NVIDIA is extending its platform into the next bottleneck before power constraints can slow hardware deployment. The counterargument is that higher efficiency reduces the number of accelerators required for any fixed amount of AI output. The outcome depends on demand elasticity.

CoreWeave · $CRWV

The power thesis is central. Any technology that increases revenue-producing compute inside existing energized sites could be financially meaningful. Capital intensity and financing remain the core risks.

Nebius · $NBIS

Nebius is explicitly solving for time-to-power through infrastructure, efficiency and behind-the-meter generation. Growth depends on how quickly planned capacity becomes operational.

AMD · $AMD

A shift toward tokens per watt and cost per token creates another competitive battlefield beyond raw benchmark leadership. NVIDIA’s increasingly integrated factory architecture raises the systems-level bar.

Utilities and energy infrastructure: AI is creating one of the most consequential new electricity-demand cycles in decades. Opportunities extend into generation, transmission, transformers, switchgear, storage and cooling — but the sector also carries stranded-asset and ratepayer-allocation risks.
27

A Useful Framework for Following the AI Power Trade

MetricWhy it mattersWhat investors should watch
Contracted powerMeasures future site potential.Binding commitments versus speculative pipeline.
Connected powerShows infrastructure physically attached to available electricity.Interconnection timing and delays.
Active powerPower currently supporting live IT equipment.Quarter-over-quarter expansion.
PUEMeasures facility overhead.Cooling and electrical efficiency.
GPU utilizationDetermines whether installed hardware is economically productive.Cluster reliability and scheduling.
Tokens per megawattConnects electricity to AI output.Hardware + software + power optimization.
Revenue per active MWConnects infrastructure to financial output.Pricing, customer mix and utilization.
Capex per MWMeasures buildout efficiency.Equipment inflation and density.
Time to powerDetermines speed from investment to revenue.Grid queues, onsite generation and permitting.
Cost of capitalAI infrastructure is heavily financed.Interest expense and debt structure.
28

The Bigger Picture: Intelligence Is Becoming an Industrial Product

There is a deeper reason the language of AI infrastructure is changing.

Artificial intelligence began as software. At scale, it is becoming an industrial process.

An AI factory accepts physical inputs — electricity, semiconductor hardware, cooling capacity, data and capital — and converts them into an output: computational intelligence.

That makes the industry increasingly resemble other forms of large-scale manufacturing.

A semiconductor fab converts electricity, chemicals, wafers and equipment time into chips. A refinery converts feedstock and energy into fuels. A steel mill converts ore, electricity and heat into metal.

An AI factory converts electrons and capital into tokens.

Once viewed this way, NVIDIA’s “AI factory” language stops sounding like marketing and begins to describe an economic reality.

Factories are optimized around throughput. Their economics depend on utilization. Energy matters. Downtime matters. Logistics matter. Capital intensity matters. Marginal efficiency matters.

That is the framework investors increasingly need for AI infrastructure.

29

The Constraint Has Moved — and That Is Usually Where Value Moves Next

Every major technology cycle shifts bottlenecks.

At first, the bottleneck was accelerator availability. NVIDIA captured extraordinary value because GPUs were scarce.

Then networking became critical because thousands of GPUs had to communicate efficiently. High-speed interconnect and networking became strategic.

Then memory became a constraint, pushing HBM supply and packaging capacity into the spotlight.

Now power is becoming the next constraint.

Value tends to migrate toward whoever solves the current bottleneck.

That does not necessarily mean utilities become the next NVIDIA. Electricity infrastructure is regulated, capital intensive and often lower margin. But power scarcity changes bargaining power across the entire technology stack.

Sites with available electricity become more valuable. Generation projects gain strategic importance. Transformers become critical equipment. Cooling innovation becomes more valuable. Software that turns static loads into flexible loads can improve interconnection economics. Efficient chips become more valuable not only because they lower operating costs but because they unlock capacity that otherwise does not exist.

The AI trade is broadening.

30

Final Assessment: From GPU Scarcity to Megawatt Scarcity

The launch of the AI Energy Management Alliance and the first production-style results from NVIDIA DSX are important because they formalize a change that has been building quietly for several years: electricity is no longer a background input to AI. It is becoming a primary strategic resource.

Google, NVIDIA and Emerald AI are trying to create a framework in which AI data centers can connect to the grid more like intelligent, controllable systems than conventional static loads. NVIDIA is simultaneously building the software and reference architecture needed to manage power inside the facility. Lambda’s early results suggest dynamic allocation can recover meaningful stranded capacity. Silicon Valley Power’s flexible-load program shows that at least some AI workloads can respond automatically to utility signals at commercial scale.

CoreWeave and Nebius show why this matters financially. Both are racing to transform gigawatts of capacity into usable compute. Their growth depends not only on obtaining chips but on turning secured power into live infrastructure quickly enough to meet customer demand and earn an acceptable return on enormous capital commitments.

That creates a new hierarchy in AI infrastructure.

The first-order question remains whether customers want more AI.

The second is whether semiconductor companies can supply enough computing performance.

But the third — and increasingly decisive — question is whether the physical world can support the scale that the software world wants to build.

That is where electricity enters the story.

If the next generation of AI factories can dynamically allocate rack power, shift workloads, cooperate with utilities, deploy behind-the-meter generation, reduce cooling overhead and extract more tokens from each energized megawatt, the industry may be able to scale faster than conventional grid timelines would otherwise permit.

If it cannot, billions of dollars of GPUs, customer backlog and planned campuses may spend longer waiting for power than investors currently assume.

GPU count still matters. Model performance still matters. Networking still matters.

But increasingly, the infrastructure winners will be the companies that can answer a harder question: how much economically useful intelligence can you manufacture from each scarce megawatt you control?

That is the emerging battle behind NVIDIA DSX. It is the reason Google and NVIDIA are organizing around flexible data centers. It is why CoreWeave reports power capacity alongside customer demand. It is why Nebius is deploying onsite generation. And it is why investors following AI infrastructure should begin treating electricity not as an operating expense buried deep inside a data-center model, but as one of the fundamental units of the AI economy.

Follow the infrastructure, not only the chip

Merlintrader tracks AI, power, defense and space infrastructure through dedicated Stock Hubs, event coverage and long-form research. Continue with the latest company-specific research and the Space, Defense & AI Event Calendar.

Explore Merlintrader Stock Hubs

Primary & institutional research

AI-assisted research disclosure. Merlintrader uses artificial intelligence to assist with research organization, source comparison, drafting and editorial production. Material facts are checked against primary or institutional sources where practicable. AI assistance does not eliminate the possibility of errors, omissions or misinterpretation, and this publication is not an academic verification service.
Disclaimer. This material is provided for educational and informational purposes only. It is not investment advice, a recommendation, an offer or solicitation to buy or sell securities, or a substitute for independent professional advice. Company statements, projections, benchmarks and forward-looking claims remain subject to uncertainty and may change. Readers should independently verify material information from original sources and consult appropriately licensed financial, legal, tax or other professional advisers before making decisions. Merlintrader may update this report as new information becomes available.