Loading...
Research & Insights

Aurora's 96-MW Proof: Inside the Certification Bet on Software-Defined Grid Capacity

September 23, 2026 · Jason Ellis

Aurora's 96-MW Proof: Inside the Certification Bet on Software-Defined Grid Capacity

In late October 2025, five organizations that control large pieces of the AI and power economies made an announcement together: NVIDIA, the grid operator PJM Interconnection, the Electric Power Research Institute (EPRI), data center developer Digital Realty, and a startup called Emerald AI said they would build the world's first "power-flexible AI factory," a 96 MW facility called Aurora in Manassas, Virginia, designed to prove that a data center can act less like a passive power drain and more like a controllable grid asset.

The promise is not incremental. Emerald AI and its partners say software orchestration can cut a data center's electricity draw by roughly a quarter during grid emergencies while still meeting performance commitments to AI customers, and that if new data centers broadly adopted the model, the United States could unlock on the order of 100 GW of existing grid capacity without waiting for new transmission and generation to catch up. For a power market where capacity auction prices have risen roughly ninefold in two years, largely on data-center-driven demand, anyone paying an electric bill has a stake in whether that claim survives contact with reality.

The engineering turns out to be the least contested part. The real question is institutional: whether grid operators, federal regulators, and state utility commissions will treat software-shifted compute as something firm, verifiable, and payable, the same standing a gas peaker plant enjoys. Aurora is the testbed for that conversion.

What Aurora is meant to prove

A power-flexible AI factory is a simple idea with a hard implementation. Instead of demanding full power at all hours regardless of grid conditions, the facility commits, in advance and verifiably, to shaping its electricity draw along boundaries the grid operator can count on during stress.

The Aurora facility, under construction by Digital Realty and slated to open in the first half of 2026, is the first building designed against what the coalition describes as a new reference design and certification standard for power-flexible infrastructure. Northern Virginia is the world's largest data center market, served by Dominion Energy within PJM's footprint, and grid constraints there are no longer theoretical; interconnection timelines and transmission bottlenecks have become the binding constraint on AI expansion.

Aurora's job is to host demonstration testing under EPRI's DCFlex initiative, which studies how data centers can provide grid services. The tests, as described in the coalition's announcements, will simulate conditions such as summer heatwave demand spikes and sudden drops in renewable generation, measuring whether the facility can ramp power down to support the grid while workload orchestration preserves quality of service for AI training and inference jobs.

In March 2026, NVIDIA raised the stakes. At CERAWeek, it unveiled the Vera Rubin DSX AI Factory reference design, including a software library called DSX Flex built to connect AI factories to power-grid services, alongside a coalition that added energy companies AES, Constellation, Invenergy, NextEra Energy, Nscale Energy and Power, and Vistra. The reference architecture also contemplates "hybrid" factories that start on co-located generation and storage as bridge power, then later use those resources to supply the grid. That detail matters: the coalition's own language concedes that new generation remains necessary in many places. Flexibility accelerates interconnection; it does not manufacture energy.

How does the software shed load without breaking its promises?

The mechanism is a control stack that sits between the grid operator and the GPU cluster. Emerald AI's GridLink product reads the grid's real-time needs; its Conductor platform then allocates power limits across running jobs, integrating with NVIDIA's software stack, including NIM microservices and Mission Control. NVIDIA's March 2026 announcement describes Conductor coordinating compute flexibility alongside onsite generation, batteries and other behind-the-meter resources so operators can meet power targets while protecting priority workloads.

Close-up of tower servers in a data center with blue and red lighting.
Photo by panumas nikhomkhai on Pexels

The actual knobs are more mundane than the marketing. According to software and pseudocode Emerald AI released on GitHub from its Arizona field demonstration, Conductor profiles each job for throughput across different GPU allocations, sorts workloads into buckets based on how much performance loss they can tolerate, and then applies interventions: DVFS power capping that lowers GPU clock and voltage, pausing and checkpointing of interruptible jobs, reallocation of GPU counts, and fair-share distribution of reductions across buckets so no single tenant absorbs the whole cut. A June 2026 paper the Emerald-led team submitted to IEEE adds performance-aware geographic shifting: migrating work from a cluster in a stressed region toward clusters where the grid has slack.

Writing in Utility Dive as an Emerald AI adviser, Arushi Sharma Frank lists the capabilities the platform is designed to deliver:

  • Peak reductions of 20 to 30 percent held for multi-hour windows, without the rebound spike that traditionally plagues demand response events
  • Sustained curtailments up to 10 hours
  • Multiple discrete events in a single day at different depths and durations
  • Smooth ramps down and up over 5, 15, or 30 minutes
  • Response on roughly 10 minutes' notice, or up to 2 hours' planned notice
  • Carbon-aware operation that tracks a 5-minute marginal CO2 signal
  • Mapping of wholesale price signals into power targets and day-ahead bid curves
  • Replays of historical grid emergencies, including the Polar Vortex, as proofs

That menu is the ambition, not yet the certified product. Frank's piece is candid about the context: it quotes data center veteran David Mytton's verdict that demand response research has historically been "completely disconnected from commercial realities." The whole Aurora wager is that software has changed that verdict.

What the Phoenix experiment proved, and what it has not

High-voltage electrical substation with steel structures and power lines at dawn.
Photo by Robert So on Pexels

The strongest evidence that flexible compute is real comes from the Salt River Project desert, not Northern Virginia. In May 2025, Emerald AI, NVIDIA, and Oracle ran a field demonstration on a 256-GPU cluster with Arizona public power utility SRP, under EPRI's DCFlex umbrella. As the Emerald-led IEEE submission recounts, the cluster cut electricity consumption by 25 percent during simulated peak demand periods while maintaining service levels for its jobs. The team also re-enacted the August 2020 California heat event, the one that led grid operators to call rotating outages, replaying CAISO's net-demand curve against the cluster's measured power draw. Unusually for a vendor demo, the underlying data and key orchestration code were published on GitHub, allowing outside scrutiny.

The newer paper also reports a real-world deployment on a 130 kW GPU cluster, demonstrating rapid load reduction, sustained curtailment, carbon-aware operation, and load shifting between geographically distributed clusters. Emerald AI has also announced a UK demonstration with National Grid and NVIDIA under the U.S.-U.K. Tech Partnership.

Two qualifications deserve equal billing with the results. First, scale: Aurora is roughly 700 times larger than the 130 kW test cluster, and nothing about orchestration fairness logic changes the physics of whether a 96 MW facility full of paying tenants behaves as politely as a curated experiment. Second, provenance: the published results so far come from teams led by Emerald AI, with NVIDIA, Oracle, EPRI, and National Grid personnel as co-authors. NVIDIA is not a neutral party either; Emerald AI is an NVentures portfolio company, disclosed in the October announcement. None of this invalidates the data, which is unusually open, but Aurora's DCFlex tests are the first results that will carry independent certification weight. They remain the moment of truth.

Where does the 100-gigawatt figure come from?

The headline number is the coalition's own estimate. The October 2025 announcement asserted that the new power-flexible reference design, if adopted nationwide, could unlock an estimated 100 GW of capacity on the existing electricity system. The logic underneath is straightforward. U.S. power systems are built to serve peak demand but sit underutilized during most hours of the day. Treat a data center as a firm load and it triggers worst-case planning; treat it as a flexible load that commits to curtailment during a small fraction of peak hours, the argument goes, and existing headroom can absorb far more connection requests.

The coalition's framing, repeated in NVIDIA's March 2026 release, is that power-flexible AI factories "can help unlock up to 100 gigawatts of capacity across the U.S. power system by combining optimized infrastructure design with efficient use of existing assets and, where needed, new-build generation." The October announcement translated that as roughly 20 percent of annual U.S. electricity consumption.

Independent work points in a similar direction, if with wider error bars. A widely cited February 2025 Duke University study by Tyler Norris and colleagues at the Nicholas Institute, analyzing hourly load data from the largest U.S. balancing authorities, found that the system could host new large loads totaling on the order of tens to a hundred-plus gigawatts if those loads accepted curtailment during a small percentage of hours.

Retro industrial control room with vintage panels and dials in an old power plant.
Photo by Pixabay on Pexels

Treat the number as an estimate with a scope, not a discovery. It measures potential peak-hour interconnection headroom, not guaranteed deliverable power. It says nothing about whether enough energy can be generated at acceptable cost across all hours; adding large round-the-clock loads still requires energy supply, which is why six power producers signed the March coalition. And grid headroom is relentlessly local: spare capacity in Ohio does not help a project in constrained Northern Virginia. The 100 GW figure is defensible as an order-of-magnitude indicator that flexibility buys speed to power. It is not a promise of 100 GW of served load without any new wires or plants.

Would PJM pay a data center the way it pays a gas plant?

This is the gate on which the entire thesis turns, because capacity markets are where reliability is converted into money.

PJM's Reliability Pricing Model pays generators and demand resources to be available in future delivery years, and three recent base auctions tell the story of the AI crunch in numbers. Clearing prices went from roughly $29 per MW-day for the 2024/25 delivery year to about $270 for 2025/26 and around $329 for 2026/27, according to PJM's auction results, with PJM's market monitor and outside analysts attributing much of the increase to forecast data center demand colliding with generator retirements.

A gas peaker earns its capacity payment by accepting an obligation: perform whenever PJM declares emergency hours, with steep penalties for failure under the Capacity Performance rules. Demand response has earned money in this market for two decades, but in recent years PJM has tightened the bar, eliminating its limited "Base" capacity product and requiring resources to perform across more of the year.

So the question for an Emerald-orchestrated data center is not whether it can shed 25 percent in a May demonstration in Phoenix. It is whether it can sign a commitment the way a power plant does: available on every emergency hour PJM might declare, including a December 2022-style winter storm like Elliott, when the grid operator ran emergency operations and generator failures triggered penalty assessments lasting days, not the 10-hour sustained curtailment the Emerald capability list claims. Capacity accreditation is increasingly calculated on marginal contribution during the tightest hours, and a curtailment promise that expires at hour 11 of a 60-hour cold snap accredits differently than a turbine.

The money at stake is real but modest against AI economics. If 24 MW of Aurora's 96 MW were certified as firm curtailment, the 2025/26 clearing price would imply capacity revenue on the order of $2.4 million a year, an illustration, not a forecast, since rules would determine the actual payout. Against the revenue density of a modern AI cluster, that is a secondary income stream. The far larger prize is the one flexibility was actually invented to win here: permission to connect years earlier, which a capacity payment alone would never justify.

That tension defines the design problem. Markets must pay enough to make the firmness credible, and planners must credit enough to make interconnection faster, but a load paid for both its connection speed and its capacity raises double-recovery questions that market monitors exist to police. PJM's independent market monitor has a long institutional memory of demand response measurement disputes, and any software-mediated megawatt will be examined for whether the baseline, the counterfactual of what the facility would have consumed, is being gamed by design.

Demand response, Bitcoin miners, and batteries: the scaling precedent

Software-flexible compute is not the first resource to ask the grid to trust behavior instead of steel. Three precedents mark the trail.

A power plant with smokestack on a sandy coastline, showcasing industrial architecture.
Photo by Joseph Russo on Pexels

Traditional demand response scaled through exactly this machinery. FERC's Order 745 in 2011 required wholesale markets to compensate demand reductions comparably to generation; the D.C. Circuit struck it down; and the Supreme Court restored it in FERC v. Electric Power Supply Association in 2016. FERC's 2020 national assessment put wholesale demand-response capability on the order of 30 GW. That took on the order of fifteen years of rulemaking, litigation, and measurement fights to accumulate.

Behind the meter, FERC's Order 2222 in 2020 told grid operators to open markets to aggregated distributed resources, and implementation then crawled through compliance proceedings, jurisdictional disputes with states, and repeated deadline extensions. The lesson is not that aggregation works; it is that multi-jurisdiction market plumbing takes years even after the legal authority is clear.

Texas has shown what the faster version looks like when the deal is connection speed rather than market revenue. ERCOT's registration requirements for large flexible loads and its curtailment arrangements with crypto miners effectively trade interruptibility for grid access, and miners curtailing during storms earned credits worth more than running flat out. In 2023, Google signed what it described as its first demand-response agreement for a U.S. data center, with Omaha Public Power District, shifting compute during grid stress. These were ad hoc bilateral deals. Aurora's certification attempt is, at bottom, an effort to turn ad hoc flexibility into a standardized, auditable product class.

Batteries set the floor test for value. A grid battery can inject power, absorb surplus, respond in milliseconds, and earns accredited capacity under well-established rules, typically for two to four hours of duration. Software-shifted compute can do none of the injection side, and it cannot respond without affecting paying workloads, but it can, if the claims hold, sustain reductions across 10-hour windows that would exhaust a standard 4-hour battery, at near-zero incremental capital cost, using infrastructure the grid never had to buy. The honest comparison is that compute flexibility and storage are complements, with batteries able to firm up exactly the failure mode, the long multi-day emergency, where compute curtailment is weakest.

The harder gate: interconnection and federal rules

Electrician installing a home energy battery storage unit. Renewable and sustainable energy solution for residential use.
Photo by Elite Power Group on Pexels

If the capacity market is where flexibility earns money, interconnection is where it earns time, and time is the scarcer asset in the AI buildout. More than 2,000 GW of generation and storage projects, roughly double the installed U.S. fleet, sits in interconnection queues studied by Lawrence Berkeley National Laboratory, which is a large part of why the March 2026 coalition centers co-located "bridge power." On the load side, new data center connections in constrained regions can face years of transmission study and upgrade cycles.

This is precisely the problem EPRI designed its flexibility framework to attack. The DCFlex Flex MOSAIC framework, published in 2026, defines standardized flexibility classes mapped to grid needs such as congestion relief, system peak mitigation, prolonged low supply, and fast intra-day balancing, each specified by duration, notice period, and availability. Its stated purpose is to let a large load walk into a connection study and declare its capabilities transparently against standardized classes, rather than being modeled as a bespoke worst case. Five design principles run through it: technology neutrality, transparency, completeness, standardization, and fit with existing processes. Adoption by utilities in interconnection studies is where this document either becomes infrastructure or stays a PDF.

The federal layer moved in late 2025. The Department of Energy petitioned FERC to consider asserting jurisdiction over large-load interconnection to the transmission system, proposing standardized national procedures, with flexibility reportedly among the criteria for favorable treatment. PJM, meanwhile, spent 2025 in an accelerated stakeholder process over how to handle large loads, with concepts on the table including requirements that new data centers arrive with their own supply or commit to curtailment capability. The outcomes of those proceedings will largely determine whether flexible interconnection becomes a standard offer or remains a pilot privilege.

One more federal actor looms over all of it. NERC's reliability work has documented that data centers can create a new kind of grid risk: in 2025, discussion inside NERC's large-load task force surfaced a July 2024 transmission fault in Northern Virginia in which roughly 1,500 MW of data center load dropped from the grid nearly simultaneously as facilities shifted to backup power. Orchestrated, telemetered flexibility is arguably the cure for exactly that failure mode, but the episode pushed NERC's large-load task force toward performance and ride-through expectations that any certification regime will have to incorporate.

Who pays for the peak?

The retail side of the ledger is where the politics live, and Virginia is the front line. The General Assembly's research arm, JLARC, reported in December 2024 that data centers were on track to drive Dominion's load into territory requiring a near-doubling of generation, and estimated that under unmanaged cost allocation, typical residential bills could rise by roughly $14 to $37 a month by 2033. Dominion has since proposed moving the largest facilities into a separate rate class with long-term minimum-demand obligations, in a State Corporation Commission proceeding testing how much of the risk of overbuilding should fall on households versus shareholders.

Elsewhere in PJM, the politics have already produced guardrails. In April 2025, Pennsylvania's governor announced a settlement with PJM imposing a temporary price collar, a floor of $175 and a cap around $325 per MW-day, on coming capacity auctions, an intervention aimed squarely at blunting data-center-driven costs.

The utility sector's own rhetoric, notably, has shifted toward the flexibility thesis. Constellation CEO Joe Dominguez, quoted in NVIDIA's March 2026 announcement, put it bluntly: the country does not have a supply problem, it has a peak problem, and demand response addresses it. AES CEO Andrés Gluski said DSX Flex lets AI infrastructure "operate as a grid asset" that accelerates clients' time to power. When the companies that would sell the new generation start arguing that better use of the existing system comes first, the idea has moved rooms.

Moody view of industrial cooling towers with smoke against a twilight sky, emphasizing industry impact.
Photo by Safir Khan on Pexels

The rate-design question that remains is distributional. If Aurora earns flexibility revenue, who receives it, the data center, its tenants, or ratepayers, and does a facility that paid less for its interconnection because it promised flexibility then owe the grid that flexibility without additional capacity-market pay? Regulators will not bless a model in which data centers receive speed for free, revenue for the same attribute, and leave residential customers holding the transmission upgrades. Any durable national standard has to answer that accounting question before it answers the engineering ones.

The risks a certificate cannot paper over

The honest ledger of unresolved problems starts with evidence: none of the published flexibility results comes from Aurora's certification testing. They come from vendor-led teams at sub-megawatt scale and have not yet survived independent peer review, a process the June 2026 paper is currently undergoing.

Workload physics impose the next limit. Training jobs tolerate pause and checkpoint; batch inference shifts; latency-sensitive interactive inference does not, and cooling introduces its own inertia. A grid emergency during an inference-heavy afternoon simply contains less flexible load than a training-heavy one, which means the certified rating of any facility depends on its tenant mix, and certification will need to price that variance rather than assume it away. Then there is duration risk: winter events like Elliott and the 2021 Texas freeze ran for days, far beyond the longest curtailment yet demonstrated, and the mass-disconnection events NERC has examined show that data center load itself can behave like a contingency event.

Market-integrity risk is the third ledger item. Baseline estimation for demand reductions is a known fight; a resource paid for reductions relative to a counterfactual baseline it influences is a resource with an incentive to inflate that baseline. And the 100 GW arithmetic, however sound, is a statement about peak headroom, not about annual energy or local deliverability. Grid planners interviewed across the 2025 debates keep returning to the same point: flexibility accelerates connection, but the MWh still have to come from somewhere, and somebody still has to build that somewhere.

Finally, the institutional capture question deserves plain statement: the certification's chief evangelists include its vendor, its vendor's strategic investor, and a utility-funded research institute. EPRI's DCFlex is the closest thing the market has to an independent testbed; vendor slides alone will not carry an interconnection study. The system is pointed in the right direction. It has not yet produced the independent result.

Where things stand in the fall of 2026

The sequence is now clear. A 256-GPU Phoenix demonstration in May 2025 produced the first hard numbers. The October 2025 announcement attached a 100 GW estimate to the model and named Aurora as the certification testbed, with PJM as a partner. Late 2025 brought DOE's interconnection petition and PJM's large-load process. March 2026 turned the concept into product architecture, the DSX reference design and six power producers. June 2026 put the technical claims into the IEEE review pipeline. Aurora's DCFlex demonstration results are the next data point that decides whether the reference design is a standard or a brochure.

What would prove the concept is not another percentage from a controlled cluster. It is a small set of institutional artifacts: an interconnection agreement that grants speed in exchange for metered flexibility classes; a PJM or FERC rule that lets certified curtailable load earn accredited, penalized capacity the way a gas plant does; a retail tariff that moves flexibility value to ratepayers without double-paying the load; and a NERC-aligned performance standard the facility meets during a real event rather than a replayed one.

A gas peaker plant is paid because it is firm, measurable, and punished when it fails. Aurora is a bid to make 96 megawatts of software-orchestrated GPUs meet the same identity. The software has already shown it can bend the load. The question the certification is actually answering is whether anyone will write that behavior into a tariff, put a price on it, and collect when it doesn't show up.

Sources/References

Beautiful landscape view of Phoenix, Arizona during sunset with cityscape and desert hills.
Photo by Nicole Seidl on Pexels
Share this article

Comments (7)

  • carmen_shah Sep 23, 2026

    Can you dig into how EPRI's DCFlex certification will actually interface with PJM's capacity market rules? I'm curious whether software-shifted compute would clear alongside a gas peaker or get siloed into a separate product.

  • Rafael C. Sep 23, 2026

    My worry is that state utility commissions in PJM footprint states move on much slower timelines than federal regulators, so a facility-level certification from EPRI may not translate into anything billable for years. The gap between 'verifiable' and 'paid like a peaker' is where most of these programs quietly die.

  • victor.silva Sep 23, 2026

    That line - 'flexibility accelerates interconnection; it does not manufacture energy' - is doing a lot of quiet work in the piece, and I think it's the most honest sentence here. The hybrid factory concept with co-located generation and storage almost concedes the point, which makes me wonder how much of the 100 GW headline survives once you net out the sites that will still need new megawatts.

  • aisha.patel Sep 23, 2026

    After three demand response events on our cluster last year, the rebound spike was the worst part - we ate through any savings the moment we ramped back.

  • benc24 Sep 23, 2026

    We ran a small demand response pilot last summer cutting GPU cluster load by about a third during a peak event and watching the workloads gracefully throttle was exactly the kind of thing this article describes - the DVFS capping and checkpointing actually worked far better than I expected.

  • Quinn K. Sep 23, 2026

    The piece's framing of capacity auction prices rising roughly ninefold in two years echoes what I've been watching on the PJM side, where interconnection queue reform has stalled behind a separate but adjacent debate about whether demand-side resources should clear the same capacity market as new gas peakers. I keep coming back to the 100 GW number, though, because it's a load-weighted figure that assumes broad adoption across hyperscalers, and the history of voluntary grid programs suggests adoption curves flatten fast once operators discover the cost of interrupting inference jobs. The hybrid factory concept, where bridge generation gets repackaged as a grid asset later, feels like an honest acknowledgement that software alone won't reach the auction clearing prices regulators are starting to demand. I'm not sure Aurora alone resolves that, but it's the first facility I've seen built explicitly to generate the kind of telemetry that could move a regulator off the fence.

  • A. Hayes Sep 23, 2026

    The 20 to 30 percent peak reductions held for multi-hour windows lines up with what we measured in our own orchestration testing, where a heterogeneous mix of training and inference jobs shed about a quarter of draw over a four-hour simulated heatwave without breaching our SLA on inference latency. The detail that surprised me was how much of the savings came from pausing and checkpointing interruptible jobs rather than from DVFS alone, because clock capping on modern accelerators leaks power through the rest of the system. The fair-share distribution across buckets so no single tenant absorbs the whole cut is genuinely clever and maps to how we ended up structuring reductions internally. The carbon-aware piece tracking a 5-minute marginal CO2 signal is new to me and I'd love to see how that plays with day-ahead bidding.

Comments are reviewed before they appear.

Continue exploring

Aug 26, 2026

The AI Energy Gridlock

The artificial intelligence race is colliding with an older, slower machine: the electric grid. A frontier AI campus can require…