Throughput per GPU was the number that mattered in AI infrastructure for most of 2024 and 2025. It has quietly stopped being the number that matters. On September 16, Emerald AI, Google and NVIDIA launched the AI Energy Management Alliance (AEMA), a coalition that grades a data center on measurable grid behaviour — response speed, duration, predictability — rather than on how many accelerators it contains. On October 8, PJM approved its first three projects under its Expedited Interconnection Track: 2,115 MW of storage and gas uprate, fast-tracked because those resources can show up within three years.
Those two events describe one operating rule: the facility that can prove it will bend when the grid is stressed gets connected first. That rule is not confined to gigawatt campuses. A single RTX 3090 on a home bench is a controllable load, and the same trade — throughput for power — shows up at 200 watts as it does at 100 megawatts. I measured it on my own machine, and the result is not the one the efficiency framing implies.
What “flexible data center” actually means
Grid flexibility is not a vague sustainability claim. It is four specific moves, and only the last two are worth anything to a grid operator:
- Shave — carry a peak with on-site batteries or generation instead of drawing it from the grid.
- Shift — run deferrable compute (training, fine-tuning, batch inference, embedding jobs) during off-peak hours or when local generation is surplus.
- Shed / curtail — cut metered demand by a committed amount for a committed duration, measured against a baseline, and get paid or get a lower rate for it.
AEMA’s principles are performance-based and technology-neutral: ride-through, curtailment and contingency-response obligations defined before a facility connects, common performance metrics, and interconnection costs allocated by actual system impact. Emerald AI’s chief executive Varun Sivaram put the qualification standard plainly — facilities should get faster or larger connections only if their ability to cut demand is “verifiable and enforceable.” A flexibility promise you cannot meter is not a promise.
What happened in the last two weeks
The PJM approvals on October 8 are the first concrete output of a process FERC approved in June and that opened its application window on July 31. The three accepted projects are Engie IR Holdings’ 860 MW storage site in Crawford, Pennsylvania, Engie’s 800 MW storage site in Morrow, Ohio, and a 455 MW uprate by LS Power at the 810 MW Hunterstown combined cycle plant in Adams, Pennsylvania. PJM expects all three online in mid-2029. The track is capped at 10 requests a year, requires at least 250 MW of unreliable capacity, requires backing from the primary siting authority, and sunsets at the end of 2027.
Two details matter more than the headline megawatts. First, 1,660 MW of the 2,115 MW — about 78% — is storage, not generation, and storage is the asset that lets a facility honour a curtailment commitment without stopping its work. Second, the track exists because PJM has added little generation in recent years while data centers pushed its demand forecast upward, raising capacity costs and utility bills in the region.
The same logic runs at the transmission level. The Department of Energy’s SPARK program funds 31 transmission projects across 26 states — $1.9 billion in grants against $5.25 billion of project value, finalised between October 2026 and January 2027. Only about 7% is attributed to AI outright, but half the value sits in the “new large loads” category, the industry’s catch-all for data centers. A new power plant does not help a data center unless the transformers and local circuits can deliver those electrons to that exact site.
Why inference is the hard part
Training is deferrable. Inference is not — a job that answers a user has to run now, and inference is 80 to 90 percent of AI compute. That asymmetry is why flexibility numbers are ranges rather than commitments.
EPRI’s technical leader Arin Kaye reports surveyed data centers identified peak reduction potential of 10% to 30% depending on facility type. A Boston University PEACL lab simulation puts AI data center flexibility at 18% to 55% of average power consumption via hardware power capping and power-aware scheduling. The mechanism is simple: capping lowers clocks, scheduling moves work.
There is a subtler lever most coverage misses. Quantization is a dispatchable resource. A June 2026 paper on grid-responsive LLM inference maps each model-precision configuration to dispatchable parameters and reports a 34.3% reduction in operating cost without curtailing served token volume. In practice a serving stack can drop from BF16 to FP8 during a grid event and keep answering requests at lower power — the same lever that gives 30-40% more tokens per second at identical TDP in normal operation. The difference is that during a grid event it becomes a grid service instead of a cost saving.
Backend choice changes the arithmetic too. Continuous batching, KV cache pressure and quantization support differ enough between serving engines that the same hardware can produce meaningfully different tokens per watt; the tradeoffs are laid out in the 2026 LLM backend battle between vLLM, Ollama and llama.cpp.
What power capping does at home: a measured test
I ran the trade on my own hardware — a Ryzen 5900X with an RTX 3090 (max power limit 390 W, stock cap 370 W) and 64 GB DDR4, serving a 27B model through llama.cpp. Same prompt, same 300-token budget, three runs at the stock cap, three runs at 200 W, with power draw sampled every 1.5 seconds during each request.
At the stock 370 W cap the card sat at roughly 360 W during generation and produced about 31.6 tokens per second in steady state. At 200 W it held exactly 200 W and produced about 14.4 tokens per second. The power reduction was 38.9%; the throughput reduction was 54.5%.
The part worth arguing about: efficiency per watt got worse, not better. Uncapped the system produced 0.096 tokens per second per watt; capped, 0.072. Energy per token rose from about 10.4 joules to 13.9 — roughly 34% more energy per token. Power capping is not an efficiency measure, it is a demand measure. It buys a lower peak at the meter, which is exactly what a grid operator pays for, and it costs throughput and, on this hardware, energy per token.
One more number for anyone building an agent loop: the first request after the model loaded ran at 10.8 to 12.4 tokens per second wall-clock against roughly 32 tokens per second once warm. That gap is prefill and tokenization overhead, not decode speed. Planning a fleet around engine-reported throughput means planning around the wrong number — the agent experiences the wall-clock figure.
The honest framing for a home lab: you cannot sell flexibility, you can only avoid paying for it. A time-of-use rate with a demand charge is the home-scale version of the same bargain — run deferrable work (batch inference, embeddings, fine-tuning, indexing) when the meter is cheap and let interactive work run at full power. The scheduling logic is identical to what Emerald Conductor does at 96 MW; the price signal is just smaller.
What to measure before you claim anything
- Set the cap and confirm it holds.
nvidia-smi -pl 200sets the power limit, and it must be re-established after each driver load unless you use persistence mode or the out-of-band SMBPBI path. Verify during generation, not at idle. - Meter at the outlet, not the GPU. Board power, fans, PSU losses and the CPU are part of the bill. A plug that reports real power is what makes a reduction claim verifiable — the same standard AEMA is pushing for facilities.
- Report wall-clock tokens per second. The gap between engine-reported and wall-clock throughput is the overhead an agent loop actually pays.
Gear for measuring your own power envelope
A THIRDREALITY Zigbee smart plug with real-time power monitoring is the cheapest way to get a continuous power reading on a single machine — around $15 for the single plug, about $39 for a four-pack, and it reports current, voltage and cumulative kWh into Home Assistant or SmartThings, which is enough to build a baseline and prove a reduction. For anything you want to defend to a utility, a Fluke 117 clamp meter (about $239) measures AC current on the conductor without interrupting the circuit, which is how you separate a GPU’s reported draw from what the panel actually delivers. And a CyberPower CP1500PFCLCD at about $240 is the home-scale version of the storage asset PJM fast-tracked: it is not a grid service, but it is the reason a capped machine keeps serving through a disturbance instead of dropping. If you are choosing the machine itself first, the hardware tiers are covered in the best hardware for running local LLMs in 2026, and the rack-scale power problem that made all this necessary is in how 120 kW racks rewrote data center power.
What this means
The metric is moving from throughput per GPU to tokens per watt, because the binding constraint is now the grid connection, not the silicon. AEMA and PJM’s fast-track are the same policy at two scales: prove you can bend, get connected faster. On a single consumer card the trade is measurable in an afternoon — 38.9% less power for 54.5% less throughput, and worse energy per token, which is the number nobody puts in the press release. Measure your own curve before adopting the efficiency framing. The operators who can prove a reduction, at any scale, get the cheaper connection.
