LLM-generated.
TL;DR Link to heading
Having spent a career reasoning about where the real limits of a machine sit, I read the ML.ENERGY tutorial at NeurIPS 2025, “Energy and Power as First-Class ML Design Metrics”, as a marker of a shift the architecture community has seen coming: electrical power, not chips or algorithms, is now the binding constraint on AI scaling. Once you accept that, the engineering agenda writes itself: measure GPU power and energy without fooling yourself, optimize performance when the real budget is watts rather than clock cycles, and reduce energy by moving a workload along its time–energy Pareto frontier instead of blindly minimizing joules. What follows is that argument and the four systems that discharge it (Zeus, Perseus, DynamoLLM, and the ML.ENERGY Benchmark), read the way an architect reads a design: mechanism first, then the numbers that decide whether it matters.
What Power and Energy as First-Class ML Metrics means?
- Reframes energy as a first-class computing resource, on par with FLOPs and memory bandwidth: something you measure, optimize under, and reduce.
- Nails down the correct way to measure ML power/energy (RAPL for CPU/DRAM, NVML/AMDSMI for GPU/HBM) and the single most common bug that silently corrupts results: unsynchronized CUDA.
- Shows that on modern GPUs performance and power are inseparable: sustained clocks are set by a moving-average power controller reacting to workload composition and even input data.
- Establishes the time–energy Pareto frontier as the object to reason about, and surveys four systems that each pick a better point on it at a different layer of the stack.
- Grounds all of it in production reality via a panel with the vendors and hyperscalers who live inside these power budgets.
Table of contents
- 1. Why Now: Power Is the Binding Constraint
- 2. Measuring Power and Energy Correctly
- 3. Optimizing Performance Under a Power Cap
- 4. Cutting Energy Along the Time Energy Frontier
- 5. The Industry Panel
- 6. Key Takeaways
- 7. The Complete Reading List
1. Why Now: Power Is the Binding Constraint Link to heading
Opening framing. Slides.
1.1 The problem Link to heading
The opening is diagnosis rather than a research result, and the diagnosis is correct. AI has reached unprecedented scale along three orthogonal axes: Large (frontier models such as Kimi-K2 at ~1T parameters, with closed models suspected larger), Numerous (thousands of specialized models), and Everywhere (deployed worldwide, serving millions). That scale drives ballooning demand for accelerators and gigawatt-class datacenters. And the thing you actually cannot buy quickly is electrical power.
1.2 The evidence: a news timeline Link to heading
The case is a chronological wall of 2024–2025 headlines, and it is hard to argue with:
| Date | Reported move | Power/scale signal |
|---|---|---|
| Jan 2024 | Meta buying 350,000 H100 GPUs | ~600,000 H100-equivalent total compute |
| Apr 2025 | Meta plans a datacenter “the size of Manhattan” | “hundreds of billions” in AI spend |
| May 2025 | xAI’s Memphis site | proposed 1.56 GW gas plant, up to 90 turbines |
| Sep 2025 | OpenAI + Oracle + SoftBank expand Stargate | $500B, 10 GW targeted by end of 2025 |
A wry aside relabels “350,000 [GPUs]” as “350,000 refrigerators”, a fair jab at how casually we report enormous power-hungry purchases as bare device counts, with no sense of the load behind them.
The core structural tension:
Next-generation frontier models “were due yesterday,” but large-scale power procurement (siting, permitting, approval, construction) takes years. That cadence mismatch is the bottleneck the rest of the tutorial attacks.
1.3 Why it matters Link to heading
Three constituencies are on the hook: AI builders/providers (power capacity gates both research and serving capacity), hardware vendors (“no one can buy new hardware without the energy to run them”), and users/the public (grid stability and electricity prices). The conclusion I draw is straightforward: GPU power draw, energy-aware scheduling, and measurement tooling now belong on the same design dashboard as raw FLOPs.
2. Measuring Power and Energy Correctly Link to heading
Session 1. Slides.
2.1 Power vs energy Link to heading
- Power (watts, J/s) is the instantaneous rate delivered by the grid, bounded by a hardware/datacenter power ceiling. It is the constraint.
- Energy (joules or Wh) is the time-integral of power: the area under the power-vs-time curve. It is the cost.
Conflating the two is the root of many bad benchmark claims. Power is the constraint; energy is what you are billed for.
Because datacenter capacity is itself provisioned in power units (e.g., 1 GW), the natural efficiency metric collapses cleanly:
throughput / power = (work/s) / (J/s) = work / energy
which is why the field increasingly reports tokens per joule rather than raw watts.
2.2 The tools: RAPL, NVML, AMDSMI Link to heading
Accurate system-level accounting means stitching two instrumentation families together, using vendor counters rather than external meters for the compute components:
| Component | Tool | What it exposes |
|---|---|---|
| CPU package | Intel RAPL (via sysfs) | energy/power; supported on most Intel and AMD |
| DRAM | Intel RAPL (DRAM domain) | energy; only some Intel CPUs |
| NVIDIA GPU / HBM | NVML (C++ + Python bindings) | power draw (1-s average or instantaneous) and cumulative energy since driver load |
| AMD GPU / HBM | AMDSMI | the same counter families as NVML |
Two subtleties matter in practice:
- NVML reports power at whole-module (e.g., a Grace Hopper superchip), per-GPU, or per-HBM-stack granularity.
- For a windowed energy number, subtract two reads of the cumulative-energy counter rather than time-integrating sampled instantaneous power; it is the more trustworthy path.
Even in an 8×A100-80GB server, GPUs are the largest slice of provisioned power but not the only one: fans and CPU are non-trivial, as an ASPLOS 2024 characterization of LLM-datacenter power made clear.
2.3 The number-one pitfall: asynchronous kernels Link to heading
CUDA kernel launches from the CPU (e.g., PyTorch) are asynchronous. So the CPU-side code that reads an NVML counter can execute before or during the wrong slice of actual GPU execution unless you explicitly call torch.cuda.synchronize() around your region of interest.
The most dangerous power-measurement bug is invisible in the data itself: an unsynchronized NVML read produces a perfectly clean-looking power trace; it is just measuring the wrong interval.
All of this is the hygiene check that must precede any measurement; the tutorial pairs it with a hands-on Colab exercise (measuring_energy.ipynb) so the discipline sticks.
3. Optimizing Performance Under a Power Cap Link to heading
Session 2. Slides.
This is the most hardware-grounded material in the tutorial: a vendor-side view built on real H100/GB300 measurements, and the one section that lands closest to home for an architect. Its thesis: you cannot optimize GPU performance without being aware of power.
3.1 Performance is a max, energy is a sum Link to heading
The organizing framework:
- Performance =
max f(x)over the critical path; only the bottleneck work sets end-to-end latency. - Energy =
sum f(x)over all work; every joule anywhere adds up. - Efficiency = Performance / Power is the bridge between them.
This asymmetry is why energy optimization is always worthwhile, while performance optimization is selective. Speeding up non-bottleneck work does nothing for latency but still burns energy.
3.2 Static vs dynamic power Link to heading
Dynamic Power = Dynamic Energy / Execution Time
Dynamic Energy = energy-per-{op,byte} × #{ops,bytes}
The consequence is a clean taxonomy of optimizations:
- Reduce work (kernel fusion, fewer ops/bytes) cuts both time and energy.
- Reduce cycles / raise utilization (e.g., compute–communication overlap) cuts time but not energy (same ops, just less idle).
3.3 Power limits are moving averages Link to heading
A GPU’s Total GPU Power (TGP) limit is enforced as a moving average over a short Electrical Design Point window (hundreds of µs to tens of ms), not an instantaneous cap. So brief bursts above TGP are allowed: GB300 is cited at 1400 W TGP with excursions up to 1820 W. Because dynamic power ≈ utilization × frequency:
- Compute-bound phases (FP8 MatMul) throttle to a lower sustained clock (topping out ~3.2–3.5 PFLOPS while pinned near TGP).
- Memory-bound phases (MemRead) sustain a higher clock without saturating TGP (reaching ~7.2–7.4 TB/s).
Utilization drives power, and power dictates sustained frequency.
And the counterintuitive corollary, validated against a real DeepSeek-V3 transformer-block power trace: interleaving low-power memory phases lets MatMul bursts run faster. A duty-cycle sweep (MatMul 100% → 10%) raised normalized MatMul throughput from 1.00 up to ~1.19 (≈19% faster). Real workloads beat pure-MatMul microbenchmarks.
3.4 Pick kernels at the power cap, not the clock Link to heading
An H100-SXM GEMM study (M×N×K = 1024×128000×1024) comparing a 128×128 tile (K1) against a 128×256 tile (K2):
| Constraint | K1 (128×128) | K2 (128×256) | Outcome |
|---|---|---|---|
| Fixed 1400 MHz | 277.9 µs @ 598 W | 291.0 µs @ 510 W | K1 4.5% faster, but K2 draws 17.3% less power |
| Fixed 700 W cap | 267.2 µs @ 1462 MHz | 267.4 µs @ 1530 MHz | effectively tied |
| Fixed 400 W cap | 370.1 µs @ 1045 MHz | 335.8 µs @ 1215 MHz | K2 is 10.2% faster |
A kernel that looks 4.5% slower unconstrained becomes 10.2% faster once power-limited, because moving fewer L2 bytes lowers dynamic energy, which buys a higher sustainable clock. Benchmark and select kernels at the deployment power cap.
3.5 Profile with real data Link to heading
FP8 MatMul power depends on input data distribution (switching activity / bit flips). Zero-valued test data reached ~4.6 PFLOPS versus ~2.6 PFLOPS on representative random data at the same power; the slide labels this ~75% more FLOPS at the same power for zero data.
Zero data underestimates power and overestimates achievable clocks and throughput. Always profile with production-representative distributions.
3.6 Serving efficiency and rack power Link to heading
- TensorRT-LLM ships built-in energy instrumentation via NVML’s
GetTotalEnergyConsumption, auto-logged throughpynvmlin NGC containers, reporting tokens/Joule and average GPU power to find Pareto-optimal serving configs. (“You can’t optimize what you can’t measure.”) - The serving throughput-vs-interactivity tradeoff carries a real cost: a ~20× gap in tokens/sec/Watt between throughput-optimized and latency-optimized operating points.
- At the rack level, GB300’s PSU-integrated energy buffering cut measured AC-input peak power on a real Megatron LLM powershelf trace from 20 kW down to 13 kW vs the prior GB200 design, letting operators provision for average rather than peak power.
4. Cutting Energy Along the Time Energy Frontier Link to heading
Session 3. Slides.
4.1 The one idea Link to heading
Every knob that saves energy (lower clock, different parallelism, different batch size) also moves job completion time. So the right object to reason about is not “minimum joules” but the time–energy Pareto frontier: the same computation, run different ways, traces a curve, and the goal is to pick the best point under whatever hard constraint actually binds: a latency SLA for serving, an accuracy target for training, a critical path in parallel training.
Energy/time decisions exist at every layer of the stack simultaneously: algorithm (MHA vs GQA), job (batch size, resolution), framework (parallelism, prefill/decode disaggregation, scheduling), cluster (instance count, GPU type), and hardware (frequency, power cap). They jointly decide where you land on the frontier.
Four systems, one per layer, make the frontier concrete. The map before the details:
| System | Venue | Layer of the stack | Main knobs | Headline result |
|---|---|---|---|---|
| Zeus | NSDI 2023 | single / recurring training jobs | batch size + GPU power limit | 15.3–75.8% less training energy |
| Perseus | SOSP 2024 | pipeline-parallel training | per-operation GPU frequency | up to 30% energy cut, ~0% slowdown |
| DynamoLLM | HPCA 2025 | inference cluster | instance count + tensor-parallel + frequency | 53% energy, 38% carbon, 61% cost |
| ML.ENERGY Benchmark | NeurIPS 2025 D&B | measurement + recommendation | config sweep on the frontier | 21–44% via config; TDP overestimates by up to 4.1× |
4.2 Zeus: batch size and power limit for training Link to heading
NSDI 2023. arXiv:2208.06102 · USENIX.
Problem. DNN training is energy-intensive (GPT-3’s run consumed 1,287 MWh, ~120 years of an average US household’s electricity), and GPUs are ~70% of training energy. Yet the field’s default playbook (max batch size, max power limit) is systematically energy-suboptimal because GPUs are not power-proportional: drawing max power gives diminishing throughput returns.
Approach. Zeus defines Energy-to-Accuracy (ETA) = TTA × AvgPower, the energy analog of time-to-accuracy, and a single user dial:
C(b, p; η) = η · ETA(b, p) + (1 − η) · MAXPOWER · TTA(b, p) # η ∈ [0,1]
The key trick is analytical decoupling: since the number of epochs to reach target accuracy doesn’t depend on the power limit p, the joint (batch, power) search collapses to a batch-size-only outer search plus a cheap just-in-time online profiler that sweeps power limits during the first epoch (doing real training work, so overhead is ~0.01–0.6%). Batch size is then a multi-armed bandit (arms = batch sizes, trials = job recurrences) solved with Gaussian Thompson sampling, which also elegantly handles concurrent job submissions because its choices are randomized.
Results.
- 15.3–75.8% end-to-end energy reduction vs the default baseline across six diverse workloads (V100).
- For DeepSpeech2/LibriSpeech, the energy-optimal config was
b=32, p=100 Wwhile the time-optimal config wasb=48, p=250 W: genuinely different operating points, not a small perturbation. - Grid search accumulates up to 72× more regret than Zeus before converging.
- On a 1.2-million-job Alibaba cluster trace, Zeus saved 7–52% energy with at most +16% time (and often less time). Zeus is the foundation of the whole ML.ENERGY line (github.com/ml-energy/zeus).
4.3 Perseus: removing energy bloat from pipeline parallelism Link to heading
SOSP 2024. arXiv:2312.06902 · ACM.
Problem. “Not all energy consumed during training directly contributes to end-to-end throughput.” That surplus is energy bloat, from two sources: intrinsic (pipeline stages are never perfectly balanced, so off-critical-path stages finish early and idle) and extrinsic (in data-parallel training, a straggler (thermal throttling of 10–50%, or I/O up to 4× compute) forces every other replica to wait).
Approach. Perseus assigns a GPU frequency to every forward/backward computation to minimize energy subject to a fixed iteration-time ceiling. Its universal prescription is T_opt = min(T*, T′), where T* is the energy-minimizing iteration time. It captures that being too fast wastes energy just like being too slow. The exact problem is NP-hard (Knapsack reduction, and NP-hard to approximate to any constant factor), so Perseus fits each op’s (time, energy) tradeoff to a smooth exponential curve and reduces “shrink the critical path by unit time τ” to an s-t min-cut / max-flow problem, tracing the entire Pareto frontier in polynomial time. At runtime a straggler event is just an instant lookup into the cached frontier.
Results.
- Up to 30% energy reduction with negligible/zero throughput loss.
- On real hardware, up to 28.5% measured savings (T5-3B, 8-stage A40 pipeline, 0% slowdown), consistently beating the prior EnvPipe heuristic.
- It captures 74% (A100) / 89% (A40) of the theoretical maximum savings, with a one-time overhead of ~13 min profiling + ~6.5 min planning (only 87 s even at an emulated 8,192-GPU scale, since only one pipeline needs solving).
Because time is held fixed at
T_opt, Perseus’s energy savings are a genuine free lunch: no accuracy, convergence, or wall-clock penalty.
4.4 DynamoLLM: energy management for inference clusters Link to heading
HPCA 2025. arXiv:2408.00741 · IEEE.
Problem. LLM inference is heterogeneous (prefill is compute-bound and TTFT-sensitive; decode is memory-bound and more latency-tolerant; requests vary wildly in length) and dynamic (diurnal load; the Coding trace peaks at 34.6× its valley). A single static “always max performance” config wastes enormous energy, but the config space (instance count × tensor-parallel degree × GPU frequency) is large and reconfiguration isn’t free.
Approach. Split traffic into 9 per-request-type pools (short/medium/long input × output, routed by a BERT-based output-length predictor) and control three knobs at three timescales so no expensive global optimizer is needed:
Results.
- Headline: 53% energy, 38% operational carbon, and 61% cost saved while meeting TTFT/TBT SLOs.
- Robust to prediction error: a 40% output-length misclassification rate costs only +13% energy and +7.3% TTFT, because the hierarchy detects and corrects mis-routes.
- A sharp economic caveat worth internalizing:
At current cloud pricing, GPU rental dominates cost by ~300×: DynamoLLM’s savings were $1362.7/hour in GPU cost (fewer servers: 40 → 24.6 average) versus only $4.4/hour in electricity. “Saving energy” and “saving money” are not the same lever today, but the carbon and power-provisioning case for energy stands on its own.
4.5 The ML.ENERGY Benchmark and Leaderboard Link to heading
NeurIPS 2025 Datasets & Benchmarks (Spotlight). arXiv:2505.06371 · poster · leaderboard.
Problem. GPUs are 50–70% of provisioned datacenter power and inference is 80–90% of ML compute demand, yet energy is usually estimated (often from GPU TDP) rather than measured, and prior efforts each fail on portability, batch-size realism, per-request granularity, or reproducibility.
Approach. Four principles: software-only GPU measurement via Zeus/NVML (no physical power meter); production stacks (vLLM, Diffusers, A100/H100) at steady state; per-response granularity (energy per complete generation, since per-token hides verbosity differences); and actionable output that sweeps configs, builds the time–energy Pareto frontier, and recommends the energy-optimal config for a user’s latency target.
Results.
- 40 model architectures across 6 tasks (LLM chat/code, VLM visual chat, text-to-image, text-to-video, image-to-video).
- TDP-based estimation overestimates measured energy by up to 4.1× (worst case CodeGemma-2B on H100), a warning for any carbon/grid-planning number built on non-measured figures.
- Recommendations: targeting 100 ms TPOT for Llama-3.1-8B lands at 77 ms and 44% less energy than the latency-minimizing config; a 5 s target for Stable Diffusion 2.1 lands at 3.63 s and 21% less energy.
- Architecture can matter more than size: Phi-3 Small (7B, GQA) can use less energy than Phi-3 Mini (3.8B, MHA) because MHA’s 3× larger KV cache caps the batch size it can amortize over. Reasoning models cost 1–2 orders of magnitude more energy per generation.
- Diffusion models draw near-TDP (compute-bound) while LLM decode draws far below TDP (memory-bandwidth-bound), so an A100→H100 upgrade barely helps LLM serving energy but substantially helps diffusion.
5. The Industry Panel Link to heading
A panel spanning NVIDIA, Google, and Meta/MLCommons.
No papers here, just a production-ops discussion structured around four prompts:
- How is the power bottleneck manifesting in infra planning and operation?
- Short- and long-term projections for power/energy issues?
- If you could magically change one thing in the ML stack, what would it be?
- The one thing you’d want ML researchers and practitioners to take away?
It closes the loop on the opening thesis: the organizations that actually provision gigawatts confirming that power, not silicon, is the planning constraint.
6. Key Takeaways Link to heading
- Treat energy as a first-class metric. Report tokens/Joule (throughput/power ≡ work/energy), not just latency and throughput.
- Measure, don’t estimate. Use RAPL for CPU/DRAM and NVML/AMDSMI for GPU/HBM; prefer differencing the cumulative-energy counter; and
cuda.synchronize()before every counter read; an unsynchronized read looks clean and is wrong. TDP-based estimates can be off by up to 4.1×. - Power, not clock, is the modern constraint. Sustained frequency is set by a moving-average power controller reacting to workload composition and input data. Benchmark and select kernels at the deployment power cap, and profile with representative data.
- Optimize on the time–energy frontier, under a real constraint. Don’t minimize joules blindly. Slack is the resource you spend: a latency SLA, an accuracy target, or straggler-induced idle time.
- The biggest wins are “free” or nearly so. Off-critical-path work (Perseus), non-power-proportional GPUs (Zeus), and workload heterogeneity (DynamoLLM) all offer double-digit energy savings without hurting the thing you care about.
- Energy ≠ dollars (yet). Today most of the cost savings come from needing fewer GPU-hours; electricity is a ~300× smaller line item. The carbon and power-provisioning case is the durable one.
7. The Complete Reading List Link to heading
Presentations (all four talks + panel):
| # | Talk | Affiliation | Slides |
|---|---|---|---|
| 0 | Opening: Why Power and Energy? And Why Now? | ML.ENERGY (U-Michigan) | |
| 1 | Accurately Benchmarking Power & Energy | ML.ENERGY (U-Michigan) | |
| 2 | Optimizing Performance Under Power Constraints | NVIDIA | |
| 3 | Optimizing Energy with Performance in Mind | ML.ENERGY (U-Michigan) | |
| 4 | Industry Panel: Power & Energy in Production ML | NVIDIA · Google · Meta/MLCommons | n/a |
Papers (the four cited systems):
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency, HPCA 2025 · arXiv:2408.00741 · IEEE
- The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization, NeurIPS 2025 D&B · arXiv:2505.06371 · poster
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training, NSDI 2023 · arXiv:2208.06102 · USENIX
- Reducing Energy Bloat in Large Model Training (Perseus), SOSP 2024 · arXiv:2312.06902 · ACM
Tools & resources: ML.ENERGY Benchmark · Leaderboard · Zeus · MLPerf Power · hands-on Colab · tutorial site · recordings