LLM-generated.

TL;DR Link to heading

Having spent a career reasoning about where the real limits of a machine sit, I read the ML.ENERGY tutorial at NeurIPS 2025, “Energy and Power as First-Class ML Design Metrics”, as a marker of a shift the architecture community has seen coming: electrical power, not chips or algorithms, is now the binding constraint on AI scaling. Once you accept that, the engineering agenda writes itself: measure GPU power and energy without fooling yourself, optimize performance when the real budget is watts rather than clock cycles, and reduce energy by moving a workload along its time–energy Pareto frontier instead of blindly minimizing joules. What follows is that argument and the four systems that discharge it (Zeus, Perseus, DynamoLLM, and the ML.ENERGY Benchmark), read the way an architect reads a design: mechanism first, then the numbers that decide whether it matters.

What Power and Energy as First-Class ML Metrics means?

  1. Reframes energy as a first-class computing resource, on par with FLOPs and memory bandwidth: something you measure, optimize under, and reduce.
  2. Nails down the correct way to measure ML power/energy (RAPL for CPU/DRAM, NVML/AMDSMI for GPU/HBM) and the single most common bug that silently corrupts results: unsynchronized CUDA.
  3. Shows that on modern GPUs performance and power are inseparable: sustained clocks are set by a moving-average power controller reacting to workload composition and even input data.
  4. Establishes the time–energy Pareto frontier as the object to reason about, and surveys four systems that each pick a better point on it at a different layer of the stack.
  5. Grounds all of it in production reality via a panel with the vendors and hyperscalers who live inside these power budgets.

Table of contents


1. Why Now: Power Is the Binding Constraint Link to heading

Opening framing. Slides.

1.1 The problem Link to heading

The opening is diagnosis rather than a research result, and the diagnosis is correct. AI has reached unprecedented scale along three orthogonal axes: Large (frontier models such as Kimi-K2 at ~1T parameters, with closed models suspected larger), Numerous (thousands of specialized models), and Everywhere (deployed worldwide, serving millions). That scale drives ballooning demand for accelerators and gigawatt-class datacenters. And the thing you actually cannot buy quickly is electrical power.

1.2 The evidence: a news timeline Link to heading

The case is a chronological wall of 2024–2025 headlines, and it is hard to argue with:

DateReported movePower/scale signal
Jan 2024Meta buying 350,000 H100 GPUs~600,000 H100-equivalent total compute
Apr 2025Meta plans a datacenter “the size of Manhattan”“hundreds of billions” in AI spend
May 2025xAI’s Memphis siteproposed 1.56 GW gas plant, up to 90 turbines
Sep 2025OpenAI + Oracle + SoftBank expand Stargate$500B, 10 GW targeted by end of 2025

A wry aside relabels “350,000 [GPUs]” as “350,000 refrigerators”, a fair jab at how casually we report enormous power-hungry purchases as bare device counts, with no sense of the load behind them.

The core structural tension:

Next-generation frontier models “were due yesterday,” but large-scale power procurement (siting, permitting, approval, construction) takes years. That cadence mismatch is the bottleneck the rest of the tutorial attacks.

flowchart TD A["AI scale grows: larger, more numerous, everywhere"] --> B["Compute demand explodes"] B --> C["Gigawatt-scale datacenter buildouts"] C --> D["Power procurement is the bottleneck"] D --> E["Siting, permitting, construction take years"] E --> F["Energy becomes a first-class computing resource"] F --> G["Session 1: measure it"] F --> H["Session 2: optimize performance under it"] F --> I["Session 3: reduce it"]

1.3 Why it matters Link to heading

Three constituencies are on the hook: AI builders/providers (power capacity gates both research and serving capacity), hardware vendors (“no one can buy new hardware without the energy to run them”), and users/the public (grid stability and electricity prices). The conclusion I draw is straightforward: GPU power draw, energy-aware scheduling, and measurement tooling now belong on the same design dashboard as raw FLOPs.


2. Measuring Power and Energy Correctly Link to heading

Session 1. Slides.

2.1 Power vs energy Link to heading

  • Power (watts, J/s) is the instantaneous rate delivered by the grid, bounded by a hardware/datacenter power ceiling. It is the constraint.
  • Energy (joules or Wh) is the time-integral of power: the area under the power-vs-time curve. It is the cost.

Conflating the two is the root of many bad benchmark claims. Power is the constraint; energy is what you are billed for.

Because datacenter capacity is itself provisioned in power units (e.g., 1 GW), the natural efficiency metric collapses cleanly:

throughput / power = (work/s) / (J/s) = work / energy

which is why the field increasingly reports tokens per joule rather than raw watts.

2.2 The tools: RAPL, NVML, AMDSMI Link to heading

Accurate system-level accounting means stitching two instrumentation families together, using vendor counters rather than external meters for the compute components:

ComponentToolWhat it exposes
CPU packageIntel RAPL (via sysfs)energy/power; supported on most Intel and AMD
DRAMIntel RAPL (DRAM domain)energy; only some Intel CPUs
NVIDIA GPU / HBMNVML (C++ + Python bindings)power draw (1-s average or instantaneous) and cumulative energy since driver load
AMD GPU / HBMAMDSMIthe same counter families as NVML

Two subtleties matter in practice:

  • NVML reports power at whole-module (e.g., a Grace Hopper superchip), per-GPU, or per-HBM-stack granularity.
  • For a windowed energy number, subtract two reads of the cumulative-energy counter rather than time-integrating sampled instantaneous power; it is the more trustworthy path.

Even in an 8×A100-80GB server, GPUs are the largest slice of provisioned power but not the only one: fans and CPU are non-trivial, as an ASPLOS 2024 characterization of LLM-datacenter power made clear.

2.3 The number-one pitfall: asynchronous kernels Link to heading

CUDA kernel launches from the CPU (e.g., PyTorch) are asynchronous. So the CPU-side code that reads an NVML counter can execute before or during the wrong slice of actual GPU execution unless you explicitly call torch.cuda.synchronize() around your region of interest.

The most dangerous power-measurement bug is invisible in the data itself: an unsynchronized NVML read produces a perfectly clean-looking power trace; it is just measuring the wrong interval.

flowchart TD CPU["CPU and DRAM"] --> RAPL["Intel RAPL via sysfs"] GPU["GPU and HBM"] --> LIB["NVML or AMDSMI"] LIB --> PWR["Power draw: 1s average or instantaneous"] LIB --> ENE["Cumulative energy since driver load"] ENE --> DIFF["Subtract two reads to get a window"] APP["PyTorch launches kernels asynchronously"] --> KERNEL["GPU runs the kernel later"] KERNEL --> READ["NVML power or energy read"] APP --> WRONG["Without cuda synchronize, the read spans the wrong interval"]

All of this is the hygiene check that must precede any measurement; the tutorial pairs it with a hands-on Colab exercise (measuring_energy.ipynb) so the discipline sticks.


3. Optimizing Performance Under a Power Cap Link to heading

Session 2. Slides.

This is the most hardware-grounded material in the tutorial: a vendor-side view built on real H100/GB300 measurements, and the one section that lands closest to home for an architect. Its thesis: you cannot optimize GPU performance without being aware of power.

3.1 Performance is a max, energy is a sum Link to heading

The organizing framework:

  • Performance = max f(x) over the critical path; only the bottleneck work sets end-to-end latency.
  • Energy = sum f(x) over all work; every joule anywhere adds up.
  • Efficiency = Performance / Power is the bridge between them.

This asymmetry is why energy optimization is always worthwhile, while performance optimization is selective. Speeding up non-bottleneck work does nothing for latency but still burns energy.

3.2 Static vs dynamic power Link to heading

Dynamic Power  = Dynamic Energy / Execution Time
Dynamic Energy = energy-per-{op,byte} × #{ops,bytes}

The consequence is a clean taxonomy of optimizations:

  • Reduce work (kernel fusion, fewer ops/bytes) cuts both time and energy.
  • Reduce cycles / raise utilization (e.g., compute–communication overlap) cuts time but not energy (same ops, just less idle).

3.3 Power limits are moving averages Link to heading

A GPU’s Total GPU Power (TGP) limit is enforced as a moving average over a short Electrical Design Point window (hundreds of µs to tens of ms), not an instantaneous cap. So brief bursts above TGP are allowed: GB300 is cited at 1400 W TGP with excursions up to 1820 W. Because dynamic power ≈ utilization × frequency:

  • Compute-bound phases (FP8 MatMul) throttle to a lower sustained clock (topping out ~3.2–3.5 PFLOPS while pinned near TGP).
  • Memory-bound phases (MemRead) sustain a higher clock without saturating TGP (reaching ~7.2–7.4 TB/s).

Utilization drives power, and power dictates sustained frequency.

And the counterintuitive corollary, validated against a real DeepSeek-V3 transformer-block power trace: interleaving low-power memory phases lets MatMul bursts run faster. A duty-cycle sweep (MatMul 100% → 10%) raised normalized MatMul throughput from 1.00 up to ~1.19 (≈19% faster). Real workloads beat pure-MatMul microbenchmarks.

flowchart TD MM["MatMul phase: high utilization"] --> P["Instantaneous GPU power"] MR["MemRead phase: low utilization"] --> P P --> AVG["Moving-average power window"] AVG --> Q{Average vs TGP} Q -->|near TGP| LO["Throttle to a lower clock"] Q -->|below TGP| HI["Raise clock toward max"] LO --> PERF["Sustained performance"] HI --> PERF

3.4 Pick kernels at the power cap, not the clock Link to heading

An H100-SXM GEMM study (M×N×K = 1024×128000×1024) comparing a 128×128 tile (K1) against a 128×256 tile (K2):

ConstraintK1 (128×128)K2 (128×256)Outcome
Fixed 1400 MHz277.9 µs @ 598 W291.0 µs @ 510 WK1 4.5% faster, but K2 draws 17.3% less power
Fixed 700 W cap267.2 µs @ 1462 MHz267.4 µs @ 1530 MHzeffectively tied
Fixed 400 W cap370.1 µs @ 1045 MHz335.8 µs @ 1215 MHzK2 is 10.2% faster

A kernel that looks 4.5% slower unconstrained becomes 10.2% faster once power-limited, because moving fewer L2 bytes lowers dynamic energy, which buys a higher sustainable clock. Benchmark and select kernels at the deployment power cap.

3.5 Profile with real data Link to heading

FP8 MatMul power depends on input data distribution (switching activity / bit flips). Zero-valued test data reached ~4.6 PFLOPS versus ~2.6 PFLOPS on representative random data at the same power; the slide labels this ~75% more FLOPS at the same power for zero data.

Zero data underestimates power and overestimates achievable clocks and throughput. Always profile with production-representative distributions.

3.6 Serving efficiency and rack power Link to heading

  • TensorRT-LLM ships built-in energy instrumentation via NVML’s GetTotalEnergyConsumption, auto-logged through pynvml in NGC containers, reporting tokens/Joule and average GPU power to find Pareto-optimal serving configs. (“You can’t optimize what you can’t measure.”)
  • The serving throughput-vs-interactivity tradeoff carries a real cost: a ~20× gap in tokens/sec/Watt between throughput-optimized and latency-optimized operating points.
  • At the rack level, GB300’s PSU-integrated energy buffering cut measured AC-input peak power on a real Megatron LLM powershelf trace from 20 kW down to 13 kW vs the prior GB200 design, letting operators provision for average rather than peak power.

4. Cutting Energy Along the Time Energy Frontier Link to heading

Session 3. Slides.

4.1 The one idea Link to heading

Every knob that saves energy (lower clock, different parallelism, different batch size) also moves job completion time. So the right object to reason about is not “minimum joules” but the time–energy Pareto frontier: the same computation, run different ways, traces a curve, and the goal is to pick the best point under whatever hard constraint actually binds: a latency SLA for serving, an accuracy target for training, a critical path in parallel training.

Energy/time decisions exist at every layer of the stack simultaneously: algorithm (MHA vs GQA), job (batch size, resolution), framework (parallelism, prefill/decode disaggregation, scheduling), cluster (instance count, GPU type), and hardware (frequency, power cap). They jointly decide where you land on the frontier.

flowchart TD ALGO["Algorithm: MHA vs GQA, dense vs sparse"] --> F["Time-energy Pareto frontier"] JOB["Job: batch size, sequence length, resolution"] --> F FW["Framework: parallelism, batching, PD disaggregation"] --> F CL["Cluster: instance count, GPU type"] --> F HW["Hardware: frequency scaling, power capping"] --> F F --> OPT["Pick the Pareto point under a latency or accuracy target"] OPT --> Z["Zeus: 15 to 76 percent training energy"] OPT --> PE["Perseus: up to 30 percent energy bloat"] OPT --> DY["DynamoLLM: up to 53 percent serving energy"] OPT --> BM["ML.ENERGY Benchmark: measure and recommend"]

Four systems, one per layer, make the frontier concrete. The map before the details:

SystemVenueLayer of the stackMain knobsHeadline result
ZeusNSDI 2023single / recurring training jobsbatch size + GPU power limit15.3–75.8% less training energy
PerseusSOSP 2024pipeline-parallel trainingper-operation GPU frequencyup to 30% energy cut, ~0% slowdown
DynamoLLMHPCA 2025inference clusterinstance count + tensor-parallel + frequency53% energy, 38% carbon, 61% cost
ML.ENERGY BenchmarkNeurIPS 2025 D&Bmeasurement + recommendationconfig sweep on the frontier21–44% via config; TDP overestimates by up to 4.1×

4.2 Zeus: batch size and power limit for training Link to heading

NSDI 2023. arXiv:2208.06102 · USENIX.

Problem. DNN training is energy-intensive (GPT-3’s run consumed 1,287 MWh, ~120 years of an average US household’s electricity), and GPUs are ~70% of training energy. Yet the field’s default playbook (max batch size, max power limit) is systematically energy-suboptimal because GPUs are not power-proportional: drawing max power gives diminishing throughput returns.

Approach. Zeus defines Energy-to-Accuracy (ETA) = TTA × AvgPower, the energy analog of time-to-accuracy, and a single user dial:

C(b, p; η) = η · ETA(b, p) + (1 − η) · MAXPOWER · TTA(b, p)      # η ∈ [0,1]

The key trick is analytical decoupling: since the number of epochs to reach target accuracy doesn’t depend on the power limit p, the joint (batch, power) search collapses to a batch-size-only outer search plus a cheap just-in-time online profiler that sweeps power limits during the first epoch (doing real training work, so overhead is ~0.01–0.6%). Batch size is then a multi-armed bandit (arms = batch sizes, trials = job recurrences) solved with Gaussian Thompson sampling, which also elegantly handles concurrent job submissions because its choices are randomized.

flowchart LR U["Recurring training job"] --> MAB["Thompson-sampling bandit picks batch size"] MAB --> JIT["JIT profiler picks the optimal power limit"] JIT --> RUN["Train with early stopping"] RUN --> FB["Observe time, power, energy"] FB --> UP["Update the bandit belief for that arm"] UP --> MAB

Results.

  • 15.3–75.8% end-to-end energy reduction vs the default baseline across six diverse workloads (V100).
  • For DeepSpeech2/LibriSpeech, the energy-optimal config was b=32, p=100 W while the time-optimal config was b=48, p=250 W: genuinely different operating points, not a small perturbation.
  • Grid search accumulates up to 72× more regret than Zeus before converging.
  • On a 1.2-million-job Alibaba cluster trace, Zeus saved 7–52% energy with at most +16% time (and often less time). Zeus is the foundation of the whole ML.ENERGY line (github.com/ml-energy/zeus).

4.3 Perseus: removing energy bloat from pipeline parallelism Link to heading

SOSP 2024. arXiv:2312.06902 · ACM.

Problem. “Not all energy consumed during training directly contributes to end-to-end throughput.” That surplus is energy bloat, from two sources: intrinsic (pipeline stages are never perfectly balanced, so off-critical-path stages finish early and idle) and extrinsic (in data-parallel training, a straggler (thermal throttling of 10–50%, or I/O up to 4× compute) forces every other replica to wait).

Approach. Perseus assigns a GPU frequency to every forward/backward computation to minimize energy subject to a fixed iteration-time ceiling. Its universal prescription is T_opt = min(T*, T′), where T* is the energy-minimizing iteration time. It captures that being too fast wastes energy just like being too slow. The exact problem is NP-hard (Knapsack reduction, and NP-hard to approximate to any constant factor), so Perseus fits each op’s (time, energy) tradeoff to a smooth exponential curve and reduces “shrink the critical path by unit time τ” to an s-t min-cut / max-flow problem, tracing the entire Pareto frontier in polynomial time. At runtime a straggler event is just an instant lookup into the cached frontier.

flowchart TD DAG["Iteration as a DAG of forward and backward ops"] --> PROF["In-vivo profiler sweeps frequencies"] PROF --> FIT["Fit an exponential time-energy curve per op"] FIT --> CUT["Trace the frontier via min-cut max-flow"] CUT --> CACHE["Cache schedules indexed by iteration time"] STR["Infra reports a straggler slowdown"] --> TOPT["Compute Topt = min of T-star and T-prime"] TOPT --> CACHE CACHE --> CTRL["Async NVML frequency controller"] CTRL --> OUT["Same iteration time, less energy"]

Results.

  • Up to 30% energy reduction with negligible/zero throughput loss.
  • On real hardware, up to 28.5% measured savings (T5-3B, 8-stage A40 pipeline, 0% slowdown), consistently beating the prior EnvPipe heuristic.
  • It captures 74% (A100) / 89% (A40) of the theoretical maximum savings, with a one-time overhead of ~13 min profiling + ~6.5 min planning (only 87 s even at an emulated 8,192-GPU scale, since only one pipeline needs solving).

Because time is held fixed at T_opt, Perseus’s energy savings are a genuine free lunch: no accuracy, convergence, or wall-clock penalty.

4.4 DynamoLLM: energy management for inference clusters Link to heading

HPCA 2025. arXiv:2408.00741 · IEEE.

Problem. LLM inference is heterogeneous (prefill is compute-bound and TTFT-sensitive; decode is memory-bound and more latency-tolerant; requests vary wildly in length) and dynamic (diurnal load; the Coding trace peaks at 34.6× its valley). A single static “always max performance” config wastes enormous energy, but the config space (instance count × tensor-parallel degree × GPU frequency) is large and reconfiguration isn’t free.

Approach. Split traffic into 9 per-request-type pools (short/medium/long input × output, routed by a BERT-based output-length predictor) and control three knobs at three timescales so no expensive global optimizer is needed:

flowchart TD REQ["Incoming request"] --> PRED["Output-length predictor"] PRED --> CM["Cluster manager: instance count, ~30 min"] CM --> PM["Pool manager: tensor-parallel degree, ~5 min"] PM --> IM["Instance manager: GPU frequency, ~5 s"] IM --> GPU["H100 pool serves the request under its SLO"] PROF["Offline energy-performance profiles"] --> CM PROF --> PM PROF --> IM

Results.

  • Headline: 53% energy, 38% operational carbon, and 61% cost saved while meeting TTFT/TBT SLOs.
  • Robust to prediction error: a 40% output-length misclassification rate costs only +13% energy and +7.3% TTFT, because the hierarchy detects and corrects mis-routes.
  • A sharp economic caveat worth internalizing:

At current cloud pricing, GPU rental dominates cost by ~300×: DynamoLLM’s savings were $1362.7/hour in GPU cost (fewer servers: 40 → 24.6 average) versus only $4.4/hour in electricity. “Saving energy” and “saving money” are not the same lever today, but the carbon and power-provisioning case for energy stands on its own.

4.5 The ML.ENERGY Benchmark and Leaderboard Link to heading

NeurIPS 2025 Datasets & Benchmarks (Spotlight). arXiv:2505.06371 · poster · leaderboard.

Problem. GPUs are 50–70% of provisioned datacenter power and inference is 80–90% of ML compute demand, yet energy is usually estimated (often from GPU TDP) rather than measured, and prior efforts each fail on portability, batch-size realism, per-request granularity, or reproducibility.

Approach. Four principles: software-only GPU measurement via Zeus/NVML (no physical power meter); production stacks (vLLM, Diffusers, A100/H100) at steady state; per-response granularity (energy per complete generation, since per-token hides verbosity differences); and actionable output that sweeps configs, builds the time–energy Pareto frontier, and recommends the energy-optimal config for a user’s latency target.

Results.

  • 40 model architectures across 6 tasks (LLM chat/code, VLM visual chat, text-to-image, text-to-video, image-to-video).
  • TDP-based estimation overestimates measured energy by up to 4.1× (worst case CodeGemma-2B on H100), a warning for any carbon/grid-planning number built on non-measured figures.
  • Recommendations: targeting 100 ms TPOT for Llama-3.1-8B lands at 77 ms and 44% less energy than the latency-minimizing config; a 5 s target for Stable Diffusion 2.1 lands at 3.63 s and 21% less energy.
  • Architecture can matter more than size: Phi-3 Small (7B, GQA) can use less energy than Phi-3 Mini (3.8B, MHA) because MHA’s 3× larger KV cache caps the batch size it can amortize over. Reasoning models cost 1–2 orders of magnitude more energy per generation.
  • Diffusion models draw near-TDP (compute-bound) while LLM decode draws far below TDP (memory-bandwidth-bound), so an A100→H100 upgrade barely helps LLM serving energy but substantially helps diffusion.

5. The Industry Panel Link to heading

A panel spanning NVIDIA, Google, and Meta/MLCommons.

No papers here, just a production-ops discussion structured around four prompts:

  • How is the power bottleneck manifesting in infra planning and operation?
  • Short- and long-term projections for power/energy issues?
  • If you could magically change one thing in the ML stack, what would it be?
  • The one thing you’d want ML researchers and practitioners to take away?

It closes the loop on the opening thesis: the organizations that actually provision gigawatts confirming that power, not silicon, is the planning constraint.


6. Key Takeaways Link to heading

  1. Treat energy as a first-class metric. Report tokens/Joule (throughput/power ≡ work/energy), not just latency and throughput.
  2. Measure, don’t estimate. Use RAPL for CPU/DRAM and NVML/AMDSMI for GPU/HBM; prefer differencing the cumulative-energy counter; and cuda.synchronize() before every counter read; an unsynchronized read looks clean and is wrong. TDP-based estimates can be off by up to 4.1×.
  3. Power, not clock, is the modern constraint. Sustained frequency is set by a moving-average power controller reacting to workload composition and input data. Benchmark and select kernels at the deployment power cap, and profile with representative data.
  4. Optimize on the time–energy frontier, under a real constraint. Don’t minimize joules blindly. Slack is the resource you spend: a latency SLA, an accuracy target, or straggler-induced idle time.
  5. The biggest wins are “free” or nearly so. Off-critical-path work (Perseus), non-power-proportional GPUs (Zeus), and workload heterogeneity (DynamoLLM) all offer double-digit energy savings without hurting the thing you care about.
  6. Energy ≠ dollars (yet). Today most of the cost savings come from needing fewer GPU-hours; electricity is a ~300× smaller line item. The carbon and power-provisioning case is the durable one.

7. The Complete Reading List Link to heading

Presentations (all four talks + panel):

#TalkAffiliationSlides
0Opening: Why Power and Energy? And Why Now?ML.ENERGY (U-Michigan)PDF
1Accurately Benchmarking Power & EnergyML.ENERGY (U-Michigan)PDF
2Optimizing Performance Under Power ConstraintsNVIDIAPDF
3Optimizing Energy with Performance in MindML.ENERGY (U-Michigan)PDF
4Industry Panel: Power & Energy in Production MLNVIDIA · Google · Meta/MLCommonsn/a

Papers (the four cited systems):

  1. DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency, HPCA 2025 · arXiv:2408.00741 · IEEE
  2. The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization, NeurIPS 2025 D&B · arXiv:2505.06371 · poster
  3. Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training, NSDI 2023 · arXiv:2208.06102 · USENIX
  4. Reducing Energy Bloat in Large Model Training (Perseus), SOSP 2024 · arXiv:2312.06902 · ACM

Tools & resources: ML.ENERGY Benchmark · Leaderboard · Zeus · MLPerf Power · hands-on Colab · tutorial site · recordings