<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Llm Inference on Cagri</title><link>https://cagri.dev/tags/llm-inference/</link><description>Recent content in Llm Inference on Cagri</description><generator>Hugo</generator><language>en</language><lastBuildDate>Wed, 09 Sep 2026 09:00:00 -0700</lastBuildDate><atom:link href="https://cagri.dev/tags/llm-inference/index.xml" rel="self" type="application/rss+xml"/><item><title>Power and Energy as First-Class ML Metrics</title><link>https://cagri.dev/posts/mlenergy-neurips25-tutorial/</link><pubDate>Wed, 09 Sep 2026 09:00:00 -0700</pubDate><guid>https://cagri.dev/posts/mlenergy-neurips25-tutorial/</guid><description>&lt;blockquote>
&lt;p>LLM-generated.&lt;/p>&lt;/blockquote>
&lt;h2 id="tldr">
 TL;DR
 &lt;a class="heading-link" href="#tldr">
 &lt;i class="fa-solid fa-link" aria-hidden="true" title="Link to heading">&lt;/i>
 &lt;span class="sr-only">Link to heading&lt;/span>
 &lt;/a>
&lt;/h2>
&lt;p>Having spent a career reasoning about where the real limits of a machine sit, I read the &lt;a href="https://ml.energy/tutorials/neurips25/" class="external-link" target="_blank" rel="noopener">ML.ENERGY tutorial at NeurIPS 2025&lt;/a>, &lt;em>&amp;ldquo;Energy and Power as First-Class ML Design Metrics&amp;rdquo;&lt;/em>, as a marker of a shift the architecture community has seen coming: &lt;strong>electrical power, not chips or algorithms, is now the binding constraint on AI scaling.&lt;/strong> Once you accept that, the engineering agenda writes itself: &lt;em>measure&lt;/em> GPU power and energy without fooling yourself, &lt;em>optimize performance&lt;/em> when the real budget is watts rather than clock cycles, and &lt;em>reduce energy&lt;/em> by moving a workload along its &lt;strong>time–energy Pareto frontier&lt;/strong> instead of blindly minimizing joules. What follows is that argument and the four systems that discharge it (&lt;code>Zeus&lt;/code>, &lt;code>Perseus&lt;/code>, &lt;code>DynamoLLM&lt;/code>, and the &lt;code>ML.ENERGY Benchmark&lt;/code>), read the way an architect reads a design: mechanism first, then the numbers that decide whether it matters.&lt;/p></description></item></channel></rss>