Stupid metric. It's not because a model is better performing that it necessarily requires more energy or compute.
> the NVIDIA B200 achieves 1.6× to 2.3× higher intelligence per joule than the APPLE M4 MAX across QWEN 3 and GPT-OSS model variants
The B200 = "cloud", M4 = "local".
So "cloud" does even better in energy than it does in power compared to "local". Or, to flip it, "local" is both slower and more expensive than "cloud".
Watt per Intelligence means that you have a fixed, deterministic, measure of intelligence, and you calculate how many watts it takes to get there.
If your goal is to measure which model can reach a specific outcome with the least energy possible (which is what GP says the goal is for this metric), then you cannot have a variable outcome, which is what intelligence per watt describes. As opposed to watt per intelligence, where the outcome is fixed and the numerator defines how much energy expenditure is needed to reach this fixed outcome.
> Fuel efficiency can be expressed in terms of the volume of fuel to travel a given distance, such as in litres per 100 kilometres, or through its inverse, the distance traveled per unit volume of fuel consumed, as in kilometres per litre.
If you have a car that consumes 1 liter of fuel per 100 kilometers, it's the same as saying it travels 1 kilometer per 0.01 liter of fuel. They are equivalent, describing the same relationship and proportion. It makes no difference whether the quantity is a "deterministic" or "variable" measurement. You can use either unit, depending on the aim of your calculation.
Stupid metric. It‘s not because you spend more time that you travel farther.
s/
That’s surprising, almost unbelievable, due to batching. Local is usually not batched.
You misread it. From the abstract:
> local accelerators achieve at least 1.4× lower IPW than cloud accelerators running identical models
That's "intelligence per watt". They also have IPJ, per Joule.
So, they find local is 40% "dumber" than cloud for the same power or 40% more power for the same "intelligence".
Tables 13 and 14 summarize their IPW and IPJ metrics.
But, to your actual point, I think the "local is 40% dumber per watt than cloud" message is still an understatement. And maybe this is something I failed to find in the paper but they seem to ignore the "idle baseline" costs and talks about explicitly focusing on the power consumption of just the accelerator under load.
There is a large baseline power consumption just to support the accelerator. CPUs, memory, PS losses, network, fans, general environment cooling. This "cost floor" is different for data centers and a "random local computer" and I think must be in favor of data centers which are designed and built with efficiency in mind.
Idleness should also be considered. My local GPUs at $WORK and home are idle more than they are used. Idle time energy in real world scenarios should be somehow attributed to those brief, punctuated times when LLM functions are actually active on the accelerator. Actual, local LLM usage of a GPU is brief (assuming one user per PC). Even with my heavy usage developing s/w I'd guess I heat up a GPU about one hour per day total, sometimes much less. If that is local then one must pay 23 hours of idleness for that 1 hour of "intelligence". Of course a local PC is used for other things and the idleness penalty must somehow account for that. OTOH, data centers try to maximize utilization so their idle time penalty would be much less, perhaps close to zero, by construction.
I'd argue the idleness penalty only applies if you wouldn't have otherwise had equivalent hardware. If you already have the exact same dGPU for gaming and it doubles up for inference then the only inefficiency is power consumption (both by the GPU and potentially by AC for your living space).
Conversely I think we should consider that privacy, distributed compute that you can access without an intermediary, and a more distributed power grid all provide net benefits to society at large.
This is only true if your local model is already resident in RAM / VRAM.
There could be a service that works in reverse where if someone needs a heater for a few months, they could rent a portable server (e.g. using older repurposed GPUs) with a built-in 5G modem that would run inference on LLM queries. As an incentive perhaps renting itself could be free (or you could earn money?), but you'd still have to pay your electricity bill.
I should only have to pay 1/3 of what it adds to my bill, given that it's 1/3 as efficient as a heat pump. And that coefficient is to be adjusted as outside temp changes (assuming air source heat pump; water source would have a stable COP).
Just a reminder that heat pumps can consume 300W of electricity to provide 1200W of heat.
But yes
(Tangent: I wonder why we haven't seen deployment of organic Rankin cycle generators in AI data centers, the exhaust temperature should be compatible and that could yield a 10-20% energy bill saving).
https://eu-mayors.ec.europa.eu/en/news/stockholm-sweden-heat...
I'm talking about making electricity back from the heat (using a low-temp thermodynamic cycle). It has a low yield (due to the low input temperature) but it's usually economically viable when using heat that would end up in the heavens anyway.
Let's say your incoming water temperature is 18C and you want it preheated to 50C, which is 32C degree differential, which means you'll need 32 * 1.5 * 160 = 7680 Wh, or 25 hours straight to heat the buffer tank from scratch.
You'll need to purchase a small water-to-water heat exchanger ($50), two pumps ($100 each), a power supply for said pumps, hose and/or copper pipe and fittings, and various other sundries, plus the cost of a buffer tank ($600ish), so figure all in roughly $1000, plus the cost of electricity to run the pumps.
At $0.22/kWh you're saving roughly $450 a year with this setup in foregone water heating, but because it's not 100% efficient you're spending $500 in electricity to run your GPU 24/7/365, and that's the maximum you can possibly save with the above assumptions. Scale up for more GPUs and down accordingly for less usage as you see fit.
Alternatively, use air as the heat conductor by placing the GPU laden machine in the same space as a hybrid heat pump water heater.
Because there was never a long term plan for AI data centers. It's an AI market capture and cash grab scheme that ends when local models eat their lunch.
Did the rise personal computing make data centers and super computers obsolete?
Any advance in inference that allows local models to do the job will also benefit hyperscalers. Imagine the sheer amount of compute they could throw at problems if each 32GB of VRAM was enough for frontier reasoning.
Please correct me if I am wrong.
A bit slow for agentic coding of course but fine for any chatbot use-case.
For some tasks, yes. For most of my deeper work they're not even close to my subscriptions.
> It takes less time for local model to take the first action on your task than it does for Claude to validate your login, put you into queue and start issuing the commands.
I have some decent LLM hardware here and I strongly disagree with this. Claude responds quickly. Using Fable or Opus it will deliver a working result faster than my local models because it gets there in fewer tokens. That's just how it is.
> During the winter time the GPU also doubles as a 300W in-house heater.
This is a curse in the summer. I'm feeling it right now.
Some of the items like model routing, if you do it per request instead of per session, can break down on the cloud from an energy and cost POV since one of the best things you can do for both is to maintain the KV cache which both reduces time component of energy and the quite expensive prefill energy.
I am keen on the future where we have local/cloud hybrid serving which is cache aware. I do think that could be the best use of energy resources for AI.
[0] https://arxiv.org/pdf/1911.01547
We have no dang clue what intelligence is, nor how to measure it.
my very naive question to him back then was how close we were to understanding human cognition. we were both fans of grand strategy games (though the few hundred hours of Stellaris I played vs his thousands in EU4 paled in comparison) and I was asking if it was possible to map human cognition to the same array of interdependent logical chains-of-reasoning that games like that could be boiled down to
his answer, in short, was 'we are so, so, so far away and no, that's, at best, a reductive mental model of intelligence'
I keep that conversation in mind whenever I hear about all this talk of AGI - that realistically we're so far away from actual AGI in the same way that the inventor(s) of the wheel were from a gas-powered car, and there's many paradigm shifts to go in how we even understand what the nature of intelligence is before we get there
https://pmc.ncbi.nlm.nih.gov/articles/PMC5230747/
roughly, it reviews common techniques in neuroscience, and comes to the conclusion that they would not be able to understand even simple computing platforms that we have perfect information for (and can perfectly stimulate any internal connection, can perfectly read out the values on any internal connection, etc).
- Sam Altman[0]
"Can you define intelligence?"
"Yes, it is this many moneys."
[0] https://www.businessinsider.com/sam-altman-ai-utility-electr...
Broadly in laymen's terms, no, we don't. We do have some ideas though. We have lots and lots of different tests of different facets of intelligence. How high can you count? (crows can count up to 30), how long can you remember things? (elephants can remember things from up to 50 years ago), Memory games, Reasoning. "The Science of Human Intelligence" , by Haier, Colom, and Hunt. is a good read.
>Total parameter count governs storage: at FP4, MIXTRAL-8X7B fits within 24 GB GDDR6 (Quadro RTX 6000), and GPT-OSS-120B fits within 128 GB unified memory (Apple M4 Max).