This isn't necessarily a bad or losing strategy, BTW, it just is what it is. Data paths are often very hot and don't scale linearly in many dimensions, just like power usage doesn't. Using smaller bus widths and improving bus clock in return is a very legitimate strategy to improve overall performance, it's just one of those tough tradeoffs.
Rule of thumb: take any flagship GPU and undervolt it by 30%, saving 30% power/heat dissipation, and you'll retain +90% of the performance in practice. My 3080 is nominally 320/350W but in practice I just cliff it to about 280W and it's perfectly OK in everything.
[1] Some people might even be positively surprised, since a lot of the "leaks" (bullshit rumors) were posting astronomically ridiculous numbers like 800W+ for the 4090 Ti, etc.
Every time this comes up, people forget that:
1) Flagship cards like the 3090 Ti and 4090 consume significantly more power than typical cards. You should not buy a flagship card if power is a concern.
2) You don't have to buy the flagship card. Product lineups extend all the way down to tiny single-fan cards that fit inside of small cases and silent HTPCs. It's a range of products and they all have different power consumption.
3) You don't have to run any card at full power. It's trivial to turn the power limit down to a much lower number if you want. You can drop 1/3 of the power and lose only ~10-15% of the performance in many cases due to non-linear scaling.
4) The industry has already been shipping 450W TDP cards and it's fine.
If power is an issue, buy a lower power card. The high TDP cards are an option but that doesn't mean that every card consumes 450W
I was undervolting and underclocking my Titan XP for a long time (with a separate curve for gaming) but lost the config and have been lazy to go back. I mainly did this because it has a mediocre cooler (blower design that is not nearly as good as the current ones but also not terrible) and even at low usage it was producing a lot of heat and making some noise (the rest of my system I specifically designed to have super low noise).
What you describe is accurate, the 3090 Ti produces a tremendous amount of heat under load in my experience and I would expect the same with these new cards.
Physics is telling you that you need to let the chip "shrink" when you shrink. If you keep piling on more transistors (by keeping the chip the same size) then the power goes up. That's how it works now. If you make the chip even bigger... it goes up a lot. And NVIDIA is increasing transistor count by 2.6x here.
Efficiency (perf/w) is still going up significantly, but the chip also pulls more power on top of being more efficient. If that's not acceptable for your use-case, then you'll have to accept smaller chips and slower generational progress. The 4070 and 4060 will still exist if you absolutely don't want to go above 200W. Or you can buy the bigger chips and underclock them (setting a power limit is like two clicks) and run them in the efficiency sweet spot.
But, everyone always complains about "NVIDIA won't make big chips, why are they selling small chips at a big-chip price" and now they've finally gone and done a big chip on a modern node, and people are still finding reasons to complain about it. This is what a high-density 600mm2 chip on TSMC N5P running at competitive clockrates looks like, it's a property of the node and not anything in particular that NVIDIA has done here.
AMD's chips are on the same node and will be pretty spicy too - rumors are around 400W, for a slightly smaller chip. Again, TDP being more or less a property of the chip size and the node[0], that's what you'd expect. For a given library and frequency and assuming "average" transistor activity: transistor count determines die size, and die size determines TDP. You need to improve performance-per-transistor and that's no longer easy.
[0] an oversimplification ofc but still
That's the whole point of DLSS/XeSS/Streamline/potentially a DirectX API, get more performance-per-transistor by adding an accelerator unit which "punches above its weight" in some applicable task and pushes the perf/t curve upwards. But, people have whined nonstop about that since day 1 because using inference is a conspiracy from Big GPU to sell more tensor cores, or something, I guess. Surely there is some obvious solution to TAAU sample weighting that doesn't need inference, and it's just that every academic and programmer in the field has agreed not to talk about it for the last 20 years, right?
It is kind of mind boggling how this can happen for a country with almost infinite supply of hydro power, but here we are.
Of course, Norway being Norway, since the start of September we now get a 90% refund for everything above 0.7NOK (based on the average price of each 24 hours) so for most of us will survive.
Oh, and since most power plants are owned by the public this doesn't affect budgets as badly as it could have done.
(It is even more complicated but I don't have more time now. Someone please fill in if I got something really wrong. Also, yes, this is a wild tangent but thinking I know HN I guess it will the day a little less boring for someone : )
That’s not true, my contract is for 0.41 (EUR, but that’s kinda the same right now) and I’ve seen far higher prices. Unless you are talking wholesale, which might be true but is useless for a consumer comparison.
In California with PGE we pay $0.49/Kwh on peak if you're using above "baseline"
https://www.pge.com/pge_global/common/pdfs/rate-plans/how-ra...
Our power bill for a <2000sqft single-family home is about $800 per month right now, and CA + PGE keep trying to do things like this to prevent solar from being worth it: https://www.solarpowerworldonline.com/2022/08/cpuc-proposed-...
At that level you're really in the territory where people will trip breakers in their home. This is "you have to hire an electrician and rewire part of your house" level of power consumption, which doesn't seem sustainable? Or at least "don't turn the stove on at the same time", which seems like an awfully disruptive way to ask people to live.
Heck, at that level for homes/cities with older electric service at lower amperages, the computer is now a significant fraction of the total service capacity of your connection to the city grid.
What's the response then? "Sorry you're not living in a new build home, forget about gaming"?
For me the wall-power isn't a huge issue but now you have to think about air conditioning. 1-2KW dissipating in a closed room heats the room up fast. This is already a problem with the RTX 2000-series that I'm running now and it seems like it'll be even worse.
"Glad you want to game, now please evaluate your home's electrical and HVAC systems to ensure they are compatible with a giant box pumping 2KW of power and converting it directly to heat"
While reading this I felt a tinge of nostalgia for back when I had to ask my mother to stay off of the phone while I was online.
Maybe that's the genius of their plan? The chip designers continue to put out better & hotter processors and expect someone else to figure out cooling. Thermal design failures are not attributed to Intel and Nvidia, instead to OEMs and consumers.
Offering a more powerful GPU/CPU for people and use-cases that demand it is not a problem. Most games these days run well even on mid-tier cards of the 20 series, that's two generations ago. So people have all the choice in the world.
Just because a 1000 bhp LaFerrari exists in the market, doesn't mean you cannot buy a ~100 bhp Corolla. People who buy the Ferrari will need to take special care of their car to get the most out of it. That's a pleasure in itself for some.
At the high-end, Nvidia is making more powerful cards (HP) with the same fuel efficiency as before. Maybe they are making more efficient cards at the lower end but those don’t get attention on HN it seems.
No, perf/w is still climbing every generation. It'll climb hugely this generation too.
NVIDIA's marketing number is 2x perf at the same wattage as last gen, obviously needs to be validated by third party testing but they're shrinking two nodes at once here, it's a bigger node jump than Pascal was.
People have been primed to react negatively by a couple of twitter rumormongers like kopite7kimi. Yeah 4090 has a high TGP, but it's more like 400W TGP (450W TBP) and not the 800W TGP / 900W TBP that he was shouting from the rooftops about. There is still a huge gain in efficiency here, it's just also a very large (and expensive) chip on top of the higher efficiency.
But people don’t discard their incorrect mental frames when the information used to build them is falsified. Actually they just tend to retrench and dig deeper (“450W is still too much even if it is 2x perf/w!”). Ada is a power hog, “everybody knows it”.
https://videocardz.com/newz/flagship-nvidia-ada-sku-might-fe...
People said they wanted a balls-out big chip on a competitive node, that's exactly what the 4090 is, it's highly efficient and very large, this is much closer to the cutting edge of what the tech can deliver than Samsung 8nm junk or little baby 450mm2 dies. Now people are complaining about the TDP and the cost, even if it's more efficient that's not good enough. Unfortunately the TDP is a matter of physics, thermal density is going up every generation and a big chip on a modern node pulls a lot of power.
There are lower models in the stack too and those will maintain that efficiency benefit. It's a larger-than-pascal node shrink, efficiency is going up a lot here.
This is what Intel did for a while with their "tick tock" strategy https://en.wikipedia.org/wiki/Tick%E2%80%93tock_model
You'll be hard pressed to find a generation where power efficiency actually decreased. That didn't happen with the RTX 30xx series, for example. Even though power skyrocketed, performance increased by more than power did. Meaning efficiency still increased. That's been true for every generation I can think of, certainly all the more recent ones anyway. Maybe AMD had a few like that where they just re-rev'd the same architecture multiple times in a row.
Some people are probably in the target audience for the 4090. Others may prefer the 4080 models, which have a slightly lower TDP than the 3080 models but still get a nice performance boost from much higher clock rates.
With watercooling and a big rad, it can get even quieter.
the bigger issue is where all that additional heat is going and how. I'll be paying close attention to the SPL metrics for this generation.