Nvidia investigating reports of RTX 4090 power cables burning or melting
theverge.com
theverge.com
I believe it is entirely possible to implement a solid and safe 12VHPWR connector, and its reputation is being destroyed by a shoddy implementation... but that's very much nVidia shooting themselves in the foot.
I don't personally like EVGA, had a bad experience trying (and failing) to get warranty replacement for a defective PSU, but I can certainly follow GP's thinking about how EVGA's involvement with this launch might have made a difference.
I'm sure Nvidia will get it sorted in time, but they're trying to do a lot right now, and it's noticeably bumpy.
The linked article gives very good insights in what’s going on.
Unfortunately NVidia has really screwed the pooch with this as this dangerous adaptor is the first encounter most people will have with the 12VHPWR form factor
However, you're missing the point. The spec should be designed such that even "non-reputable" manufacturers with sloppier manufacturing methods wouldn't create a fire hazard. These connectors are extra-delicate even though they're supposed to carry higher loads. They will be less robust in practice, for very little gain. That's a poor tradeoff, which makes it poor design.
These new adaptors are included with products Nvidia themselves ship, with themselves as OEM. They know this standard is not common, and they know they need to ship adaptors with them, If they can’t guarantee the quality of it’s manufacturing, then it says there’s possible more worrying things going on behind the scenes.
70 Celsius is NUTS for that usage, not ‘oh, that’s ok then’
It’s well outside normal expected operating temperatures for wire.
This data point isn't exhaustive, but it does indicate that the actual safety margin is not as tight as is assumed from this news cycle.
That brand new connectors and wire that are custom specced wouldn’t catch on fire (but come close!) at only 2x the power draw means it’s very likely any damage to a connector, wire, etc. especially with normal wire and connectors likely would cause damage and a fire, at much less than that power draw.
Which is what we’ve been getting reports on.
And that is just from minor physical damage it seems, ignoring corrosion or fraying wires over time which is usually the bigger problem.
The truth is the plug is more demanding and expensive than lower density options and it's dangerous when corners are cut. This is always true in power electronics, but the risk surface area is higher with a new and demanding connector.
Even with connectors and cables without those early problems, there are often longer term problems that show up over time - cables fraying at the connectors due to movement, plugs and connectors building up corrosion (and hence having higher resistance), connections loosening or getting bent, etc.
The link you posted, and your later statement seems to support the same view. I'm just repeating it so the folks dismissing this as just an issue with the badly designed adapters Nvidia distributed (in some cases?). It's overall a connector and spec that's on the edge of the performance envelope with some 'obvious' types of field failure modes that don't seem to be properly addressed.
Honestly, I'm a bit shocked they went with 'mountains of parallel power feeds' solution instead of... I don't know, doing 6 gauge stranded, which seems like the obvious choice to me? Parallel power feeds like this are always the source of endless fussing and headaches due to exactly the problems we're seeing. Trying to save a couple cents by using more of the same materials they have on hand I'm guessing?
Yes, if "the right materials" are used and the cables are treated "properly" (not bent out of spec, not too many connection cycles) then it's "OK", but those shouldn't be requirements designed into something that cost-conscious consumers are going to be assembling.
Does the old way of building PCIe cards still make sense when the GPU now draws more power than all of the other system components combined? GPU sag (i.e. the card weighs so damn much) is such an issue that one manufacturer actually put a level in their card.[0] The 4090 probably needs three slots just to keep it attached to the motherboard.
When someone buys a desktop these days, it's very often because they need a discrete GPU. You buy a dedicated desktop GPU because the thermal envelope of a laptop can't support a card that would draw 600W - even if it were possible to shrink the card into a laptop (or make the laptop big enough to fit the 600W card). The top of the line Intel processors draw in the neighborhood of 150W but that doesn't necessitate going to a different form factor. Future motherboard and south bridge designs should mount the GPU independently of the other components.
[0] https://www.tomshardware.com/news/geforce-rtx-4090-has-a-spi...
In your opinion, if you could force power supply makers to add another voltage rail, which voltage would you pick?
Solving for the problems of higher current, lower voltage DC we use in ATX format supplies today requires higher gauge copper, but the cable run between the GPU and PSU is so shot that the cost of extra copper could very well be less than the increased cost of a higher voltage supply.
where is this typical? in my experience, typical is 200A service for residential. is mains service amps dependent on provider/region/etc?
Often, those transformers on the pole have some wiggle room in what they can deliver, and just because you upgraded from 100a to 200a service doesn’t mean you’ll actually see a real world change in usage. You usually don’t. Utility companies can over provision that capacity because not everyone on a street switches on their microwave at the same time, etc.
I would bet the cause coming down to minor manufacturing defects and testing that was “too good” at plugging things in not accounting for sloppy end user assembly.
Combined with good old inadequate safety margins.
> things you never thought of worrying about.
In my experience, it's nearly always the connectors that you suspect first.
It is weird that in their push to all time high TDP they didn’t spend more time thinking about it. That said, cutting edge often means bleeding edge, something something.
For example, the RTX 4090 has 16,384 cores, so you have to divide 50+ billion by that number. Then inside every core you have lots of similar repetition, etc.
The biggest challenge is to create a die that size with little to no defects, which is of course the foundry's achievement.
A modern great CPU or GPU chip is not that simple to design. It is not just an ALU copy pasted thousands of times. Otherwise there would be more competition.
Not really, the cores are really similar units.
> A modern great CPU or GPU chip is not that simple to design.
Of course not, for starters there's not just the shader cores (a 4090 also has 512 TMUs, 176 ROPs, 128 RTX, and 512 tensor cores), you need to design all of these before you can replicate them, then you need to design the common components.
But the point was that the 50 billion transistors were not designed for individually, there is a large amount of block reuse. Your analogy would hold if windows inlined everything.
NVidia clearly has spectacular software and digital hardware designers.
Mechanical, analog, and power engineering are entirely different disciplines.
It doesn’t add anything to OP’s point that it took so much effort to engineer a GPU only to be messed up by a stupid connector.
We've seen these failure stories time and again to various repercussions, but they seem to all come down to some form of greed whether it be reputational or monetary.
https://en.m.wikipedia.org/wiki/Space_Shuttle_Challenger_dis...
So, yeah, how about building a rocket designed to send three humans to the moon, but failed because of some <$1 part?
Kind of like some 3d printers have.
The issue is internal: because the top PSU rail is only 12V, high-power devices (mostly GPU) need to push very high amperage totals. And because PC internal layouts, they can’t do that through large wires with huge connectors.
Also a legacy of cheap friction fit.
I was wondering where you'd get that from, and then I remembered that you're probably american so your baseline is only 1800W, in which case... yeah that's true. If you're installing a really powerful PC in the US you probably want either a dedicated circuit or a 20A circuit (or both).
In North America, you cant pull 1800W continuous either, it is limited to 1440 by regulation.
"This regulation is called the maximum continuous load, which is 80% of the total wattage computed. "
This is why space heaters and such are 1200 watt.
So while 120v * 15A is 1,800W no electrician would wire a setup like this.
if you need 1800W they would use probably use a 20A circuit and wiring rated for 20A. 80% load of this would be 1920W.
NOTE: Not an electrician.. if i am mistaken please let me know.
...in areas that use 220/240V. In north america appliances are limited to around 1500W, because the mains voltage is 120V.
Those machines ran 24/7 for at least 36 months before I had to start servicing stuff like GPU fans and heat sinks.
This stuff can be designed NOT to melt the dang 6 pin 12v rail
Those machines ran 24/7 for at least 36 months before I had to start servicing stuff like GPU fans and heat sinks.
This stuff can be designed NOT to melt the dang 12v rail
You plug both PSUs in, and DC output appearing on the primary PSU switches this auxiliary one on and off.
Also the power supply to a GPU is DC anyway, so the wall voltage probably doesn't matter much.
In my next house, I think I'll request a 220V outlet in my kitchen with a UK style plug, so I can have my fast tea kettle.
They do this so that the PSU can be made up of multiple redundant modules, and to integrate some batteries that can cover a generator starting up. This way they don't need to pay for a dedicated UPS battery with chargers and inverters just to convert it back to lower-voltage DC near the server.
They can make it so that for example 5-8 PSUs spread over different electrical circuits can share the load of the servers, even if 2-3 of them fail (so they, for example, only need 2 extra on top of the 5 that can power the servers and the 1 that recharges the battery (62.5% load if everything works and the servers are running full-power), or even relying on the servers occasionally running below full power to spread the full rated server power over 6 modules with 2 spares (75% load if everything works and the servers are running full-power)).
A big issue in this situation here is though that Nvidia skimps heavily on board space for the 4090 Founders Edition cards, seemingly to use the back for heat sink fins/air-flow cross-section. It would take up some board area, but in theory, they could add a 48V power option that makes it all sleek and pretty, possibly with a 12V->48V adapter for the people who don't already have a 48V supply.
https://static01.nyt.com/images/2016/08/05/us/05onfire1_xp/0...
https://twitter.com/hms1193/status/1585257428291325958
This isn't like running kA of current where the threat of thermal runaway is always present, even with active cooling. I worked on a stellarator that had a coil feed blow (years before I started). Pushing even 10% more current meant a complete overhaul of the feeds. That's what thin margins looks like.
https://www.youtube.com/watch?v=EIKjZ1djp8c
In short: They couldn't replicate it, no conclusion yet, further testing needed. What's weird is that all their adapter cables are clearly different (better) than the one from "Igor's Lab".
Back in the early 00's, I've had 20A breakers trip with ~8 high-end PCs on the same circuit. Today, I think we would struggle to handle half the number of PCs. My entire panel is only good for 150A of service, so I am wondering if I could even fill up a battlefield match before my main trips.
According to this setup their under load power consumption of the entire system was just shy of 700Ws https://pcper.com/2014/10/nvidia-gtx-980-3-way-and-4-way-sli...
Back then components didn't tend to be as power bursty so I think it matters whether you compare peak power consumption or average power consumption under "typical" load.
A wall load from a gaming PC with a 4090 and 13900k is definitely going to burst passed 700Ws, but average while playing a game is probably lower than that.
That said, that review mentioned several very power hungry SLI setups.
Keep in mind, while the US uses 120V at the outlet, virtually all residential service is 240V (typically two phases of 120V), so in theory your home has 240V@150A=36 kilowatts of peak capacity. With some high gauge extension cords to distribute the load to different circuits in your house, you'll likely run out of space or friends before you run out of power.
It's still single phase, one AC sine wave. The transformer supplies power using three current carrying conductors, two hots, and a neutral. The voltage, or potential difference, between the two hots is 240V. The neutral is a center tap on the transformer, and the voltage between each of the hots and the neutral is 120V. This allows circuits to use either or both 120/240V.
You are correct, however, in your math for total capacity.
Users report Nvidia RTX 4090 GPUs with melted 16-pin power connectors
https://news.ycombinator.com/item?id=33327515
(116 points, 98 comments)
Nvidia’s hot adapter for the GeForce RTX 4090 with a built-in breaking point - https://news.ycombinator.com/item?id=33373729 - Oct 2022 (274 comments)
Users report Nvidia RTX 4090 GPUs with melted 16-pin power connectors - https://news.ycombinator.com/item?id=33327515 - Oct 2022 (101 comments)
But to answer your question on "why adding pins": electrical regulations regarding fire risks eschew PSUs that don't trip over current protection when you hit like 20-ish A on a single output, due to the risk of starting a fire if you were to dump that much power into a small place and happen to have enough resistance in that "short" to not trip the over current protection. That's why GPUs feed multiple power rails from the PSU and control themselves to balance their draw across their input rails to not overload any single input.
And IMO as long as we let people build computers in cases that won't protect the room/building from e.g. a broken capacitor on the GPU's 12V side drawing 600W until the GPU is burnt to a crisp and the case fan has blown the flames out the back against the e.g. wooden shelf on/near the desk, we maybe should keep up the regulations that such fire-non-containing cases can't use PSUs where individual power rails willingly deliver enough power for an electrical fire to escape the case.
https://www.crucial.com/support/articles-faq-ssd/dangerous-m...
https://old.reddit.com/r/buildapc/comments/9mq6e4/molex_to_s...
That was cause for concern so I removed them all and used splitters instead so I didn't need any of those adapters.
They are also rumored to be more power efficient which would help things
Here there is a vid about what is the issue
I thought the whole point of going direct to the manufacturer or someplace less shady than Amazon like Monoprice.
(I haven't ordered from them in a while, but while they don't ship as fast as Amazon I never had a failure, versus Amazon constantly selling me shit that breaks, then demanding I go print a shipping label and send them their less than twenty dollar hunk of whatever -- it felt like a very purposeful strategy to exhaust folks who were more than willing to do some research.
Then again, I have an uneasy alliance with "the gamers" since I switched to being a film snob in 2009 -- that's why they took out all the payphones -- it's a long story, but apparently someone formally complained to the FBI that I said they should ring the Vatican with Italian tanks and start shelling if they don't join NATO and pay the victims, so they're putting off legalizing Jenkem and taxing it for revenue for another 69 or so years.)
Should I tell him this is a possibility if he is?
How about the families of the 5 people in that house? What should they be told?
"Sorry folks. I knew this GPU has issues and could possibly cause a fire, but people on the internet decided I was a bad person for doing the right thing, so I ignored the potential issue and didn't let the landlord know."
I hope you have a great life.
Mind your own business.
Doesn't matter how many watts. All that matters is that there is resistive heating happening.
So, I have been involved via the landlord asking my opinion due to my knowledge of technology being a bit better than his.
But here people are, acting like they are somehow 'teaching me a lesson' while acting holier than thou, when they aren't. lol
And for the record. I did offer to go do a quick check with him to help rectify the situation so that everyone is being treated fairly, since he isn't that great at calculating things like power usage.
But instead he said no, cause he figures he can do the job well enough on his own. But now that this news has come up, I thought about the situation again, and figured I would leave my comment.
Clearly a mistake. I clearly should have just done what I thought was right, and not ask a bunch of people on the internet their opinion. That was dumb of me.