DRAM thermal issues reach crisis point
semiengineering.com
semiengineering.com
We don't know and it is hard to know - but I don't blame JEDEC and would not call it a "compromise" on their part like it was a superior option.
Imagine if the failure rate was raised by as little as 1%. RAM Failure is not uncommon compared to other parts - I've had it happen before and render a system unable to boot, that's why we have Memtest86 and not CPUtest86 or SSDtest86. A 1% increase in failure over 5 years would have effects just as unbelievable as the power saving that increasing the temperature would be. How many smartphones would be junked? How many PCs would be thrown out for not working by people who are average Joes who can't diagnose them, and the extra waste that generates from both disposing the old PC and purchasing a new one? Perhaps the new PC is more efficient, but which is better, greater emissions or more eWaste in the ground due to the new PC being likely more efficient than the old one?
The point is that it is not a clear win. With further research it might be, and I might be all for it. I'm only nitpicking the description of it as being a "compromise" as though it were obvious.
[@picture: I'm at my posting limit for the day because HN is, well... I'll leave their censorship policies for another day. I would agree with you if the RAM with the 90C limit were strictly ECC RAM because that is most often used in data centers and not consumer parts. Maybe we have non-ECC/85 RAM and ECC/90 RAM options...]
In the operating range, you’re not going to have any measurable change in operations or failure rate - if you do the parts are defective. All of the stories you hear about this and that are conjecture.
The reasons this doesn't work for datacenters is two-fold, I think: First, they won't see efficiency improvements just by keeping their CPUs and GPUs (or RAM) cooler because the power levels/tables for the chips are baked-in, and operators aren't going to the trouble of tweaking voltages themselves. Second, even if they did tweak voltages, the cost of sustaining lower temperatures with better cooling likely won't outweigh the savings resulting from lower power draw for the chips.
Still, this raises the question of whether designing hardware for higher operating temperatures is always the right move. At some point there's going to be a cost in performance and/or efficiency that outweighs the savings from allowing higher temperatures. Ideally these tradeoffs should be balanced as a whole.
For example take the Epyc 7742, at 225W it sounds super power hungry. But the 64-core chip only boosts to 3.4ghz max (2.25ghz base). That's less than the base clock of almost any of the Ryzen consumer CPUs. And if you look at the lower frequency end of https://images.anandtech.com/doci/16214/PerCore-1-5950X.png there's not a whole heck of a lot of efficiency gains likely to be had below that ~3-3.4ghz mark. They're already basically sipping power at something like 3w per CPU core or less. 225w / 64c = 3.5w/c, but the IO uncore isn't exactly cheap to run and iirc sits more like in the 50-70w range. So subtract that out and you're at more like 2.5-2.7w/c. I don't think throwing cooling at this is really going to get you much of a gain.
You can undervolt (or overclock) one specific chip, individuals have been doing this at home for basically ever, but there's (almost) always a system-specific validation process you then do to make sure the system is stable for a specific workload with the new parameters.
And these parameters differ between batches of chips, or even between chips within a batch.
It's also significantly harder to drastically reduce the temperature of the chips inside of a single server, given the machine density of a typical data center.
It’d be fun, however, if we could dynamically adjust that according to workload. If workload is light, you could consolidate load into fewer sockets/memory sticks and power down everything in that socket.
We’d see serious latency spikes with high volume increases. And high volume would correlate to high I/O on shared storage exacerbating the problem. You’d get a bunch of transactional failures from the spontaneous latency hit.
Same was true with older virtualization/partitioning of hardware. You’d even see apps back up into memory and swap themselves to death.
Now, the apps were awful: inefficient and intolerant. But that’s the reality for the vast majority of business software. Churned out by devs with little understanding and tight timelines for features.
Reliability being an NFR kills me.
When I was on an architecture team that consolidated ~80 datacenter to 3 circa 2010, this was a key dollar driver. We raised the temperature ~6 degrees from the average temp, which meant kicking out a few vendors initially. The cost savings for doing that was essentially the total operational costs of 5 datacenters.
The annual failure rates for the hardware did not change at all by any metric. Number of service impacting hardware failures went to zero due to the consolidation.
In general, if you operate within the operating ranges of your hardware, you won’t have failure. You will have complaints from employees, because computers will operate at temperatures not comfortable for humans.
Based on ASHARE there are two guidelines from HPE and IBM: https://www.chiltrix.com/documents/HP-ASHRAE.pdf https://www.ibm.com/downloads/cas/1Q94RPGE
We found (by measurement) that some places in the datacenter with suboptimal airflow are well over the average or simulated temperature so one can leave the safe temperature envelope if one isn't careful.
The hyper scale people take this to the next degree.
[0] https://www.google.com/about/datacenters/efficiency/ [1] https://www.cnbc.com/2022/04/13/google-data-center-goal-100p...
edit: this makes me wonder what is the ideal humidity in a data centre? Is too dry a thing?
I’m not in this game anymore, but I remember Yahoo built a datacenter in NY off Lake Erie that wasn’t cooled. They aligned the building lengthwise on a ridge above the lake and just blew air through it with no AC. No idea how it operated though.
I don't know what kind of magical hardware you ran, but from the datacenters I've interacted with this is absolute nonsense. Higher temps and/or humidity have an immediate effect on failure rates, to the point where we could plan extra time for them next week if the temp raised even .5 degrees.
Cheaper to have a slightly higher failure rate, or have the computers throttle their clock speed, than to pay for extra air conditioning.
Of course you'll need some data for calibrating the model, but if you got that?
The groundbreaking work here was by Michael Pecht back in the early nineties: https://apps.dtic.mil/sti/pdfs/ADA275029.pdf
Downtime isn't free.
Depending on design, if the DC isn't very monolithic, then you can run with components failing and not causing much trouble, replacing when available, if your apps are designed to run on a zillion small CPUs, as I believe Google and Amazon are, instead of a giant monolithic computer, 1970s style.
Could have heavier AC for critical components, the main routers and disk farms, for instance.
Details of the calculations unknown to me.
Even in most servers the accommodation for RAM cooling has basically just been orientating the DIMMs to line up with airflow. They are still packed together with minimal clearance.
Peak power consumption of DDR4 is around 375mW/GB @ 1.2V, DDR5 drops this about 10% but also increases the maximum density of a DIMM by 8x to 512GB which is like, 150W for a single DIMM.
https://media-www.micron.com/-/media/client/global/documents...
Looking a bit closer I think you are right that it's 12V but only for Registered DIMMs.
If I had to guess, Servers would not go any further than some sort of memory-backplane where the memory for multiple channels was integrated onto a single PCB.
Even then, IIRC hot-swapping of memory modules is a thing for some servers, so that will have to be handled somehow.
Isn't that server cooling in a nutshell? Ram high volumes of airflow through the chassis with stupidly loud fans and hope the parts stay cool?
I still like this a lot, but now the top down fan and some kind of ducting to help direct the air out the top / side vent makes more sense. There's so much heat these days everyone needs the baffles inside of a case.
Cost, noise, and physical box size are to be minimized, and various compute/storoage/io capacity are to be maximized. Component placement and orientation relative to one another, baffles, fan characteristics heat sink design, all play a role.
They don't just throw parts in a box and add fans until all components are within operating temperature. Gross air flow is one lever, but it has costs and limits like all the others. Designs are modeled with heat and fluid simulations.
I know this may not be a cheap solution, but why not start selling pre-built computers with active cooling systems? Refrigerant liquids like those used in refrigerators or water cooling could be an option. The article addresses this:
> Although it sounds like a near-perfect solution in theory, and has been shown to work in labs, John Parry, industry lead, electronics and semiconductor at Siemens Digital Industries Software, noted that it’s unlikely to work in commercial production. “You’ve got everything from erosion by the fluid to issues with, of course, leaks because you’re dealing with extremely small, very fine physical geometry. And they are pumped. One of the features that we typically find has the lowest reliability associated with it are electromechanical devices like fans and pumps, so you end up with complexity in a number of different directions.”
So instead of integrating fluids within the computer, build powerful mini-freezers for computers and store the computer inside. Or split the warm transistors from the rest of the build and store only those inside the mini freezer, with cables to connect to the rest of the computer outside.
But yeah, I'd be just as or more concerned about the amount of power it would take to run a freezer like that ... I'm already drawing as much as 850 watts for my PC, with a max of a couple hundred watts for my OLED TV and speakers, and don't forget the modem and router, and a lamp to top it all off; would a powerful enough mini freezer to cool my PC even fit on the circuit?
Actually, it's even worse because I've got an air purifier running there too ... but I could move that, I suppose.
Still, I wouldn't really go for that. Knowing how noisy the average compressor and fan for that size is. I much prefer my nearly silent fan cooled machine...
In a typical data center, all this does is decentralize your cooling. Now you have many smaller (typically less robust) motors to monitor and replace, and many drain lines much closer to customer equipment and power.
Those units take up a lot more space, too, because of the insulation.
These problems start to read like problems from nuclear power, where sufficiently uniform flow is a huge deal so that various materials aren't compromised in the reactor.
In theory you can eliminate condensation.
But in practice, there's a difference between theory and practice.
I am sure there is a lot more to it than that though... ha
... probably because they do the bare minimum to keep up with CPU design and routinely get busted for cartel price fixing and predatory pricing.
That's not going to be viable for servers ror two big reasons:
a) it would big a major capacity limitation; you're not fitting 8-16 DIMMs worth of ram ontop of the CPU. Sure, not everyone fills up their servers, but many do.
b) if you put the ram on top of the cpu, all of the cpu heat needs to transit the ram, which practically means you need a low heat cpu. This works for Apple, their laptop cooling design has never been appropriate for a high heat cpu, but servers manage to cool hundred watt chips in 1U through massive airflow, so high heat enables more computation.
Heatspreaders may make their way into server ram though (although not so big, cause a lot of servers are 1U)
Otoh, the article says
> ‘From zero to 85°C, it operates one way, and at 85° to 90°C, it starts to change,'” noted Bill Gervasi, principal systems architect at Nantero and author of the JEDEC DDR5 NVRAM spec. “From 90° to 95°C, it starts to panic. Above 95°C, you’re going to start losing data, so you’d better start shutting the system down.”
CPUs commonly operate in that temperature range, but RAM doesn't pull that much power, so it doesn't get too much above ambient as long as there's some airflow, and if ambient hits 50C, most people are going to shutdown their severs anyway.
Maybe instead of 32 RAM packages and 4 CPU packages, we could have 16 CPU packages each with onboard RAM?
Obviously this is still a lot of heat in a small space but it does mean the cooler gets to have good coupling with the die rather than going all the way through some DRAM first.
However there are ways to prevent and detect leaks in current system with negative pressure: https://www.youtube.com/watch?v=UiPec2epHfc
With the angle, you can also place cable connectors and whatnot on the bottom of the board so they don't obstruct airflow as much.
Basically, optimize PV = nRT inside the computer case at no extra cost other than a redesign.
It seems like we're reaching a point where a new ATX standard is required to ensure the memory and GPU can make contact with a large heatsink similar to how the trashcan Mac Pro and XBox Series X are designed. Doing so would also cut down on the ridiculous number of fans an overclocked gaming PC needs these days, my GPU and CPU heatsinks have 5 80mm fans mounted to them.
ATX is great but it seems like only minor improvements to power connectors and whatnot have been made since it was introduced in 1995.
The series X GPU then is considered equivalent to a desktop 3070, and laptop 3080s exist and are also considered equivalent to a desktop 3070, so don't require anything particularly novel in terms of cooling solutions (3080 laptops are loud under load, but so is the series X).
Overclocked components are so heavy in cooling needs as they're being run so far outside their most efficient window to get the maximum performance - which is why datacenters which care more about energy usage than gamers tend to use lower clocked parts.
A similar single heatsink design for high end PCs would need to be much larger than either of those designs but considering how much empty space is in an ATX case, I don't think it would be much larger than current PCs.
Consider that the best PC cooling solutions all look like this: https://assets1.ignimgs.com/2018/01/18/cpucooler-1280-149617...
Or pass liquid through a radiator with comparable volume, standardizing the contact points for a single block heatsink with larger fans would make computers more efficient and quiet.
And additionally, there are only a few key components of a motherboard that need cooling. Most of the passive components like the many many decoupling capacitors don't generate significant heat. The components that do require access to cool air are already fitted with finned heat sinks and even additional fans. They interact with air enough to where a slight tilt cannot make a meaningful difference.
Basically just adding a small piece of aluminum to key areas will work better than angling the whole board
I think the fan, internal baffle, and vent positions dominate the airflow conditions inside the case. So, rather than tilting a motherboard, wouldn't you get whatever you are after with just a slight change in these surrounding structures?
Furthermore, I've never seen a case, either desktop or rackmount, that allows one to angle the fans at anything other than a 90 degree angle or parallel to the board.
None of this makes sense in terms of fluid dynamics.
But, given that boards do not have completely standardized layouts, it seems like you eventually need to assume a forest of independent heat sinks sticking up in the air. You lose the commodity market if everything has to be tailor made, like the integrated heat sink and heat pipe systems in laptops.
Cables also don't really obstruct the airflow like at all.
Before going into water cooling, a change in form factor to allow for better airflow (and mounting of larger heat sinks) would be in order.
Water cooling would require a water cooling block, not sure how it would work with the current form factor.
> So instead of integrating fluids within the computer, build powerful mini-freezers for computers and store the computer inside. Or split the warm transistors from the rest of the build and store only those inside the mini freezer, with cables to connect to the rest of the computer outside.
That's impractical. You are heat exchanging with the air, then you are cooling down the air? Versus exhaust the hot air and bringing more from the outside. You just need to dissipate heat, active cooling is not needed.
Since then, lots and lots of work has gone into making clean engines simpler. Open the hood of a gas-powered car these days and you'll find relative simplicity.
For example, humanity hasn't been able to find a single appropriate material for a superconductor at room temperature/atmospheric pressure despite significant research, but a civilization living below 100 K has a myriad of options to choose from. Superconductors are high technology to us, but if your planet is cold enough then superconducting niobium wire would be a boring household item like copper wire is for us.
The Anthropic Principle is not luck.
We are lucky that those interesting things are possible. We are also unlucky that many interesting things are not possible. But given that they are possible, it was almost inevitable that most of them would be possible around us.
It’s lucky some available material worked the right way to make a transistor.
It’s lucky some person smart enough to make that work got to work on that.
History is full of lucky coincidences like that. How many Einsteins have died out in the jungle, without access to our scientific knowledge or a way to add to it? For most of history and partly still today, being a scientist wasn’t possible for just about anyone, you had to be from the right family. It’s all about luck.
There are ICs and components built for operating in extreme environments, like drilling. You can get SiC (silicon carbide) chips that operate above 200°C (473K), if that's important to you. There are also various semiconductors that are worse than silicon at handling high temperatures, like germanium. Old germanium circuits sometimes don't even work correctly on a hot day.
If we lived at 200K, I'm sure that there's a host of semiconductor materials which would be available to us which don't work at 300K.
"And it’s not just servers. With about 8 billion transistors on a single die, a mobile phone can get so hot it might need to spend a few minutes in a refrigerator. When that happens, the apps will fail to function correctly."
Such a weird way to phrase this! All I can think is ".. and it's hard to use your phone when it's in the fridge"?
And then when I do the same thing in the winter I have to be careful about turning the heater on because the hot air will make my phone auto shut down from the heat.
This review of Lenovo's new Thinkpad X1 Yoga is what a heat problem looks like:
>Unfortunately, the laptop got uncomfortably hot in its Best performance mode during testing, even with light workloads.
https://arstechnica.com/gadgets/2022/07/review-lenovos-think...
A laptop with two fans that gets too hot to comfortably touch, even under light workloads, unless you set it to throttle all the time? That's a heat problem.
One can draw different thresholds, ie. 100% performance for x amount of time before throttling happens etc..
But the M2, throttling included, was still 10-20% faster than the M1 when it never throttled. Sorry but I don’t remember the exact numbers.
So basically it’s sort of a non-issue. You were promised a faster laptop and you got one. Yes it’s not up to it’s theoretical potential but it’s not like it’s at 100% or less what the M1 did.
The 13” is kind of a weird computer anyway. The 14” and 16” will be a much bigger test. We’ll see how that fairs.
This whole thing just doesn't make sense, according to the youtubers the M1 Air was half computer and half the Second Coming of Christ (and it's a lovely machine, I have one), and then the next generation comes, IMPROVES both the performance and the thermal behavior, and suddenly there's an issue.
It could be interesting in the future to talk about platform-constrained performance. In other words, take two 28W chips in 15W platforms. Both can boost for short periods. But how do each perform when throttled to 15W?
Why isn't it also true that you can only make something warmer, by cooling something else by a larger amount?
The movement of electricity generates waste heat, why isn't that process reversible? Making the heat disappear into a cold wire, rather than just dissipating into the atmosphere? (not suggesting it's would be easy or even practical).
[1] https://www.britannica.com/science/Seebeck-effect
[2] https://www.amazon.com/Peltier-Cooler/s?k=Peltier+Cooler
With present thermoelectric effects, using a Seebeck junction to generate current for a fan is hopelessly ineffective. But is that necessarily the case for all designs which could help to hold a system under a critical temperature when heat spikes.
IIRC TSMC's 135MBit 5nm example is 79.8mm^2, although that's got other logic.
In the abstract, a 0.021 square-micrometer-per-bit size [1] says you'd need about 21mm^2 for a gigabit (base 10) of 5nm SRAM, without other logic.
Micron claimed 0.315Gb/mm^2 on their 14nm process, [2] so somewhere between a factor of 6 and 7.
That said, my understanding is that there is some sort of wall around 10nm, where we can't really make smaller capacitors and thus the limitation on things. (This may have changed since I last was aware however.)
(There is also the way than 'nm' works these days... but I'm not qualified to speak on that)
Also, AFAIK SRAM is still broadly speaking more power hungry than DRAM (I may be completely out of date on this though...)
[1] - https://fuse.wikichip.org/news/3398/tsmc-details-5-nm/
[2] - https://semiengineering.com/micron-d1%CE%B1-the-most-advance...
> (as a standard metric, about once every 64 milliseconds)
64 milliseconds? wow.. I thought they'd need refreshing way more often