Intel plans immersion lab to chill its power-hungry chips
theregister.com
theregister.com
There's a coolant that's widely used, non-toxic, environmentally benign, cheap, abundant and non-flammable. Yeah, water. So it's not dielectric so needs some engineering. But humanity has a decent track record of building systems with pipes, hoses, heat exchangers and so forth. Same can't be said for cleaning up Superfund sites.
It's a pretty huge leap to go from "closed loop CFC cooling system for a computer" to "superfund sites".
What am I missing? If we're building a system with pipes and heat exchangers, why can't the coolant be a low-impact CFC rather than water? It's not a system where you just vent those cooling liquids in the atmosphere.
Sulfur hexafluoride, used in high voltage circuit breakers, has a half-life of 3200 years, has a global warming potential 22800 times that of CO2.
So you don't want to vent them, but any accident/leak can be considered a catastrophe.
That's just the physics of it: highly dielectric + stable often makes for a big greenhouse gas offender.
This is much worse and the scale the computer consumer market is huge.
Semiconductor lifetime goes down with temperature, roughly doubling failure rate for every 10C increase.
So if you want to heat something, it can only be to a low temperature. Heating buildings can work, but pasteurization requires temperatures near 100 C which reduces chip lifetime too much to be economical.
Quote from googling: “Single phase immersion cooling fluids can come under several categories which include: hydrofluoroethers, hydrocarbons, silicon oils and water/glycol. Single phase immersion cooling has benefits over 2 phase immersion cooling, in that they tend to be less expensive both due to the liquid itself and the system used to contain them. The ease of implementation was highlighted by Varma [132] who compared Novec 7000, a 2-phase fluorocarbon-based fluid, to GRC Electrosafe, a single phase hydrocarbon based fluid”
[1] https://www.datacenterknowledge.com/archives/2012/09/04/inte...
But most importantly, they seem to be pricy. If this stuff is going to be successful in the datacenter industry, a lot of it will be required. So intel will likely go with the cheaper option -- mineral oil.
No but seriously why don’t they build a 128-core atom server. That’s really all anybody wants. I don’t need the fastest most immersed cpu ever, just a bunch of decent ones at 30W or less.
They were an uncomfortable middle ground though, between normal CPUs and GPUs. My benchmarks showed that there wasn’t much of an advantage over 20-ish normal xeon cores (for my HPC workloads).
(Memory is a little fuzzy - that was 4-6 years ago).
https://www.intel.com/content/www/us/en/products/details/pro...
Though, unless if you 100% need X86, there is the Ampere Altra 128 core Cortex-N1 chip.
30 watts is low power mobile and "edge" compute.
After M1 and Graviton, there's a big rush for making denser server CPUs, pretty much bifurcating the market. Much more powerful than old Atom cores, but more power-efficient than regular ones, and targeted mostly at hyperscalers. AMD should have a dense variant of Zen 4 out next year, going up to 128 cores and reducing per-core power significantly. Intel will likely have a competitor out in 2024, and AMD might hit 256 cores then, although rumors of things that far out aren't very specific. On the ARM side of things, NVIDIA is launching CPUs like this next year. Qualcomm bought Nuvia, which was working on similar server CPUs, but probably switched focus to PCs. There are some other oddballs like Tenstorrent as well.
One interesting part with pushing per-core power so low is that communication between cores starts to be a serious part of total power consumption.
They keep designing chips that use less and less per core. Why not cram as many cores as possible into a single server? Lower density means more wasted material, and once you go below 2GHz you're not saving very much power any more.
Manufacturing yields
Not that it matters either way for this discussion? Whether it's "one server" in that 2U box, or "twenty servers" in that box, you're still shoving in lots of cores and lots of watts to get that density up.
Part of that, I think, was lack of parallelism in applications: in order to fully take advantage of those cores, you need to have a nearly-embarrassingly-parallel problem. Otherwise, you're not going to get the performance that you'd expect (but you'll get power efficiency!).
The fact of the matter is, that's not "all anybody wants."
But as the article points out, if 40% of your DC's power consumption is in cooling, then you'd be foolish not to target that slice.
Liquid and immersion cooling allows higher power density, which all things being equal (I know there's a lot of heavy lifting being done by this...) will be preferred. Why distribute your components over a rack if you could fit it into a single 4U board? Why distribute your components over an aisle if you could fit it into a rack?
https://www.google.com/search?channel=fs&client=ubuntu&q=chi...
Not sure we need that, except in niches. At scale you often want at least some efficiency, which is certainly not max TDP per core (because the best efficiency point is with lower frequencies and higher width, not the max freq you can achieve). So remains the question of large number of cores, but at some point the area of silicon also goes stupid high. And you can put multiple packages, without sacrificing overall system density too much, and without departing from simpler, and probably lower TCO pollution.
For small systems it depends, but you actually often have even more limited thermal budget, except again in niches if you are ready to tolerate the drawbacks (stupid power req it even becomes hard to have just a few machines on a basic electrical network in standard homes or offices, high noise under load, obviously high TDP so heating up a lot). But you have less space constraints so if you really want absurd systems you already can.
So do we really need to e.g. double or triple the (electrical/thermal) power density at scale? Do we need 2 kW chips? Do we need to sacrifice the efficiency now, and increase the nominal consumption now, instead of waiting just a few years for node improvements? (And I could even ask: do we really need that much increase of processing power, shouldn't we start to optimise for the total ecological cost instead? and I've not tried to do some prospective in that area but maybe this would mean slowing down the processing power growth...)
Will we ever be able to double or triple our (general purpose) compute per cpu any other way? Moore's law is essentially over. Node improvements aren't really happening outside of TSMC, which doesn't have enough manufacturing ability to supply everyone, and even then those node improvements are getting more and more incremental.
And regarding power consumption, I think we really need to be consuming more energy across most sectors of human activity. The most likely explanation for the "great stagnation" is that our energy consumption has basically flatlined since the 70s. It appears that on a civilizational scale, reaching greater levels of development and expression simply requires more Joules. If you disagree, I highly recommend the book Where is my Flying Car?
This future probably won't happen, but it should.
[0] https://whatisnuclear.com/blog/2020-10-28-nuclear-energy-is-...
[1] https://www.sciencedirect.com/science/article/abs/pii/S00945...
> computers are incomparibly more efficient
Yet to be clear the total world energy consumption for computers has increased, which is sad because we could certainly cope with 1/2 of the current total speed capacity but way more efficiency. Trying to get very high TDP chips and/or density is likely going in the other direction (but I could be wrong for the datacenter).
Linus Torvalds said that ARM needs to be widespread on the desktop, because you need a critical mass of developers targeting ARM. ARM server vendors don't want that critical mass, they want special deals with a handful of big companies which is obviously doomed to fail.
Alder Lake and M1 Pro are good demonstrations of those two approaches.
This graph shows a factor of 100 between the highest-performing and most-efficient systems: https://en.wikipedia.org/wiki/Performance_per_watt#Examples
As a note, I went down the single-phase immersion cooling rabbit hole (and designed a C-shaped tank for the build) but gave up when I read the instructions for disposal of the coolant (3M's Novec) and found them a little bit too complicated for my taste. I wouldn't want a fish tank of that on my living room.
Electricity by wire is of course much more convenient to use but I still wonder could there be a use cases for this.
Unfortunately it's hard to find it again.
What's wrong with water and cooling blocks? I'm sure they could develop some quick connect hardware, paired with sensors and valves, so that any leaks could be auto-stopped.
You could build the connectors such that pressing and holding the release button causes the whole loop to drain by suction, for near zero dripping as long as you wait a few seconds first.
Meanwhile immersion cooling is the opposite - Just build a tank and drop your stuff in it. Modern immersion cooling fluids aren't like mineral oil so you don't end up with components permanently coated in oil.
There's no pressure to push anything off it's connector, and in the event of anything failing and disconnecting in a major way, the loop starts draining the other way by atmospheric pressure.
Water pumps are pretty cheap, you could have triple redundant pumps per rack for an irrelevant cost.
Air pumps are also cheap. You could even enclose your connections to heat blocks in a second layer of slight negative pressure, that would both contain any drop-an-hour leaks that somehow happened, and funnel it to a humidity sensor.
If they use anything like Fluorinert, I'm pretty sure a whole rack of this kind of tech could cost less than a single gallon of that stuff.
I'm actually kind of surprised consumer water cooling setups aren't like this. It might be a pretty fun project, and you could probably use 3D printed parts depending on how much you trust the active controls.
Keep in mind there are other hot chips besides the CPU. Memory, network interfaces, and power supplies all need cooling too. Most systems are hybrid: water for the few hottest chips and air for everything else.
I assume gaming PCs are where you'd start?
I've never actually had a gaming desktop or anything similar, so I'm not sure what that market is like or if anyone would pay several hundred dollars for one, or if any of those companies would ever hire someone for that without a background in this stuff.
I'm mostly in embedded controls and prop building, and have lots of incidental experience with water handling and compressed air, but I've never done anything related to data centers or anything significant with desktops.
The waterblock, just acrylic with some fittings, alone is in the 130 euro[0] retail prices. So, there is some market for enthusiast but it's not large, of course.
[0]. https://www.ekwb.com/shop/water-blocks/cpu-blocks/velocity2
In this case you would need to sell the whole kit already sealed by the factory to guarantee the absence of operator error. Perhaps the OEM could install the cooling system from the get go with high quality quick connect ports integrated into the case.
Don't forget to polarize the damn connectors so that you can't mix up the inlet and outlet ports.
I was thinking to not even use separate inlets and outlets at all. The connectors would be 3-port in/out/vacuum air, with the vacuum port meant to be very slightly leaky.
You'd wrap the whole entire connector in a silicone sleeve, so that any small leaks in the connector(If the negative pressure on the water itself failed) would be contained and sucked back through the vacuum hose and trigger shutdown.
For a prototype you'd just 3D print some connectors and wrap them with tape, and let a mechanical engineer figure it out if your ever get to production, since poor quality, leaky parts are a good thing in a prototype, where the goal is to prove that the server stays dry even when things fail.
That seems to have immersion properties, without having the same toxicity.