How thermal management is changing in the age of the kilowatt chip
theregister.com
theregister.com
https://www.youtube.com/watch?v=pzyZpauU3Ig
The engineering to handle that power density is insane. Technically its less power per mm^2, but the chip is the size of a dinner plate.
EDIT: The video was taken down, but looks like the web archive got it:
https://web.archive.org/web/20230812020202/https://www.youtu...
As well as Vimeo (thanks morcheeba): https://vimeo.com/853557623
There are reposts with that same link.
Maybe it revealed too much and was taken down?
I wonder, how do all the contacts in the 20,000A power distribution plate that is bolted on top of the wafer-scale die line up? The engineering involved in just making that part work must be crazy.
If we presume the wafer consumes a large percentage of that (say 20kW out of the 24kW max) and that they are feeding the "wafer" with DC at 1v, then they /do/ need to feed in 20,000 amps to deliver 20kW of power at 1v.
So yes, 20kamp is a lot of current, but it is within the "power budget" the device seems to express in its marketing material.
For comparison, one of the workstation AMD EPYC processors uses ~400W under peak load, and would use approximately ~320A peak current. It's only ~30x more current... modern CPUs use an enormous amount of current these days.
We have EPYC chips with ~100 cores at ~200w TDP. So each core is around 2W in the AMD chips. Core voltages are ~1v with modern CPUs. So that's 2A per core.
850k cores at 20kA is very much lower than the AMD chips. Must be massively parallel, lower-performing cores. But it's quite feasible that it needs 20kA.
The CS-2 system on their website specs out at 23kW peak. So all this lines up with each other.
As far as benchmarks and utility of such systems, I am not sure if they've proven it out.
I guess this is a process limitation? i.e, the stuff you would need to make an appropriate step down isn't compatible with the other stuff they need?
1. Process limitations (no high voltage devices, poor analog characteristics, limited resistor choice, etc)
2. The skill set for power device/analog IC design is very different than digital design (and harder to recruit for as the talent base is relatively small).
3. On chip power converters universally suffer from poor inductor quality which trashes your efficiency (thus increasing cooling demands as well).
From a business perspective it would be quite risky and likely not cost effective.
As for skillset, there are a bunch of IP companies with silicon proven designs available. I'm sure they didn't design their SERDES or PLL(s) in-house either.
They also have another design advantage here. I believe they're not IO limited with their bumps.So they can use most of it for power, which is much better than any DC-DC solution.
Cerebras has evidently scaled them up a bit (and repatented them, according to the video).
Interesting early application (if not the first) from the 1970s: http://www.hp9825.com/html/hybrid_microprocessor.html
Weirdly - CFD has been zeitgeisting me here on HN the last couple of days - I have been talking about FluidX3D - and have been attempting to compile it this AM locally on windows with failures (about to see if I can plop it into a docker)
Never thought youd be doing CFD calc to keep a stable temp flow over 1.6 trillion transistors to keep them evenly cooled - did ya?
--
In watching that above CFD vid - I was led to start thinking if CFD could be applied to the ways AI models/GPTs communicate or calculate.
I wonder if one could use an CFD anaolgies to the data flows through AI models/systems such as OpenAI.
It would be interesting to look at the OpenAI GPT Store's entanglements through the lens of CFD and determine where relations might be made for how stacking GPTs might communicate through their Tapestry.
An CFD-heads care to dive in?
I wonder if one were to treat 'token flow' in a CFD manner in which one can visualize how tokens are assigned attention scores.
GPT claims to not be able to visualize the attention score matrix for the tokens - but assuming it could - it would seem as though it should be easy to visualize attention matrices in a CFD visual.
https://i.imgur.com/jzZ1wsP.png
--
Or I sound like an idiot. Lets ask GPT. haha
---
(also - in super expensive machines , why aren't sheets of AeroGels used as gaskets if you want thermal separation.
Imagine taking an AeroGel powder and mixing it with Silicone - and having a super thin, flexible material, such as D3 - which has a melting point of 134c/273f....
So - mixing aerogel with D3 as a gasket would be good, and with D3 being a non-newtonian it works well for shocks. A space gasket as it were.
My largest design as approximately 23 x 12 inches. Maintaining thermal efficiency and uniformity across the entire surface is where the real challenge lies. We had an extremely tight uniformity specification (0.1 deg C). This could only be achieved through a complex design process that entailed writing genetic algorithms to evolve and test solutions using FEA. In the end it was a combination of sophisticated impingement cooling and other techniques that did the job. That project was seriously challenging. I like projects that absolutely kick my butt. This one definitely did.
I can tell you that we worked very hard to try and see if we could accomplish the objectives using forced air (fans). That effort involved laser-welded fins with sophisticated airflow management, techniques to break-up the boundary layer (which impedes optimal heat transfer) and powerful centrifugal fans. It worked well, yet it was large and sounded like a jet engine.
Ultimately, while still complex, fluid-based thermal management offered a far more compact solution that could exist in a room with people not having to wear hearing protection. In addition to that, with a fluid-based system you can move the hot side to a different room.
It will be interesting to see if commercial data centers re-fit with cooling water transport under the floor or above the machines (riskier). This will be a challenge for DCs without a lot of space under there, presumably they could boost the floor height after a door transition. Still how many of them have 20 - 40kW of power allocated per rack.
Its one of the few times I miss being at Google because they approached this sort of problem very creatively and with an effectively unlimited budget to try different things. I'm sure their data centers are very much different from my time there!
IIRC they mainly put power hungry compute nodes for the clusters in this new datacenter and I remember that servers full of GPUs had crazy power draw. The water then goes through an heat exchanger to help generate hot water to heat the campus and for the taps.
I have a massive surplus of solar in summer, and on my local auction sites some old miners became available. I considered picking them up and putting them to use using up my surplus "free" electricity.
I gave up when the mining ROI calculation came back with multiple years, even when you factor in a $0 electicity cost.
This might've worked 3-5 years ago (maybe more), but I don't see that happening now to be honest.
https://fortune.com/crypto/2023/12/21/bathhouse-nyc-bitcoin-...
https://www.microsoft.com/en-us/research/publication/the-dat...
Here is an example of large scale project from Finland. https://www.fortum.com/media/2022/03/fortum-and-microsoft-an...
And in the summer months here in the northern hemisphere we'll send our training jobs down to the southern hemisphere.
Now that chiplets are maturing that is a little less far-fetched.
So what's the limit of a system you could run from your own house (or apartment)?
A single standard outlet yields 1500-1800 watts but there are higher voltage/amperage outlets in many houses.
Assuming you are in Europe, and you have 2.5mm² cabling (which is the standard for residential applications) then you are indeed limited to 16A per group.
However, there is nothing preventing you from using multiple groups for one appliance. This is actually typical for high-power appliances, such as induction cooking.
Ultimately it is your main fuse that limits your total power consumption, which for most European countries this is typically rated at 25A (5750W), but on request you can usually have this raised to 35A, 50A or even 80A, if supply is sufficient.
The only limit to household wiring would be the capacity of the distribution coming into your home.
A few years later I got to know the person whose house it was installed in. And when the homeowner was talking about it he complained that the plumber didn’t install a big enough one and he had to have it redone.
On the other hand, tankless electrics often cause a service upgrade unless planned for new construction. Great for us electricians, less great for the homeowner.
Our all-electric home's max instantaneous draw over the last a couple months peaked around 44 kW.
In retrospect, it's insane to get an electric tankless hot water heater from a climate perspective. Gas tankless makes infinitely more sense if you need limitless hot water. If you don't need limitless, heat pump would be the most emissions efficient.
I'll try to get my usage data together and update this comment.
Edit: Cost of usage electric and gas over the last five years. The big spike last winter was due to our remodel and heating the whole house while it was neither sealed nor insulated in the middle of winter with just the 10kWh furnace backup heater which ran almost non-stop. Units are USD. The two gaps in data are due to meter replacements, I think. https://imgur.com/AVTFwpP
It's kind of hard to compare the data since we re-insulated, went from gas furance to heat pump, and added a hot tub.
Most houses would be able to upgrade to 35A without extensive reworking of cables, getting to a 24kW maximum.
Also 50A and 63A are available for consumers in most locations, but would require re-evaluating the cabling coming to the house.
I'm guessing it would be the same in Slovenia.
Any circuit breaker will not trip at the rated current, though. They're designed not to. So, you can run all 48kW indefinitely without tripping a circuit breaker, assuming everything else is sized appropriately (i.e., wire, interconnects, etc).
If you want to stay with your normal residential circuit, in the US they're commonly 100 or 150A. 200A isn't uncommon, but you might have to pay for an upgrade.
That leaves you with 200A*240V=480kW minus whatever you need for normal house things.
So probably more compute than you have physical space for.
You’re off by an order of magnitude: 200 * 240 = 48000
> If you're willing to pay, you can probably convince your utility to hook you up to three phase power, which is typically reserved for industrial use.
Three-phase power isn’t only for industrial use, (in the United States) small commercial buildings will have a 208v three-phase service drop and larger commercial buildings will have a 480v service drop, or 13.8kV medium voltage drop if it’s big enough. Large enough industrial customers will have dedicated substations.
Also, in Europe 3 phase power afair is fairly common.
[1]: https://lists.debian.org/debian-ai/2023/12/msg00031.html
So I assume they are using around 0.9V and 2.2kA meaning ~2kW.
In a 2U 2Node system, this is a potential of 1024 vCPU in a single server.
There are niche expensive datacenters with higher power density, but as it stands, exotic multi-kW hardware at scale makes sense if you either save a ton on per-node licensing, or you need extreme bandwidth and/or low latency.
There’s always some limiting factor, and there’s always some (possibly crazy expensive) way to resolve it and get a bit more power until you run into the next limiting factor.
Datacenter: seems to cap out at around 850 MW [2]
Same ballpark I guess? Probably both are limited by inexpensive power availability + other connectivity factors (road/rail, fiber).
[1]: “Therefore, a 300-tonne, 300 MVA EAF will require approximately 132 MWh of energy to melt the steel, and a "power-on time" (the time that steel is being melted with an arc) of approximately 37 minutes.” via https://en.m.wikipedia.org/wiki/Electric_arc_furnace
[2]: https://www.racksolutions.com/news/blog/how-many-servers-doe...
How many backup generators does that need??
I think that was the case in 2020;
>By 2020, that was up to 8–10 kW per rack. Note, though, that two-thirds of U.S. data centers surveyed said that they were already experiencing peak demands in the 16–20 kW per rack range. The latest numbers from 2022 show 10% of data centers reporting rack densities of 20–29 kW per rack, 7% at 30–39 kW per rack, 3% at 40–49 kW per rack and 5% at 50 kW or greater.
We dont have 2023 numbers and we are coming to 2024. But it is clear that demands for high power density is growing. ( And hopefully at a much faster pace )
If you train an AI, you don't need low latency/high bandwidth Internet access.
Sea water is corrosive, hard to use for cooling.
Part of the problem is consistent humidity management, but the other is that air, even really cold air, isn't dense enough to move specific heat as effectively as the same volume of a denser working fluid.
For a practical example in the other direction, look at combined cycle gas turbines which use heated air for the primary turbine and the recapture as much denser steam for the secondary.
In any case, I think you'd still need substantial cooling infrastructure to deal with the summer. Your capital costs are going to waaaay higher. While your thermal management costs will go down, your payroll will probably be higher.