The World’s Largest Computer Chip
newyorker.com
newyorker.com
https://en.wikipedia.org/wiki/Wafer-scale_integration
... as backed by Sir Clive Sinclair: Ivor Catt's Anamartic Ltd.
http://www.ivorcatt.co.uk/x5as.htm
https://en.wikipedia.org/wiki/Ivor_Catt
http://www.computinghistory.org.uk/det/8199/Anamartic-Limite...
The official page demonstrates relative size at quick glance (I guess they do use fingernails and dinnerplates but eh): https://cerebras.net/chip/
The first 1 trillion transistor chip was Samsung's 3D NAND chip, and it went with rather little fanfare.
P.S. 2 — Google is by far not the first company to do "automatic floorplanning." This is what literally every EDA does.
Here, instead of that, the circuits are overlap-printed so that a single wafer can support a set of 80 connected circuits, which are now physically cooleable because of the flat design? While they must be sacrificing some interconnection richness, because of geometrical placement, for AI applications this probably doesn’t matter so much. Very interesting.
It's not quite that direct a limit. While performance almost always wants bigger chips there are counterbalancing forces (yield and heat dissipation) that want smaller chips.
Chip manufacturing is (roughly) limited by defects until you reach the physical size of a wafer. The larger your chip the more likely that some critical circuit hits/has a defect and fails. For example, if a wafer has 10 defects, but you are producing 1500 chips then at worst you will get a yield of 99%. If you are producing 100 bigger chips, you may get a yield of 90%. If you are only producing 10 (big!) chips, you may get a yield of 0%. This drops as the square of the chip dimension.
After you manufacture the chip, chip size is limited by power distribution and heat dissipation. The bigger your chip the more it needs of both until you can't supply it or cool it anymore. This is why "3D" chips don't really win--getting power in and heat out scales as surface area--not volume.
Why wouldn't these giant chips be wired together into a cluster too?
1 wafer, doing X work, in Y time
= 1 wafer, doing 2X work, in 2Y time
= Two wafers, doing 2X work, in something still close to 2Y time
I.e. the slowness of between-wafer communication, vs. in-wafer communication, will dwarf the computing time. Obviously there is some N, where N wafers would be worth clustering, but it might be quite high.
Maybe the company is working on ways to cut down cross-wafer communication too. Vertical optical connections for instance would be awesome.
Can any materials scientists or engineers comment on if other elements will withstand higher heat better than silicon? Seems like such a large chip would be somewhat better to run at higher temperature rather than budget for huge and elaborate cooling. (This is very much a layman's question. The people who designed the chip and its cooling are far, far smarter than me!)
My kettle has 2 kW and doesn't take long to boil water from room temperature. I reckon you could fit four such kettles on the chip area (roughly). That means the chip would roughly boil water two times as fast as my kettle, were it used for that purpose.
While that does pose reasonably interesting engineering challenges regarding coolant throughput etc., I don't think there's anything particularly difficult there. You probably would want a better heat transfer medium between chip and water than my kettle has (well, I have not disassembled it), but I agree with GP that a bunch of water pipes will work well enough as a cooling solution.
Edit: Actually, screw it, we can calculate how much water we need to put through there. Warming water from 20 to 100 °C takes 334 kJ/kg. (That comes out to 167 seconds to heat 1l of water in a 2 kW kettle, for reference.) To remove 15 kW of heat with water cooling, assuming the water goes in at 20 and comes out at 100 °C, we need a throughput of 0.045 kg/s = 45 g/s = 45 ml/s.
Sure, the temperature range may be a bit optimistic, but 45 ml/s (one liter every 22 seconds) is literally "just hold it under a running tap". The main engineering challenge would be making sure that heat is removed evenly enough, I guess.
I.e. its like a kettle that isn't just radiating the stove coil's energy away, it is actually trying to keep your stove coil cool while its turned up to 10!
The same amount of energy movement, but not the same problem at all.
https://f.hubspotusercontent30.net/hubfs/8968533/Cerebras-CS...
The traditional computer included in the box is probably quite high end and power hungry too so that it can provide enough data to maintain those bandwidths. They don't appear to sell the chips by themselves.
I think the comparison is with an equivalent gpu cluster like the nvidia DGX systems or HPC CPU nodes. The DGX A100 is 6.5kW for example,
https://images.nvidia.com/aem-dam/Solutions/Data-Center/nvid...
The Cerebras system fits 15 rack units which is more than 2x larger than the DGX (6.5U). A similar 15 node HPC server with CPU is probably not that far from 15kW either (2 socket per node, 250W per CPU is already 7.5kW, then add RAM etc.) so by HPC standards it's less "full room of servers" than "single cabinet".
The trick isn't total power, it's power density on the die. In that regard, I don't think this is pushing the boundaries. It just needs custom built interfaces.
[0]: https://de.wikipedia.org/wiki/Siliciumcarbid#cite_ref-22 [1]: CREE Wolfspeed's CGHV1J070D ; datasheet: https://cms.wolfspeed.com/app/uploads/2020/12/CGHV1J070D.pdf
There are considerable efficiency gains from running silicon CMOS at LN2 temperatures instead of room temperature, but the benefits fall apart once you realize you'll have to heat-pump the electrical consumption from 77K to room temperature. Main benefits would be being able to run them faster, and a good part of the optimization would need lower dopant concentrations in the transistor channels to properly take advantage of the low temperature, which unfortunately rules out common shared-wafer prototyping runs (so testing this IRL isn't really accessible).
Cerebras still seems to have an advantage here, because they can use on-chip interconnects, which potentially allows higher bandwidth between the tiles.
Considering that he ruined someone's life, why is he revered in computing circles?
https://www.mercurynews.com/2015/11/14/gene-amdahl-father-of...
As a motorcyclist who was hit and very nearly killed by a bad driver, that has quite some, er, impact upon me. :-/
Just don't write a popular well-reviewed book that mentions in two sentences how dating in Silicon Valley is different from back home -- you'll never work again!