Cerebras Wafer Scale Engine Gen2 7nm 2.6T Transistors
servethehome.com
servethehome.com
And SC19: https://secureservercdn.net/198.12.145.239/a7b.fcb.myftpuplo...
So everything is custom-developed, including wire bonding, packaging, connectors, heat spreaders, an impressive achievement. Unfortunately it's only a three-page overview, I'd love to read more on technical details.
Which brings us to the next topic: What does the power supply look like, and how does power delivery work? Given the huge amount of power, the PSU for the chip alone is an interesting piece of electronics. Assuming its power supply is similar to a PC or server... First, the mains power is converted to a system DC voltage such as 12 V. For a typical modern cheap and boring switched-mode power supply, the standard efficiency requirement is 80%. But in this case, it means the PSU itself will waste 3 kW of energy, this is unacceptable both from an environmental and thermal perspective. A really good supply can achieve an efficiency of 95%-98%, and this is when things get interesting, at least you'll see some expensive controller chips and high-power, high-frequency transistors in such a PSU, terms like SiC, GaN or IGBT come to mind, switching 1,250 amps on and off. And it doesn't end here, the next step is converting the 12 V system voltage to the chip's core voltage, say 1.0 V, by a Voltage Regulator Module near or on the motherboard. Now, don't even mention the technical challenge of designing a VRM. I can't even imagine how they run the huge current on the motherboard, it's 15,000 amps!
Unfortunately, there's no information about power supply and power delivery.
Edit: In the slides shared above [1] (page 16), it says "12x 4kW hot-swappable universal PSUs".
[1] : https://secureservercdn.net/198.12.145.239/a7b.fcb.myftpuplo...
I had an opportunity see a laser sintered copper heat exchanger that was very very efficient. Pumping in water and pulling out steam efficient. It was pretty cool.
I think it has the same customer base as the quantum processors of today.
I wouldnt have thought this amount of sram to be practical from a size and cost standpoint compared to DRAM.
There have been other, actual eDRAM technologies that can be manufactured on some logic processes, however none of them are compatible with the state of the art TSMC 7nm logic process.
Maybe cerebras could fit well for that.
As for the models of the future getting much bigger - is the GPT family representative ? Are we seeing the same growth is in other domains ? And isn't there a big niche for smaller than GPT models ?
Top500 Supercomputer in a single Rack. Just imagine.
And if we take liberty to dream extreme, "Available on Amazon, On Demand". :D
And someone will be spending a lot of time trying to figure out how to support a wafer cut lengthwise without shattering it during processing.
[0] https://www.anandtech.com/show/14758/hot-chips-31-live-blogs...
Maybe they're meant to be used with a DL compiler like TVM [0]?
More of the software stack was described at HotChips today, covered by AnandTech: https://www.anandtech.com/show/16006/hot-chips-2020-live-blo...
I guess these are the secret sauces and we are unlikely to get in depth information.
[0] https://www.hotchips.org/hc31/HC31_1.13_Cerebras.SeanLie.v02...
Now I know its just a LinusTechTips level video, but I guess they havent heard of "Wafer Scale Engine", and I hadn't either, but now this proves the video is already obsolete…
People have been trying Wafer-Scale Integration [0] since the 1970s, there was quite some hype of building a "super chip" back at that time [1], but all efforts failed miserably. Cerebras' success is only the beginning, even if this approach is workable (which remains a question), there's still at least a decade to go from a HPC-specific chip to a general-purpose chip. Another possibility is that WSI will forever be a technology used in massively-parallel computers.
[0] https://en.wikipedia.org/wiki/Wafer-scale_integration
[1] Giant microcircuits for superfast computers, Popular Science, 1984. https://books.google.com/books?id=eAAAAAAAMBAJ&pg=PA66
> even if this approach is workable (which remains a question)
No need to repeat.
http://www.computinghistory.org.uk/det/3043/Anamartic-Wafer-...
The problem is that the real project was a massively parallel computer just like the Cerebras (scaled to 1989 technological limits) and the disk replacement was just a way to develop the needed techniques and finance further developments. If the investors had had a little more patience then computing in the 1990s might have been a bit more interesting.
If you’re paying up for an entire batch and you only need one working example out of it, it probably matter less.
For TSMC N7 Design Rules[0] the pitch is 720nm for the top two metals (1388 wires per mm) and 76nm for the middle metal layers (13157 wires per mm).
In the case of chiplets on an interposer, if we suppose that the interposer can use a process similar to the top layers of the wafer then the number of wires between chiplets is the same. But you need pads and solder balls from the chipets to the interposer (and back on the other side). These pads can be in a grid pattern (like in a BGA package) so the area is the limit instead of perimeter. Let's be optimistic and suppose a pitch of 30µm for the copper pillar micro-bumps.
With a conservative 1000 wires per mm of perimeter the pads are not the limit with a chiplet larger than 3.6mm per side. With a more aggressive 10000 wires per mm the chiplets would have to be larger than 36mm per side, which is just beyond the reticle size limit of many chip fabs.
I don't actually know, but that seems to be the route the industry is going and economics of scale will likely make it the winning approach.
https://www.hotchips.org/hc31/HC31_1.13_Cerebras.SeanLie.v02...
> Redundancy is Your Friend
> Uniform small core architecture enables redundancy to address yield at very low cost
> Design includes redundant cores and redundant fabric links
> Redundant cores replace defective cores
> Extra links reconnect fabric to restore logical 2D mesh
One day we'll find an easier way to print chips. Won't end moore's law but may shift algorithm implementation back to hardware after decades of software eating everything.
Medium term, they'll get a new lease on moore's law by switching off silicon to smaller atoms. And/or switching to something that handles heat better so we can go 3D. This is already happening in power electronics with GaN. My bet is on this tech slowly spreading into digital logic