I don't expect good yields from a chip that takes up the whole wafer. They must disable cores and pieces of SRAM that are damaged. How is this programmed?
I don't expect good yields from a chip that takes up the whole wafer. They must disable cores and pieces of SRAM that are damaged. How is this programmed?
Full disclosure: I am a Cerebras employee.
There is extensive support for TensorFlow. A wide range of models expressed in the TensorFlow will be accelerated transparently.
For everyone else: normally a wafer is divded into dies, each of which (loosely) are a chip. Yield is a percentage of good parts and it's very unlikely that an entire wafer is good. Gene Amdahl estimated that 99.99% yield is needed for successful wafer scale integration:
https://www.eetimes.com/document.asp?doc_id=1335043&page_num...
I was an SE at an hardware company and it is the first thing that you do as a product manager.
18 GiB × 6 transistors/bit ≈ .93 trillion transistors
The issue with HBM is that it's much slower, much more power hungry (per access, not per byte), and not local (so there are routing problems). You can't scale that to this much compute.
If this doesn't answer your question, I'm stuck as to what you're asking about. They use SRAM because it's the only tried and true option that works. Lots of SRAM means efficient execution of small batch sizes. If your problem fits, good, this chip works for you, and probably easily outperforms a cluster of 50 GPUs. If your problem doesn't, presumably you should just use something else.
Our startup has been working on a full Wafer Scale Integration since 2008. We are searching for cofounders. Merik at metamorphresearch dot org
Also, do you have a runtime to run the chip as a single 400,000 core CPU with some kind of memory mapped I/O so that a single 32 or 64 bit address space writes through to the RAM router through virtual memory? I'm hoping to build a powerful Erlang/Elixer or Go machine so I can experiment with other learning algorithms in realtime, outside the constraints of SIMD-optimized approaches like neural nets. Another option would be 400,000 virtual machines in a cluster, each running a lightweight unix/linux (maybe Debian or something like that). Here is some background on what I'm hoping for:
https://news.ycombinator.com/item?id=20601699
See my other comments for more. I've been looking for a parallel machine like this since I learned about FGPAs in the late 90s, but so far have not had much success finding any.
I wonder if they figure that out every time the CPU boots, or at the factory. At this scale, maybe it makes sense to do it all in parallel at boot. Or, even dynamically during runtime.
There may be edge case cores that sort of work, and then won't work at different temps, or after aging?
EDIT: to add a bit more and possibly address the original question (which I think keveman may have misunderstood), there will usually be some hardware dedicated to controlling the chip's redundancy. Part of that is often a OTP fuse-type thing that can be programmed during wafer test to indicate parts of the chip that don't work. Something (software or hardware) will read that during boot and not use those parts of the chip.
With this many cores it seems like the probability that a core dies during a multi-hour job (or in case it's used for inference, during a very long-lived realtime job) is pretty high, so the software in all layers would need to handle this kind of exception. They probably don't, today, since we haven't seen a 400k core chip before.