Chiplet ASIC supercomputers for LLMs like GPT-4
arxiv.org
arxiv.org
So they are comparing actual implementations with a theoretical implementation. Never mind that they got the A100 figures wrong, they are still in the 'wouldn't it be nice if we had 'x'' stage. This looks like a paper whose sole purpose is to raise funds for a research project that will probably ultimately go nowhere and they needed a reason that looks good on paper to increase their chances of getting funded. A100 can already be had for $0.87/hour so even their theoretical advantage is under significant pressure and assuming they got everything else right by the time the project has run the market will have overtaken them. This is what usually happens to CPUs that are application specific.
Pragmatically the prices are closer to $2/hr according to this recent post here on Hacker News: https://llm-utils.org/Nvidia+H100+and+A100+GPUs+-+comparing+...
Although again prices change on a daily basis on spot providers.
https://cloud.google.com/blog/products/compute/a2-vms-with-n...
That's as close as I got to verifying that price.
https://fullstackdeeplearning.com/cloud-gpus/
I feel there are more fair criticisms of that paper than its inclusion of the snapshot price of variable priced compute resource.
Of course, it's a first architectural study to illustrate the promise of the idea, lots more details to work out in a physical implementation and the final realized benefit is likely to be lower. But even a 3X is huge in this space.
> 2) Metrics: We use three performance metrics: (i) latency, i.e., end-to-end output generation time for a batch of input prompts, (ii) token throughput, i.e., tokens-per-second processed, and (iii) compute throughput, i.e., TFLOPS per GPU.
This is somewhat confusing to me because at least two of these three definitions should be essentially the same thing, but I don't think there's any way to interpret their claim of ~74 teraflops achieved other than ~211 tokens/second of throughput.
Put another way, 18 tokens per second is 2% flops utilization, which we are obviously capable of doing better than for bulk inference.
3x is not huge in this space because just using a 4090 instead of an A100 is a 5x gain.
This table was very helpful by the way, I didn't see that before. To me it clearly shows that 211 tok/s/A100 is very plausible and in fact kind of a poor showing because if you look at table D.4 and specifically the results for BS=256 PP3/TP8 they achieve ~150 tok/s/A100 on a model that's 3x larger than GPT-3.
A key architectural feature to achieve this is the ability to fit all model parameters inside the on-chip SRAMs of the chiplets to eliminate bandwidth limitations. Doing so is non-trivial as the amount of memory required is very large and growing for modern LLMs
...
On-chip memories such as SRAM have better read latency and read/write energy than external memories such as DDR or HBM but require more silicon per bit. We show this design choice wins in the competition of TCO per performance for serving large generative language models but requires careful consideration with respect to the chiplet die size, chiplet memory capacity and total number of chiplets to balance the fabrication cost and model performance (Sec.3.2.2) We observe that the inter-chiplet communication issues can be effectively mitigated through proper software-hardware co- design leveraging mapping strategies such as tensor and pipeline model parallelism
Also, a GPU is already an ASIC but with a fancy name.
Recent x64 chips are at about that amount of L3 cache which might be pretty similar. I've lost track of GPU hardware specs.
That proper hardware software co-design to mitigate communication? Viciously difficult bordering on imaginary.
There's still a lot of compromises and tradeoffs to be made:
> We observe that the inter-chiplet communication issues can be effectively mitigated through proper software-hardware co- design
Doubtful. Especially given it's all vapourware. Codesign is not adequately magic to handwave away this one.
SRAM is for data that needs to be read/written/used very frequently - for example, read in 1 out of 10 clock cycles.
LLM weights are certainly not this. If a GPU is calculating 200 tokens per second, then most weights are only used 200 times per second. For a 1 GHz GPU, you're only using the data for 1 cycle out of 5,000,000! The rest of the time, that SRAM is wasted power, wasted silicon area, and eventually wasted dollars.
Instead they should use SRAM for intermediate results (ie. the accumulators) of matrix multiplication - they will end up being read/written every few cycles.
Weights should be streamed in from in-package DRAM. Activations too (but they are often used multiple times in quick succession, so it might make sense to cache them in SRAM).
Costs millions per chip though.
Source: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
Look at how successful AMD's chiplet strategy has been. Chiplets sidestep the yield problems. Wafer scale amplifies them hundred or thousand fold.
Nothing in the industry is designed to work with wafer scale products, so everything has to be custom made. Yes this is a chicken-egg problem, but it's going to be expensive to get any sort of momentum. The silicon industry is extremely conservative.
It's sexy and enticing. If someone can make it work that's awesome. I will remain skeptical though.
>It's sexy and enticing. If someone can make it work that's awesome. I will remain skeptical though.
But Cerebras has made it work since 2019 as @cubefox pointed out. They're on the second generation already and they have been shipping to customers for years.
Here's a good overview of how they did it: https://www.anandtech.com/show/14758/hot-chips-31-live-blogs...
AI cannot safely and cheaply be used to drive trucks. But as soon as it can with one truck, it can with millions of trucks. All at once.
AI and robots haven’t advanced enough to replace construction workers for fluid, dexterous tasks, but as soon as they do, robots can replace millions of construction workers and surpass them in sophistication and speed.
This will happen in our lifetime, and the change will be extremely transformative.
I think modern LLMs are powerful enough now that they will still be useful in a couple years even if they aren't state-of-the-art. ChatGPT still lets you run their older model for cheaper than running GPT4, I could see a world where GPT4 is still available in 1.5 years even if there are better models out there.
This development presents a more compelling case that we are in fact on the precipice of larger LLMs being able to serve everyone for cheap. Still not really convinced by the AGI argument, but this does spook me. Overall though very cool.
Group at my school recently got a grant for 10MM for such a fantasy. All they had was an ISA - no RTL, no functional model, no compiler. Kid in my group (co-advised) is busy scrawling assembly on notebook paper lol. Suffice it to say I don't have high hopes for a tapeout anytime soon.
I'm not exactly convinced though, since all the results seem to be purely theoretical or simulated. I would've liked to see a prototype built across several FPGAs with clock speeds extrapolated for ASICs.
Also does this lower total cost depend on SRAM being available for DRAM prices?
What makes SRAM so much more expensive than DRAM?
If the design cannot serve models of this level, there will be no economic interest.
And a comparison with Jim Keller's Tenstorrent AICloud?
The extra stuff TSMC must do to pull that off are probably expensive... But I can't imagine it being, say, 10x more expensive than a wafer full of reticle sized dies (like Nvidia does). And thats setting aside the massive IO advantage of Cerebras's mega die.