Cerebras Inference: AI at Instant Speed
cerebras.ai
cerebras.ai
As near as I can tell, from the model card[1], the majority of the math for this model is 4096x4096 multiply-accumulates. So, there should be 70b/16m about 4000 of these in the Llama3-70B model.
A 16x16 multiplier is about 9000 transistors, according to a quick google. 4096^2 should thus be about 150 billion transistors, if you include the bias values. There are plenty of transistors on this chip to have many of them operating in parallel.
According to [2], a switching transition in the 7nM process node, is about 0.025 femtoJoule (10^-15 watt seconds) per transistor. At a clock rate of 1 Ghz, that's about 25 nanowatt/transistor. Scaling that at 50% transitions(a 50/50 chance any given gate in the MAC will flip), gets you about 2kW for each 4096^2 MAC running at 1 Ghz.
There are enough transistors, and enough RAM on the wafer to fit the entire model. Even if they have a single 4096^2 MAC array, a clock rate of 1 ghz should result in a total time of 4 uSec/token, or 250,000 tokens/second.
[1] https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md
[2] https://mpedram.com/Papers/7nm-finfet-libraries-tcasII.pdf
Not the entire 70b fp16 model. It'd take 148GB of RAM to hold the entire model. Each Cerebras wafer chip has 44GB of SRAM. You need 4 of them chained together to hold the entire model.
Here’s an AI voice assistant I built that uses it:
But then I asked her to integrate sin(x) * e^x and got this bizarre answer that started out as speech sounds but then degenerated into chaos. Out of curiosity, why and how did she end up generating samples that sounded rather unlike speech?
Here's a recording: https://youtu.be/wWhxF7ybiAc
FWIW, I can get this behavior pretty consistently if I chat with her a while about her voice capabilities and then go into a math question.
The underlying model is not voice trained -- she says things like "asterisk one" (reading out point form) -- but this is a great preview for when ChatGPT GAs their Voice Mode.
Llama3 with ears just dropped (direct voice token input) which should be awesome with cerebras [2]
[1]: https://kitt.livekit.io [2]: https://homebrew.ltd/blog/llama3-just-got-ears
If you have an H100 doing 100 tokens/sec and you batch 1000 requests, you might be able to get to 100K tok/sec but each user's request will still be outputting 100 tokens/sec which will make the speed of the response stream the same. So if your output stream speed is slow, batching might not improve user experience, even if you can get a higher chip utilization / "overall" throughput.
"If you care and have to ask it's not for you".
In all seriousness I've worked with and am familiar with Cerebras, Groq, etc. Let's just say GPUs still reign supreme in terms of practicality outside of usage of their hardware via cloud for nearly all use-cases.
Groq, for example, has essentially stopped selling their "real" HW directly because the borderline absurd amount of floor space, etc was found to be challenging once they hit the market. There's enough demand and more (recurring) money to be made anyway hosting services on your chips.
Similar to the Bitcoin mining ASIC game in the heyday - sure we could sell these or we could just use them to mine, develop next gen, sell previous gen, repeat.
https://liquidstack.com/blog/breaking-the-thermal-ceiling-in...
It's actually worse for the majority of GPU implementations for large models. The matrices don't fit shared memory so the model is loaded many, many times to shared memory (as tiles). Also, unless you are using Hopper distributed shared memory, CTAs can't even share across them.
It would be nice to see a Cerebras solution for pre-training and fine-tuning.
A single A100 processes at 13 t/s/u for batch 32. That costs $10k to process 39 billion tokens over 3 years = $0.25 tok/s/u. If you have batch size 420 you can do it even cheaper.
TL;DR: Cerebras are certainly advertising at a loss-leading price and will only have a viable product if they can get extraordinarily high utilisation of their system at this price. I don’t think they can, so they’re basically screwed selling tokens. Maybe this is to attract attention in the hope of selling hardware to someone willing to pay a premium for very low latency, but I suspect it’s just a means of getting one more round of funding in the hope of reducing costs in the next version.
If this business scales, they can probably afford to lower the price by a factor of 10.
So for $20,000 (in quantity) you get somewhere around 10 trillion transistors?
That's enough for about 50 4096x4096 multiply-accumulate chips. At a nice slow 1 Mhz clock rate, each would take about 3.5 watts, and give you 16 teraflops of performance. If you stepped up the power and cooling, you could likely get to 350 watts of power at 100 Mhz, and 1.6 Petaflops.
50 of those chips, for $20,000 --> $400 each
This video from Cerebras perfectly explain how they solve the interconnect problem, and why their approach greatly reduces the risk of Blackwell-type hardware design challenges.
The cooling is significantly better than what you'd see on a server platform with water cooling channels going to each row of the wafer.
https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
I had guessed that Cerebras had made some trade-offs in process in order to make it work at scale, but then they aren't actually building these devices at scale (yet).
https://cerebras.ai/press-release/cerebras-systems-smashes-t...
And actually there have been attempts to do it, I mentioned in an earlier version of my comment that Gene Amdahl had attemped to make WSE work something like 20 years ago, without success - but also without the clear profitability story of AI to attract the same mountains of cash being thrown around today.
What's surprising is not that it is hard, or that it's hard as fuck, but that given the potentially stratospheric rewards for success there have not been more attempts in this direction.
IIRC cerebras' design was originally for HPC workloads, so even it may not necessarily be optimized for LLMs
> Traditional LLMs output everything they think immediately, without stopping to consider the best possible answer. New techniques like scaffolding, on the other hand, function like a thoughtful agent who explores different possible solutions before deciding. This “thinking before speaking” approach provides over 10x performance on demanding tasks like code generation, fundamentally boosting the intelligence of AI models without additional training.
[1] https://langchain-ai.github.io/langgraph/tutorials/multi_age...
https://groq.com/12-hours-later-groq-is-running-llama-3-inst...
CS-3 system is built for single node domain scaling to 24 Trillion parameter models. I.e., they claim you can run the same code without hand-written distributed training code to reach 24 Trillion parameter models.
Does Cerebras support reliable structured output like the recent OpenAI 4o?
For example, I know the latest batch of Mistral models all have json output support.
Cerebras is a startup producing innovative AI chips. Their chips are super cool, and I personally believe Cerebras is ahead of the industry and is on the right technical path. As a matter of fact, Cerebras started with HPC chips. Then pivoted to AI like everyone else.
They are still deep in the trench for survival.
Given that, they have very little software prowess compared to AMD (which has *terrible* software stack for AI GPUs look at https://github.com/ROCm/rdc, an equivalent to NVIDIA DCGM, which virtually has no maintainer, and no one is using it), NVIDIA (the golden standard of software stack for AI GPUs); and you are referring to structured output and prompt caching which are prominently developed by LLM research institutions (OpenAI Anthropic, each of which have way more funding than Cerebras)
In the end, educate yourself, and do not put unrealistic expectation on startups.
The point remain that Cerebras is not in a position to focus on structured output or prompt caching.