There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding
There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding
Who knows if they will subsidizes it to mitigate sticker shock, but it's a safe assumption that it will be scarily expensive. However if you are in a "cost is no obstacle, speed is god" position, it will likely be pure magic.
I don't think I understand why they aren't leveraging the increased speed to do batching to serve more customers at a "normal" tok/s.
Is the limitation, even on cerberus, still that the cache can only serve so many concurrent sessions over time? Is there no scaling advantage? I genuinely do not understand how any of this works.
Also worth looking into how they do cooling for it, because that's kind of absurd and awesome as well.
Batching works because of severe memory bottleneck, but Cerebras whole thing is serving models out of "L1 cache" (?).
So, afaik, Cerebras are optimizing for ultra-low latency batch=1 inference.
https://newsletter.semianalysis.com/p/cerebras-faster-tokens... goes quite in-depth.
There's some technical hypotheses about it that other people are offering.
But also from a business perspective, it totally makes sense not to go any sort of batching play. It's really valuable and very clear to consumers to make your pitch entirely about lower latency rather than higher bandwidth.
There are so many scenarios that are latency-constrained that will be difficult or even impossible for someone even with fleets of high-bandwidth compute to compete with you on.
Very easy pitch to sell a customer who asks what differentiates you from other companies: you pay us a premium for lower latency than anyone else.
With the tool calls that can be done, you're not pricing this against an executive assistant or pocket analyst - you're pricing this against the ability to have an entire Bourne Identity style analysis room at your disposal. The limited inventory will go to the people for whom money is no object.
Otherwise it’s just lazy. I know shallow dismissals is kind of HN’s thing, but come on, a little effort please. Currently, your comment is just as much slop
One way large models are served on a bunch of cerebras chips is by essentially distributing layers' weights across chips. Few layers's weights per chip - as many as the KV cache + activations + weights will allow. You use pipelining to hide the latency of the inter-chip 150 GB/s link.
On GPUs, you amortize the cost of loading weights from HBM to SRAM across multiple users - thereby making it cheaper _per_ user. But here, there is no such amortization. The weights are already there. It is the activations that stream through.
You _could_ do batching/continuous batching, but that would just service more users at lower token/s each without any amortization of fixed cost, due to fixed cost (loading weights) being non-existent.
When you think about it, it would still be dirt cheap compared to normal way of doing things. In the old days, if you had an outage on a serious user facing system, you'd wake up people across various timezones, wake up their managers and scramble to find the root cause, identify a solution, brainstorm on possible side effects of a fix, and then rush to build it and deploy. This cycle would involve, sometimes, dozens of people, for, say, 10 man hours each. So lets make it 120 man hours per serious outage, and lets assume and average of $100 per hour - so, $12,000 per a serious outage fixed under a day, counting conservatively and not including the costs of the actual outage.
I'd guess the pricing for those ultrafast, very energy inefficient and hardware heavy models will be competing with that. Its going to be possible to get a fix out in 30 minutes, 10 of which will be tests, 5 will be the deploy, and the remaining 15 will be some unlucky guy trying to keep up with the super fast model throwing a 50 "load-bearing deferrals earning their keep" per minute :-)
The pricing on those things is competing with costs to run entire departments. I'd, for one, imagine offshore ops teams will be a thing of the past in under a year, since one gets way better initial response to anything from a model, given right setup, esp. on codebases that have been built from the ground up with agentic coding - so with good documentation and effective test coverage baked into repos.