Why?
For each new word a transformer generates it has to move the entire set of model weights from memory to compute units. For a 70 billion parameter model with 16-bit weights that requires moving approximately 140 gigabytes of data to generate just a single word.
GPUs have off-chip memory. That means a GPU has to push data across a chip - memory bridge for every single word it creates. This architectural choice, is an advantage for graphics processing where large amounts of data needs to be stored but not necessarily accessed as rapidly for every single computation. It's a liability in inference where quick and frequent data access is critical.
Listening to Andrew Feldman of Cerebras [0] is what helped me grok the differences. Caveat, he is a founder/CEO of a company that sells hardware for AI inference, so the guy is talking his book.
[0] https://www.youtube.com/watch?v=MW9vwF7TUI8&list=PLnJFlI3aIN...
I wish I could say more about what AMD is doing in this space, but keep an eye on their MI4xx line.
That is curious. Things are moving so quickly right now. I typed out a few speculative sentences then went ahead and asked an LLM.
Looks like Cerebras is responding to the market and pivoting towards a perceived strength of their product combined with the growth in inference, especially with the advent of reasoning models.
From die shots and materials I’ve seen, it even looks like ~40% of the die might be allocated to memory [1]. Given that, I’m curious about your point on “not enough die for memory” — is it a matter of absolute capacity still being insufficient for current model sizes, or more about the area-bandwidth tradeoff being unbalanced for inference workloads? Or perhaps something else entirely?
I’d love to understand this design tension more deeply, especially from someone with a high-level view of real-world deployments. Thanks again.
[1] Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Fig. 5. Die photo of 14nm ASIC implementation of the Groq TSP. https://groq.com/wp-content/uploads/2024/02/2020-Isca.pdf
This. Additionally, models aren't getting smaller, they are getting bigger and to be useful to a wider range of users, they also need more context to go off of, which is even more memory.
Previously: https://news.ycombinator.com/item?id=42003823
It could be partially the DC, but look at the rack density... to get to an equal amount of GPU compute and memory, you need 10x the rack space...
https://www.linkedin.com/posts/andrewdfeldman_a-few-weeks-ag...
Previously: https://news.ycombinator.com/item?id=39966620
Now compare that to an NV72 and the direction Dell/CoreWeave/Switch are going in with the EVO containment... far better. One can imagine that AMD might do something similar.
https://www.coreweave.com/blog/coreweave-pushes-boundaries-w...
What I’m still trying to understand is the economics.
From this benchmark: https://artificialanalysis.ai/models/llama-4-scout/providers...
Groq seems to offer near lowest prices per million tokens and the near fastest end to end response times. That’s surprising because in my understanding, speed(latency) and the cost are trade-offs.
So I’m wondering: Why can’t GPU-based providers can't offer cheaper but slower(high-latency) APIs? Or do you think Groq/Cerebras are pricing much below cost (loss-leader style)?
https://www.datacenterknowledge.com/data-center-chips/ai-sta...
https://www.semafor.com/article/12/03/2024/amazon-announces-...
What even is an AI data center? are the GPU/TPU boxes in a different building than the others?
Google does many pieces of the data center better. Google TPUs use 3D torus networking and are liquid cooled.
> What even is an AI data center?
Being newer, AI installations have more variations/innovation than traditional data centers. Google's competitors have not yet adopted all of Google's advances.
> are the GPU/TPU boxes in a different building than the others?
Not that I've read. They are definitely bringing on new data centers, but I don't know if they are initially designed for pure-AI workloads.
And I'll echo, what even is an AI data center, because we're still none the wiser.
A data center that runs significant AI training or inference loads. Non AI data centers are fairly commodity. Google's non-AI efficiency is not much better than Amazon or anyone else. Google is much more efficient at running AI workloads than anyone else.
I don't think this is true. Google has long been a leader in efficiency. Look at the power usage effectiveness (PUE). A decade ago Google announced average PUEs around 1.12 while the industry average was closer to 2.0. From what I can tell they reported a 1.1 average fleet wide last year. They've been more transparent about this than any of the other big players.
AWS is opaque by comparison, but they report 1.2 on average. So they're close now, but that's after a decade of trying to catch up to Google.
To suggest the rest of the industry is on the same level is not at all accurate.
https://en.wikipedia.org/wiki/Power_usage_effectiveness
(Amazon isn't even listed in the "Notably efficient companies" section on the Wikipedia page).
We've seen the rise of OSS Kubernetes and eBPF networking since, and a lot more that I don't have on-stack rn.
I wouldn't be surprised if everyone else had significantly closed the hardware utilization gap.
That said, the torus approach was a gamble that most workloads would be nearest-neighbor, and allreduce needs extra work to optimize.
An AI data center tends to have enormous power consumption and cooling capabilities, with less disk, and slightly different networking setups. But really it just means "this part of the warehouse has more ML chips than disks"
Thank you very much, that is the piece of the puzzle I was missing. Naively, it still seems (to me) far more hops for a 3d torus than a regular multi-level switch when you've got many thousands of nodes, but I can appreciate it could be much simpler routing. Although, I would guess in practice it requires something beyond the simplest routing solution to avoid congestion.
No one else has access to anything similar, Amazon is just starting to scale their Trainium chip.
The end of Moore's law pretty much dictates specialization, it's just more apparent in fields without as much ossification first.