Groq Raises $640M to Meet Soaring Demand for Fast AI Inference
wow.groq.com
wow.groq.com
Nvidia has tons of other products, and they have been delivering AI accelators for many years now. Groq on the other hand is basically a gamble at this point.
As a developer, I'm thrilled they'll be increasing capacity.
This comment is the epitome of that considering how Groq has near zero revenue while Nvidia brings in close to $100b annually with over 50% profit margins.
This is how both numbers can be correct at the same time.
Just because Grog is starting out smaller, doesn't mean they will actually survive to 10x their current valuation.
2.8 * 10 = 28 billion. Very few companies ever reach that point.
What models does Groq run at 500tok/sec? In what mode (e.g. batch size)?
UPD. found the model, it is Mixtral 8x7B-32k. AFAIK 8xH100 will do 100+tok/sec with batch size 1.
But that does not look too impressive for batch sizes higher than 1: https://www.baseten.co/blog/faster-mixtral-inference-with-te...
As far as I know, Grog AI accelerators are really fast because they load the model into SRAM only and spread the model out over many Grog chips. They don't use HBM or off-chip RAM at all.
That doesn't seem like a defensible moat to me.
Right now, I believe Nvidia's GPUs are geared more towards training. These startups tend to compete more in inference since it's simpler. I do think Nvidia needs to shore up their inference offerings more. Actually, buying Grog or starting their own inference-only chips aren't bad ideas.
Also, when an end market is growing as fast as the AI compute market is, multiple players can do very well for years at a time. I think Groq also has a high probability of being acquired by Google/Microsoft/Apple/Meta if the FTC allows it. But they are probably better off following the hockey stick growth for a couple more years before doing that.
At a very high level, Nvidia just needs to break out the Tensor cores into its own chips, add loads of SRAM, and run CUDA. They already have all the other datacenter stuff figured out such as interconnects.
In other words, they are competing with Nvidia, not the million startups using Nvidia.