I just experiment with some local LLMs, but the differences are pretty huge:
Llama 3 8B, Raspberry Pi 5: 2-3 Tokens/second (but it works!)
Llama 3 8B, RTX 4080: ~60 Tokens/second
Llama 3 8B, groq.com LPU, ~1300 Tokens/second
Llama 3 70B, AMD 7800X3D: 1-2 Tokens/second
Llama 3 70B, groq.com LPU, ~330 Tokens/second
There seem to be huge gaps between CPU, GPU and specialized inference ASICs. I'm guessing that right now there aren't many genius-level architecture breakthroughs, and that it's more about how much memory and silicon real estate you're willing to dedicate to AI inference.
I think groq doesn't use quantization, so the gap between your hardware and groq would be even further apart.
To my knowledge this isn't (absolutely) publicly known but users on /r/LocalLLaMA and elsewhere have provided some pretty clear examples that Groq is almost certainly quantized. Which makes sense considering their memory situation...
An entire GroqRack (42U cabinet) has 14GB of RAM which means it likely can't even reasonably run llama3 8b in BF16/FP16. Let alone 70b, Mixtral, etc.
The amount of hardware required to run their public-facing hosted product likely takes up an obscene amount of floor space, even in int4. Their docs for GrowFlow describe int8 quantization but their toolkit is heavily dependent on ONNX, which has had recent tremendous work in terms of different post training quantization strategies and precisions.
However, the power efficiency vs performance is very good, potentially to the point of being able to use very cheap datacenter/co-location space that isn't capable of meeting the power and (air) cooling densities of datacenter AMD and Nvidia GPU products.
Interestingly I have access to a GroqRack system that I'm hoping to be able to spend some time on this week.
How much RAM is required for this result? It's quite impressive that it even works as well as it does.
LM Studio will also let you do partial GPU offloads, but I've only started experimenting with that. The 1-2 Tokens/second value is what I got using GPT4All.
"The Osborne effect is a social phenomenon of customers canceling or deferring orders for the current, soon-to-be-obsolete product as an unexpected drawback of a company's announcing a future product prematurely. It is an example of cannibalization."
You're right on the Osborne effect though! Thanks for that. We are definitely not doing that.
To clarify: When we started, MI300x was not officially announced yet, so we were planning on buying MI250's. Due to everything taking longer than expected around starting the business and receiving funding, by the time we had money in the bank, it was time to buy MI300x. Going forward, we are buying MI300x today and will continue to buy AMD MI series as they are released in the future.
I've been interested to give 8x MI300x a try, as they are supposed to be cheaper per FLOPs, but it looks like your service does not provide on-demand pay-per-second instances. Any plans to change that?
It kind of makes sense since their history is only supporting the high end GPUs in their HPC solutions, where they don't use VM's. They've committed to us directly that they will fix this issue.
I updated our pricing page to note this.
Nice job on your website!
Sincerely,
Someone who recently gave you a hard time for not having info on your website.
Thank you for the response!
Renting 8 GPUs at once is fine and desired; it's not the issue.
The issue is that right now one has to commit to at least 1 week of use; my use patterns are bursty and it does not map well to the current proposition.
Right now, we are trying to attract people who are a mix of wanting to kick the tires on a new product, as well as take compute for longer term.
I did mention in the pricing section that we can store your data locally, as part of the advertised pricing. This is our effort to recognize your use case.
I also understand that you want to optimize and don't want to pay for something that you're not using. We will eventually get to that point, but honestly just not there yet.
Think also from our end, we have these GPUs and if you're not using them... then who is? We've put out the capex/opex to make these available to you at any time, so the only way to be efficient on our side is to do a week long block right now.
Regardless, if you want to reach out to me directly, please do so. Maybe there is a middle ground we can both work from. Happy to consider all options and getting in early with us will always have first mover advantages.
So really, they lose nothing. They've already booked sales of everything there is to sell. So might as well now turn attention to those who might be customers two years from now, and make them feel like the wait will be worth it.