Run the strongest open-source LLM model: Llama3 70B with just a single 4GB GPU
ai.gopubby.com
ai.gopubby.com
It doesn't mention quantization either.
I wish models sizes were published in bytes.
Well I'm definitely worried about recall and all the Microsoft nonsense, I really want to be able to run and train LMMs, and other machine learning frameworks locally.
> Of course, it would be more reasonable to compare the similarly sized 400B models with GPT4 and Claude3 Opus
No. It's completely irrelevant to the topic of the article.
The article is mostly a press release for llama 3. It also contains a few comments by the author, they aren't bad but don't save the clickbaity, buzzy, sensationalist core.
As a bonus; I also use this for a personal project where I use prompts and Llama3 to control smart devices. JSON responses from the LLM are parsed and translated into the smart device commands from a raspberry pi. I control it using speech via my Apple Watch and Apple shortcuts to the raspberry pi’s api. It all works magically and fast. Way faster than pulling up the app on my phone. And yes the LLM is smart enough to control groups of devices using simple conversational AI.
edit; here's a demo https://www.youtube.com/watch?v=dCN1AnX8txM
Considering this says it's "not designed for real-time interactive scenarios" it's probably also really slow
"Write a haiku about Hacker News mentioning AI in the title"
Here is a haiku:
AI whispers secrets
HN threads weave tangled debate
Intelligence born
eval time = 30363.04 ms / 23 runs ( 1320.13 ms per token, 0.76 tokens per second)
total time = 34294.80 ms / 33 tokens