I've talked to a couple dozen people in real time who've played with up to 30B but no one I know has the resources to run the 65B at all or fast enough to actually use and get an opinion of. None of the open source llama projects out there are using 65B in practice (despite support for it) so I think my 30B and under conclusions are applicable to the topic the article covers. I'd love to be wrong and I'm excited for this to change in the future.
Admittedly, it's quite slow and therefore not useful for chatting or real-time applications, and it's unreliable enough in its quality that I'd like to be able to iterate faster. Definitely more of a toy at this point, at least when run on CPU.
I should probably go back and try again to see if it's worth it for the extra speed, now that I've played with 65B for a while.
That said, it’s clear that replicating GPT4+ performance is within the resources of a number of large tech orgs.
And the smaller models can definitely still be useful for tasks.
LLaMA incorporated new techniques that make 65B perform way better than GPT-3's 175B so the model size argument is not very strong.
I'd agree the secret sauce for how great the newest services perform is probably in the fine-tuning. We're seeing almost daily releases of fine-tuning data sets, training methods and models (at lower and lower costs) so I'm personally pretty optimistic that we'll be seeing some big improvement in self-hosted LLM performance pretty quickly.
[1] https://ar5iv.labs.arxiv.org/html/2302.13971#:~:text=Table%2....
My experience here is pretty similar. I'm heavily (emotionally at least) invested in models running locally, I refuse to build something around a remote AI that I can only interact with through an API. But I'm not going to pretend that LLaMA has been amazing locally. I really couldn't figure out what to build with it that would be useful.
I'm vaguely hoping that compression actually gets better and that targeted reinforcement/alignment training might change that. GPT can handle a wide range of tasks, but for a smaller AI it wouldn't be too much of a problem to have a much more targeted domain, and at that point maybe the 30B model is actually good enough if it's been refined around a very specific problem domain.
For that to happen, training needs to get more accessible though. Or communities need to start getting together and deciding to build very targeted models and then distributing the weights as "plug-and-play" models you can swap out for different tasks.
And if there's a way to get 65B more accessible, that would be great too.
For 65B, GPTQ 4-bit should fit LLaMA 65B into 40GiB of memory. Currently the cheapest way to run that at an acceptable speed would be to use 2 x RTX 3090/4090s (~$2500-3000) or maybe a Jetson Orin 64GB (~$2000). I've seen people trying to run it on an M1 Max and it's just a bit too slow to comfortably use (I get a similar speed to when I try it on my 5950X - about 1-2 tokens/s), but it seems like it's within a factor or two of being fast enough, so not out of the question that it might get there just through software optimizations. I'd definitely upgrade to a 7950X/X3D or a Threadripper (w/ 96GB of DDR5-5200) if I could get 65B running at a comfortable speed all the time.
I think training is also advancing at a pretty good clip. LLaMA-adapter [2] is doing fine tuning of LLaMA 13B on a single 8xA100 system in 1h (so for ~$12 for a spot instance) and was already over 3X faster than Alpaca's training.
To me, the biggest thing limiting easy plug-and-play distribution is actually LLaMA's licensing issues, so maybe someone will offer a better open foundational model soon and the community can standardize on that. It'd be nice to have a larger context window (Flash Attention?) as well.
[1] https://github.com/facebookresearch/llama/blob/main/MODEL_CA...
a like-for-like comparison would be GPT-4 against the larger models like LLaMA 65B, but those cannot be run on consumer-grade hardware
so one ends up comparing the stuff one can run... against the top stuff from OpenAI running on high-end GPU farms, and this technology clearly benefits a lot still from much larger scale than most people can afford
the great revelation this year is how much does it get better as it get much, much bigger without a clear horizon on where will diminishing returns be hit
but at the same time, some useful stuff can be done on consumer hardware - just not the most impressive stuff
GPT-3.5 OTOH is much better, but it's also much better at producing convincing-sounding but completely incorrect answers
The output is unremarkable; it’s not significantly better than the 13B model for most uses.
GPT 3.5 is an order of magnitude better at least.
To run it properly you need a lot more than a Mac Studio, and then comparisons need to be done more or less seriously, not just a few random prompts, because anything in a black box will "cheat" and will be fine tuned to do well at popular benchmarks.