Prior to LLaMA + llama.cpp you could maybe run a large language model locally... if you had the right GPU rig, and if you really knew what you were doing, and were willing to put in a lot of effort to find and figure out how to run a model.
My hunch is that the ability to run on a M1/M2 MacBook is going to open this up to a lot more people.
(I'm exposing my bias here as a M2 Mac owner.)
I think the race is now on to be the first organization to release a good instruction-tuned model that can run on personal hardware.
And given how early the cpp port is, there is likely plenty of performance headroom with more m1/m2-specific optimization.
I suspect people saying it's not good are prompting it like ChatGPT, not realizing how much trickier a raw model is to prompt. Getting the hyperparameters for good sampling is another stumbling block. The models are very good if you do everything properly.
Personally I hope we quickly get to the stage that there's a real open llm like SD is to DALL-E. It sucks to have to bother with Facebook's core model, and give it more attention than it deserves, just because it's out there.
If facebook had actually released it as an open model, I would have said that all the credit should go to them. But instead people are doing great open source work on top of their un-free model just because it's available, and in the popular conception they're going to get credit that they shouldn't
Llama 7B is much better that something like GPT-Neo at text generation.