The past week has felt like a wake-up call to enthusiasts. Running models locally has been available for a while (even small, fairly coherent ones), and the majority of "improvements" recently have come from implementing the leaked LLaMa model.
The results from 7B are an improvement on what we had a year ago, but not by much. We're learning that there's room to optimize these models, but also that size matters. ChatGPT and 7B are both great at bullshitting, but you can feel the difference in model size during regular conversation. Adding insult to injury, it will almost always be faster to query an API for AI results than it will be to run it locally.
Analysis: Things are moving at a clip right now, but people expecting competitive LLMs running locally on their smartphone will be disappointed for quite a while. As the technology improves, it's also safe to assume that we'll find ways to scale model intelligence with greater resources, and the status quo will look much different than it does today.