Ignoring the licensing issues, there are a few other constraints that would make the model harder to go viral outside of developers who spend a lot of time in this space already:
1) Model weights are heavy for just experimentation, although quantizing it down to 4-bit might make them on par with SD FP16.
2) Requires extreme CLI shenanigans (and likely configuration since you have to run make) compared to just running a Colab Notebook or a .bat Windows Installer for the A1111 UI.
3) Again hardware: a M1 Pro or a RTX 4090 is not super common among people who are just curious about text generation.
4) It is possible the extreme quantization could be affecting text output quality; although the examples are coherent for simple queries, more complex GPT-3-esque queries might become relatively incoherent. Particularly with ChatGPT and its cheap API (timely!) out now such that even nontechies have a strong baseline on good output already. The viral moment for SD was that it was easy to use and it was a significant quality leap over VQGAN + CLIP.
I was going to say inference speed since that's usually another constraint for new LLMs but given the 61.41 ms/token cited for the 7B model in the repo/your GIF, that seems on par with the inference speed from OPT-6.7B FP16 in transformers on a T4.
Some of these caveats are fixable, but even then I don't think LLaMA will have its Stable Diffusion moment.