Can't have a stable diffusion moment if you refuse to release the weights to the general public. Stable diffusion only got to where it is because 10,000 people with otherwise zero reputation were able to play around with the code and models.
LLaMA is still only available to the elite.
It was released on 4chan recently :)
files_catbox_moe[slash]o8a7xw(dot)torrent
It would not be a HN dream for long, given the implications it has for the internet
lol good luck running a 13B model on a single GPU
You need a RTX 3090 24gb
Seeing the performance of implementations like FlexGen [1], I don't think it would be entirely unreasonable to run a 13B model on a single GPU for personal usage purposes. You are not going to a run a public service out of it, but it probably would be good enough to run your own ChatGPT or Copilot locally.