Now that I've seen that LLaMA can work I'm confident someone will release an openly licensed instruction-tuned model that works on the same hardware at some point soon.
I also expect that there are prompt engineering tricks which can be used to get really great results out of LLaMA. I'm hoping someone will come up with a good prompt to get it to summarization, for example.
For a model running on my own laptop I'm OK taking more risks. I'd like it to be able to obey simple instructions like "Summarize this text" or "Extract the names of everyone mentioned in this article" - I don't care as much about the stuff ChatGPT has to get right.
Enough programmers want this badly enough that its going to happen. Inference at 8 GB and fine tuning at 24 GB, just like stable diffusion, on a 13B model.