There are some newer models out there I have not tried yet like the open llama, GPT4all, etc. so I’m not sure how good they are. I get the sense they are still GPT-3 level but are achieving that with less RAM.
There’s a race on both for raw capability and optimization via pruning and quantization. The latter is important to make these things runnable locally without gigantic hardware. Lots of people have stuff with GPUs, fast CPUs, and 32-64G RAM. Few have huge workstations with hundreds of gigs of RAM.
Unless progress stagnates I can see something approaching GPT-4 that can run on under $5k worth of hardware in a year or so.
Open model progress seems to be lagging only 1-2 years behind big cloud hosted models.
Probably the fastest way to get started is to look into [0] - this only requires a beta chromium browser with WebGPU. For a more integrated setup, I am under the impression [1] is the main tool used.
If you want to take a look at the quality possible before getting started, [2] is an online service by Hugging Face that hosts one of the best of the current generation of open models (OpenAssistant w/ 30B LLaMa)
[0]: https://mlc.ai/web-llm/ [1]: https://github.com/oobabooga/text-generation-webui [2]: https://huggingface.co/chat