Smaller models are useless for me, because my native language is Ukrainian (it's easier to spot mistakes made by model in a language with more complex grammar rules).
As GUI, I use Page Assist[3] plugin for Firefox, or aichat[4] commandline and WebUI tool.
[1]: https://github.com/ollama/ollama/releases
[2]: https://ollama.com/
However, as far as I can tell, it's never actually clear what the hardware requirements are to get these to run without fussing around. Am I wrong about this?
I personally use an uncensored version which is another huge benefit of a local model. Mainly because I have many kinky hobbies that piss off cloud models.
It's slowly getting there.
For running them, you want a GPU. The limitation is that the model fits in VRAM or the performance will be slow.
But if you don't care about speed, there's more options.
Was playing with them some more yesterday. Found that the 4bit ("q4") is much worse then q8 or fp16. Llama3.1 8B is ok, internlm2 7B is more precise. And they all hallucinate a lot.
Also found this page, that has some rankings: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_...
In my opinion they are not really useful. Good for translations, to summaries some texts, and.. to ask in case you forgot some things about something. But they lie, so for anything serious you have to do your own research. And absolutely no good for precise or obscure topics.
If someone wants to play there's GPT4All, Msty, LM Studio. You can give them some of your documents to process and use as "knowledge stacks". Msty has web search, GPT4All will get it in some time.
Got more opinions, but this is long enough already.