I too use local 7b open-hermes and it's really good.
I too use local 7b open-hermes and it's really good.
Q5 is minimum.
https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B-...
If you have the time, could you explain what you mean by "Q5 is minimum"? Did you determine that by trying the different models and finding this one is best, or did someone else do that evaluation, or is that just generally accepted knowledge? Sorry, I find this whole ecosystem quite confusing still, but I'm very new and that's not your problem.
If you're RAM constrained, you'll also have to make trade-offs about the context length. e.g. you could have 8 GB RAM and a Q5 quant with shorter context, vs Q3 with longer, etc.
[0] https://huggingface.co/TheBloke/OpenHermes-2.5-Mistral-7B-GG...
I ran this originally on a M1 with 32GB, I run this on an Air M2 with 16GB (and mac mini M2 32GB), no problem.
I use llama.cpp with a SwiftUI interface (my own), all native, no scripts python/js/web.
7b is obviously less capable but the instant response makes it worth exploring. It's very useful as a Google search replacement that is instantly more valuable, for general questions, than dealing with the hellscape of blog spam ruling Google atm.
Note, for my complex code queries at $dayjob where time is of the essence, I still use GPT4 plus, which is still unmatched imho, without running special hardware at least.