I ran the 7B Vicuna (ggml-vic7b-q4_0.bin) on a 2017 MacBook Air (8GB RAM) with llama.cpp.
Worked OK for me with the default context size. 2048, like you see in most examples was too slow for my taste.
Worked OK for me with the default context size. 2048, like you see in most examples was too slow for my taste.
OpenAIs paid GPT4 has few restrictions and is still cheap.
... Not to mention GPT4 with browsing feature is vastly superior to any home of the models you can run at home.
For LLMs this means I am allowed their full potential. I can generate smut, filth, illegal content of any kind for any reason. It’s for me to decide. It’s empowering, it’s the hacker mindset.