I actually wrote about getting an LLM chatbot up and running a while ago: https://blog.kronis.dev/tutorials/self-hosting-an-ai-llm-cha...
It's good that the technology and models are both available for free, and you don't even need a GPU for it. However, there are still large memory requirements (if you want output that makes sense) and using CPU does result in somewhat slower performance.
There are async use cases where it can make sense, but for something like autocomplete or other near real time situations we're not there yet. Nor is the quality of the models comparable to some of the commercial offerings, at least not yet.
So I don't have it in me to blame anyone who forks over the money to a SaaS platform instead of getting a good GPU/RAM and hosting stuff themselves.
Here's hoping the more open options keep getting better and more competitive, though!