You could use CPU inference on a smaller local model (either always or after a demo budget is spent).
Don't forget to put up a "donate" button.
Don't forget to put up a "donate" button.
CPU inference for LLMs takes forever (you'll get like 1tk/s on CPU) and limits you significantly in terms of model size/quality. You'll lock up all of your cores to provide service for a single user at a snail's pace.
I don't think it should even be considered as an option
I’d rather not offer a demo at all than offer it with these parameters.