I have some ideas I’d like to release, but LLM api pricing and sudden traffic from sites like this one seem scary.
I have some ideas I’d like to release, but LLM api pricing and sudden traffic from sites like this one seem scary.
I'm considering switching to DeepSeek since it's way cheaper. I'll swap once I'm done testing out the API. You can use your own hosted LLM but it's not worth it at this moment.
That’s the tricky part about building LLM apps. I’d love to hear more from Indie devs because money is absolutely a bottle neck here.
For fun:
You don’t need an LLM for some of your calls I think. “Where is the Eiffel Tower”, Eiffel Tower is a NER that small NLP libraries can extract. Then it’s a simple long/lat lookup. You might be able to re-route 20% of your calls to a no-cost backend call.
Don't forget to put up a "donate" button.
CPU inference for LLMs takes forever (you'll get like 1tk/s on CPU) and limits you significantly in terms of model size/quality. You'll lock up all of your cores to provide service for a single user at a snail's pace.
I don't think it should even be considered as an option
I’d rather not offer a demo at all than offer it with these parameters.