Ask HN: Running LLMs Locally
What's the best way to build an app around local LLMs. How can I educate m myself on this?
https://github.com/ExtensityAI/symbolicai?tab=readme-ov-file...
I think what you’ll find is that some applications are very capable locally, like Whisper.
A lot of plugins expect to work with the llama.cpp family. Nowadays, that’s HuggingFace TGI: https://huggingface.co/blog/tgi-messages-api
So your application could speak OpenAI api, and you’d run HuggingFace TGI on your hardware for testing and comparison.
[1] https://huggingface.co/docs/text-generation-inference/en/sup...
[2] https://github.com/Ankur3107/transformers-on-macbook-m1-gpu