Can someone explain why these types of projects are compelling (running an LLM in the user's browser?)
I guess privacy and ease-of-use?
So you don't need to download llama.cpp, and run it in the terminal or something?
I guess privacy and ease-of-use?
So you don't need to download llama.cpp, and run it in the terminal or something?
Actually, given dropbox's deterioration over time, FTP+SVN is sounding pretty good to me right now.
Its also fairly easy to route a Flask server to these models with websockets, so with that I've been able to run python and pass data to the model to run on the GPU and pass the response back to the program. Again, there's probably a better way but its cool to have my own personal API for a LLM.