Llamafile is the new best way to run a LLM on your own computer
simonwillison.net
simonwillison.net
Llamafile lets you distribute and run LLMs with a single file - https://news.ycombinator.com/item?id=38464057 - Nov 2023 (273 comments)
ExLLAmA v2 + elx2 quantization, and maybe tensorrt-llm might be the contender for the top performer
I can run 120b models at 3bpw.
The larger models like this are less affected (increased perplexity) by the quantization.
Panchovix/goliath-120b-exl2 (there's a different branch for each size)
Some of them I've had to do myself eg. I wanted a Q2 GGUF of Falcon 180b
There's a guy on huggingface called "TheBloke" who does GGUF, AWQ and GPTQ for most models. For exl2, you can usually just search for exl2 and find them.
And AFAIK the Ryzen pro do have an "AI coprocessor" but unsure of the API to use them.
ollama has a way to do this and lets you play with a bunch of models without being very smart (just do ollama pull <model-name> and it downloads a model and makes it available to ollama run/ollama serve).
(Sounds like I also need to play with LM Studio).
But it's far from best for people that actually want to explore LLM and play.
> You can also download a much smaller llamafile binary from their releases, which can then execute any model that has been compiled to GGUF format
- Download and launch the app
- Use their interface to huggingface models, download the one you want
- Go to the chat window and select load
- Chat away w/ a ChatGPT like interface that includes markdown processing and easy copying of code blocks, etc
Outside of the time to download a model, it's about 30 seconds of work to get up and running
I'll give it a shot over the weekend but if anyone knows I'm curious!
The good news is that the article describes in detail how to determine for yourself if it's worth it. The author uses an M2 as well, so you can at least know it'll likely work.
Something like a Mistral-7B Dolphin finetune actually is surprisingly useful, like GPT-3.5 in some respects. I imagine it would render ~10 tok/sec on an M2 for at least short bursts until throttling sets in.
As I mentioned in another post, try this out with LM Studio. Super simple GUI that even a non-tech person could probably figure out for finding/downloading/loading models w/ a ChatGPT like interface for chatting.
Which LLM to download for general questions/info that has good accuracy?
Separately, I couldn’t figure out how to use the NVIDIA GPU with LM Studio. Only wants to use the Intel card. I use NVIDIA with Diffusion with no problems if that makes any difference.