I wish they would do a behind the scenes on how much money, time, optimisation is done to make this all work.
Also big fan of anyscale. Their pricing is just phenomenal for running models like mixtral. Not sure how they are so affordable.
I wish they would do a behind the scenes on how much money, time, optimisation is done to make this all work.
Also big fan of anyscale. Their pricing is just phenomenal for running models like mixtral. Not sure how they are so affordable.
Builds very quickly with make. But if it's slow when you try it then make sure to enable any flags related to CUDA and then try the build again.
A key parameter is the one that tells it how many layers to offload to the GPU. ngl I think.
Also, download the 4 bit GGUF from HuggingFace and try that. Uses much less memory.
1. https://www.semianalysis.com/p/inference-race-to-the-bottom-...
Even with 12 threads of my 5900X (I've tried using the full 24 SMT - that doesn't really seem to help) with the dolphin-2.5-mixtral-8x7b.Q5_K_M model, my MacBook Pro is around 5-6x faster in terms of tokens per second...
But whatever it is, it's great, and I hope that Intel and AMD will catch up.
AMD has had the APUs for awhile but I think they aren't at the same level at all as the new Mac acceleration.
Obviously still a nascent area but https://lmsys.org/blog do a good job of diving into engineering challenges behind running these LLMs.
(I'm sure there are others)