Run Mistral 7B on M1 Mac
wandb.ai
wandb.ai
Disclaimer: I have a competing universal macOS/iOS app[3] that does support SWA with Mistral models (using mlc-llm).
[1]: https://github.com/ggerganov/llama.cpp/issues/3377
[1]: https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1/...
[2]: https://huggingface.co/mistralai/Mixtral-8x7B-v0.1/commit/58...
I can run orca-mini:3b, It's dumber than just flipping a coin but it still feels cool to have a LLM running on your own computer.
Takes 5 minutes really.
Runs on GPUs, uses about 5GB VRAM. On integrated GPUs generates 1-2 tokens/second, on discrete ones often over 20 tokens/second.
Replacement of CUDA with Direct3D 11This makes these compute kernels relatively portable across GPU APIs and languages.
I think on Windows, Direct3D is often a better tech choice for GPGPU. (1) It runs on GPUs of all vendors (2) Direct3D is an essential Windows component; the OS even composes windows on desktop with D3D11. Unlike CUDA, D3D runtime libraries are already available in the OS by default, which simplifies software installation and support.
Anyone got suggestions? Would love for it to run summarization over GBs of ncbi data.
I wanted to try how local models compare with that use case of mine.
The only one that worked decently was Mixtral
/edit this was with llama cpp though not ollama
Thanks!
Though if your machine can’t keep it all in memory, then speed will still fall off a cliff.
Would it be advantageous to divide the load and run on both?
It's very easy to port it to Python requests or a similar HTTP library, at the least.
I agree that since Ollama provides a nice REST API, why not use that.