EDIT: This a 1W light bulb moment for me, thank you!
EDIT: This a 1W light bulb moment for me, thank you!
Macs are so good at it because Apple solder the memory on top of the SoC for a really wide and low latency connection.
Then the model needs to be properly quantized and formatted for GGUF (the model format that llama.cpp uses), tested, and uploaded to the model registry.
So there's some length to the pipeline that things need to go through, but overall the devs in both projects generally have things running pretty smoothly, and I'm regularly impressed at how quickly both projects get updated to support such things.
Same! Big kudos to all involved
Also, I'm not sure if we'll call it mistral-nemo or nemo yet. :-D
I think it should also run well on a 36GB MacBook Pro or probably a 24GB Macbook Air
If you're on a Mac, check out LM Studio.
It's a UI that lets you load and interact with models locally. You can also wrap your model in an OpenAI-compatible API and interact with it programmatically.
and expose an OpenAI compatible server, or you can use their python bindings