Mistral LLM on Apple Silicon Using Apple's MLX Framework Runs Instantaneously
github.com
github.com
"A fire built around a wad of kerosene soaked rags starts instantly."
Does that sentence suggest that kerosene starts rags to you, or that kerosene soaked rags help start fires instantly?
English, being a very complex language, can have the same sentence mean multiple things as well depending on how you interpret the context. There may be better ways to phrase something to remove ambiguity, but the title as presented could be read both ways.
My app also uses a very small (30MB) PyTorch model and shipping it requires an extra 100MB for PyTorch in the app. Very very stupid.
I think its important to remember that last mile inference is still pretty bespoke for most things. If we want to see gen AI stuff take off and now have the big cloud providers in charge this needs to be fixed.
Apple is in a good place to solve at least part of the equation.
Apple will definitely push for on-device AI, but even in 2030 I firmly believe that they won't be leading the industry in performance. I'd be surprised if they even supported anything other than their proprietary CoreML by then.
A100: 1248 TOPS
MI250: 362.1 TOPS
M3 Max: 18 TOPS
Yes, 18. Unless Apple has accelerated INT4 workloads but just forgot to document it.
Honestly, I’m an Apple fan, but when they go on stage and say “AI” they mean it can do speech recognition or tell a dog apart from a cat, or autofocus a camera. It can’t run ChatGPT-like things by a loooong mile.
I just went through all the commercial options for local LLM hosting and they're definitely poised correctly because they have the correct amount of memory in their machines.
they're not going to have a come to Jesus moment.
The thread suggests it doesn't even quantize the model (running it in FP16, so tons of ram usage), and that its slower than the llama.cpp Metal backend anyway?
And MLC-LLM was faster than llama.cpp, last I checked. Its hard to keep up with developments.
- B: Yeah, llama.cpp has a killer feature set. And killer integration with other frameworks. MLC is way behind, but is getting more fleshed out every time I take a peek at it.
- C: This is a pet peeve of mine, but I've never run into a local model that was really uncensored. For some, if you give them a GPT4 prompt... Of course you get a GPT4 response. But you can just give them a unspeakable system prompt or completion, and they will go right ahead and complete it. I don't really get why people fixate on the "default personality" of models trained on GPT4 data.
Personally I run my own Yi 200K DARE merge because I love the long context.
In any case, I had fun with MLX today, and I hope it implements 4 bit quantization soon.