Llama3 running locally on iPhone 15 Pro
imgur.com
imgur.com
https://apps.apple.com/app/private-llm-local-ai-chatbot/id64...
The 8B model nominally works on 6GB phones but it's quite slow on them. OTOH, it's very usable on iPhone 15 Pro/ Pro Max devices and even better on M1/M2 iPads.
Every framework: llama.cpp, MLX, mlc-llm (which I use) all only use the GPU. Using the ANE and perhaps the undocumented AMX coprocessor for efficient decoder only transformer inference is still an open problem. I've made early some progress on quantised inference using ANE, but there 're still a lot of issues to be solved before it is even demo ready, let alone a shipping product.
They've been stingy on increasing RAM compared to Android phones.
The ram tax so absurd it's bordering on criminal, but it also just seems stupid, because if they hadn't put 8gb's of ram in the new smallest macbook air m2 their whole lineup would be more than capable at running local quality LLM's, or double their gaming devices because of their awesome chipset giving them 16gb's of vram essentially, but no, not now when 25% have low ram, ie. no new OS LLM updates.
Also we can't have gaming because half their new sold devices have shit ram, so they also kind of already ditched their "gaming" plan they just got started on a year ago - all because they wan't to push products with ram levels from 10 years ago - bizarre!
They must be betting on local AI as a "pro" feature only.
Most importantly, though, we are talking about iPhones here. I can’t say I’ve ever thought to myself “gosh, I wish my phone had more RAM!” in…over a decade?
That was… *checks CV*… I left in April 2015.
I think RAM is like roads: usage expands to fill available infrastructure/storage.
That an iPhone today has as much RAM as the still-functioning Mid-2013 MacBook Air sitting in a drawer behind me is surprising when compared to the 250-fold growth from my Commodore 64 to my (default) Performa 5200… but it doesn't seem to have actually harmed anything I care about.
Swimming in badly written SPAs and cordova/whatever hybrid apps is seriously helped by eg 12GB of RAM on a mobile :)
Honestly I'm glad they don't. The PC is the last open platform out there and the last thing I'd want to see is Apple encroaching on it with their walled gardens and carbonite-encased computers.
... in any of their products.
FTFY
So yes, there is tremendous marketing value from that low starting price, although I think it's nearing the end of it's usefulness now that even fan sites are starting to call out the inadequacy.
I had a Macbook with 8GB RAM and 256GB disk as my daily driver for work until last year running Docker and my fat IDE without too many issues. It's a similar story with my phone - I bought the bigger storage version because I thought I'd need it but after 3 years of using it I'm still not close to even using 128GB.
It's almost certain that the iPhone 16 will ship with 8GB of RAM. What needs to be seen is whether iPhone 16 Pro and Pro Maxes will ship with 16GB of RAM (Like with high end M1/M2 iPad Pros with >= 1TB SSD).
12 Pro Max had 6GB RAM
15 Pro Max has 8GB RAM
16 most likely will have between 8GB and 12GB RAM.
The chat is answering at a speed of one word per few/several seconds. But still, this a nice feat.
Example recording for the curious: https://www.youtube.com/watch?v=nZEvUj-QTrI
"Next level: QLoRA fine-tuning 4-bit Llama 3 8B on iPhone 15 pro.
Incoming (Q)LoRA MLX Swift example by David Koski: https://github.com/ml-explore/mlx-swift-examples/pull/46 works with lot's of models (Mistral, Gemma, Phi-2, etc)"
APK download link: https://github.com/mlc-ai/binary-mlc-llm-libs/releases/downl...
Similarly, I am for the first time going to care about how much RAM is in my next iPhone. My iPhone 13's 4GB is, as in your case, inadequate.
And does anyone know how many tokens per second it can run?
Does anything similar exist for Android?
Increase the Ctx to more than 100 & link to any Q4 GGUF of 7B
Says device is unsupported :(