LeCun: Qualcomm working with Meta to run Llama-2 on mobile devices
twitter.com
twitter.com
Keep in mind “mobile devices” extends past just smartphones, onto wearables/headsets as well.
Ah but is it in (insert any ad business big-tech company)'s best interests to do this? Do they already have so much data that these user <> LLM interactions are only marginal that they don't care about harvesting the data ?
They have amazing chips, but Siri has been a subpar assistent since forever. Now is the time to redeem her.
I wonder what the numbers actually are for local compute on custom hardware compared to firing up the wifi antennae to make the remote request.
For comparison, a quantized mobilenetv2 takes 3.7MB disk space to solve the same image understanding task as the old VGG Caffe models which take 550MB. We’ve come a long way in six years.
> that it would be massive overkill to reset all settings (lose all wifi passwords, etc.).
I don't think that's what anyone was recommending...
Settings -> General -> Transfer or Reset iPhone -> Reset -> Reset Keyboard Dictionary is almost certainly what they were recommending.
What does resetting your keyboard dictionary have to do with your wifi passwords?
None of that explains how resetting everything would have any affect on capitalization of names if resetting the keyboard dictionary wouldn't, and you didn't say whether you tried resetting keyboard dictionary.
But the autocorrect does account for the contact list. I can personally vouch for that, as it works properly on my phone.
> > I would expect the learned dictionary to be prioritized
> After all, if the dictionary is 'above' the contact list, then resetting the dictionary wouldn't fix the problem (since the word would be coming from the contact list itself, not the dictionary).
"The dictionary" is not what I said. But, even if I had said that, then any words that aren't in the dictionary would be sourced from the contact list and other lower priority sources, wouldn't they?
What I said was the learned dictionary. The dictionary made of words learned from what the user types, when the user corrects the autocorrect. iOS does not allow you to access or modify the list of words that it has learned, only to clear them. The words that it learns seem to be prioritized above everything else, which includes special capitalizations for words (not just sequences of letters lacking capitalization). Sometimes it learns the wrong capitalization for a word, and it sticks to it. Just searching for "ios capitalizing wrong words" on google turns up a ton of results from reddit and other discussion forums that mention this exact problem with iOS learning the wrong capitalization.
FWIW, iOS 17 sounds like an all-new autocorrect system, so maybe it will fix your problem just by throwing out the data anyways, since it would make sense to me for them to start from scratch with such a radically different autocorrect algorithm.
I'm glad it works on your phone. It doesn't work on mine!
I was aware of the upcoming changes in iOS17, which is why I'm not making changes and hoping that it will be fixed. It is annoying to have to manually capitalize my own name!
I would like to see an LLM in Siri and eventually even have it interact with/control the rest of the system.
Ideally with Whisper-level speech recognition of course.
Link to the announcement timestamp: https://developer.apple.com/videos/play/wwdc2023/101/?time=1...
Whisper is phenomenal by comparison Siri and arguably even what Google and Microsoft use, and no, there is nothing that stops Whisper from being used in real time. I can run real-time Whisper on my iPhone using the Hello Transcribe app, but the last time I tried it, the app itself was too flawed to recommend besides as a demo of real-time transcription.
I am looking forward to trying out the new transcription model that Apple is bringing to iOS 17.
I've tried to talk to ChatGPT through a Siri shortcut for a day and Siri transcribed pretty much all of my requests wrong, to the point that GPT seldom understood what I want.
Even the Hey Siri … Ask ChatGPT trigger phrase fails ~50% of the time for me.
You can do this on current mobile devices with current technology (ie. 4-bit quantization).
They don't need a massive shift in the tech to make it better though. It would be cool if they improved it that way, but they have low-hanging fruit they could've addressed long time ago. Switching from one underutilized tech to another underutilized tech may not solve much without taking the whole feature seriously.
This sounds totally plausible. It will be a much smaller transformer model than ChatGPT, probably much smaller than even GPT-2.
https://www.apple.com/newsroom/2023/06/ios-17-makes-iphone-m...
Apple is also dogfooding an LLM AI tool internally also likley to gain a better understanding of how this works in practice and how people are using them. https://forums.macrumors.com/threads/apple-experimenting-wit...
In this case it's a transformer model that is not "large". So, an LM.
Unless they also made an option to run it in iCloud, but offering so many options to do a thing doesn't sound very Apple-like.
They have done it before and they will certainly do it again, especially with Apple Silicon and CoreML.
The one that needs to worry the most is O̶p̶e̶n̶AI.com as they rushed to stop the adoption of downloadable and powerful AI models for free to the regulators. That shows that O̶p̶e̶n̶AI.com does not have any sort of 'moat' at all.
They haven't mentioned LLMs at WWDC beyond keyboard autocorrect (mentioned already by your sibling comment).
I’ll be queueing at midnight for the first time ever if I’m wrong.
* User pays for the silicon and the energy used to service requests. OTOH Cook probably would love to sell you an AppleAI+ subscription…
Can someone explain to me why Meta's approach is responsible? I mean, I applaud Meta for "open sourcing" the models, but don't they contain potentially harmful data which can be accessed without some kind of filter? Let's say retrieve instructions on how to efficiently overthrow a government?
lol the examples people give for AI safety are always so ridiculous. Here I am thinking I’m bad at estimating when I misjudge a week’s worth of work for a month and this guy thinks he can overthrow a government by following a few thousand word plan written by an LLM.
Your license to use Llama can be revoked if Meta investigates and deems your action to be against the code of conduct[1]. I suppose if you continue to use the model without a license, you may be open to subsequent legal action in uncharted case-law territory.
1. https://github.com/facebookresearch/llama/blob/main/CODE_OF_...
Maybe in the future we'll get very advanced models with that number of parameters. But running current Llama2 13B on a mobile device doesn't seem too useful IMO.
A year ago, many believed that Meta was going to destroy itself as the stock went below $90 in peak fear. Now it looks like Meta is winning the race to zero in AI and all O̶p̶e̶n̶AI.com can do is just sit their and watch their cloud-based AI decline in performance and run to fix outage after outage.
No outage(s) when your LLM is on-device.
Quest runs off Qualcomm chipsets, although in terms of actual units shipped Quest is a rounding error for QC.
It doesn't matter that the Snapdragon 8 gen 2 has "AI" tensor cores or any of that. Memory bandwidth is the bottleneck for LLM. Phones have never needed HPC-like memory bandwidth and they don't have it. If Qualcomm is actually addressing this issue that'd be amazing. But I highly doubt it. Memory bandwidth costs $$$, massive power use, and volume/space not available in the form factor.
Do you know of a smartphone that has more than 1GB/s of memory bandwidth? If so I will be surprised. Otherwise I think it is you who will be surprised how specialized their compute is and how slow they are in many general purpose computing tasks (like transferring data from RAM).
For LLMs the biggest challenge is addressing such a large model - or finding a balance between the model size and its capability on a mobile device.
People are unreasonably attracted to things that are "minimal", at least 3 different local LLM codebase communities will tell you _they_ are the minimal solution.[1]
It's genuinely helpful to have a static target for technical understanding. Other projects end up with a lot of rushed Python defining the borders in a primordial ecosystem with too many people too early.
[1] Lifecycle: A lone hacker wants to gain understanding of the complicated word of LLMs. They implement some suboptimal, but code golfed, C code over the weekend. They attract a small working group and public interest.
Once the working group is outputting tokens, it sees an optimization.
This is landed.
It is applauded.
People discuss how this shows the open source community is where innovation happens. Isn't it unbelievable the closed source people didn't see this?[2]
Repeat N times.
Y steps into this loop, a new base model is released.
The project adds support for it.
However, it reeks of the "old" ways. There's even CLI arguments for the old thing from 3 weeks ago.
A small working group, frustrated, starts building a new, more minimal solution...
[2] The closed source people did. You have their model, not their inference code.
Optimization for this workload has arguably been in-progress for decades. Modern AVX instructions can be found in laptops that are a decade old now, and most big inferencing projects are built around SIMD or GPU shaders. Unless your computer ships with onboard Nvidia hardware, there's usually not much difference in inferencing performance.
Qualcomm has it's own SDK for that (used to be called SNPE), which uses GPU, DSP (hexagon),.. CPU is really only a fallback.