Transformer architecture optimized for Apple Silicon
github.com
github.com
openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that openai is a sufficient motivator.
Different directions - this is the fundamental challenge. It's hard to know what these directions are when we're flying blind at the cutting edge of technology.
We're all human. I suspect out of the 100, you'll have 95 of them going into the same as before areas. Maybe 5 truly understand the above point and explicitly strike down ideas that have been done before.
GPT-3 has been live for a while now. The industry has hundreds of story generator apps specialized for various kinds of content generation. Very few are really thinking about AGI or chain-of-thought reasoning etc as an example.
So I don't think launching 100 of them into different directions would work. Maybe 5 of them with both the expertise and direction to pursue seemingly impossible goals.
By all accounts GPT-3 is wildly inefficient in resource use; OpenAI runs like a company that’s concerned with the functionality it can achieve by calendar date, and has an almost infinite bankroll to do it. But, other actors in the field have different priorities, and the various open source or, at least, available-to-use models seem to be far more efficient than the OpenAI models of similar function (though they are behind the newest OpenAI models in function.)
We must be talking orders of magnitude differences in operational cost, not to mention completely unique features like privacy. The very definition of disruption, waiting in the wings.
Not only that, but this model of 'wait in the wings and pounce when the tech is really there and absolutely nail it' is Apple's forté, aligned completely with the cultural and strategic DNA of the company.
So much for an “unfathomable moat”.
There are definitely two tiers of Siri queries. There are queries like "set the brightness to 10%" or "set a timer for 5 minutes" which absolutely and consistently work without internet, and have for several years, and if you're legitimately having a different experience then its possible a cosmic ray hit your iPhone (or, realistically, you're running into a strange and rare bug which is not indicative of the general experience and will be fixed in three to five years or maybe never). There are also queries like "create a new note" (from the Notes app) which should be able to work offline, but don't. And, naturally, there are queries like "when did Resident Evil 4 come out" which wouldn't reasonably work without internet (but, if you're curious, she does get it right).
In other words, Siri clearly does some local guessing as to whether she can answer a query without the internet, and some queries appear to be miscategorized into the second bucket. My leading theory on why this happens, which may be incorrect, but it seems like: if an app has any Siri functionality which requires the internet to answer, all of the queries which are responded to by that app have to require the internet. It doesn't matter where the processing ends up actually happening to respond to the query, it just shuts the query down. Its weird, but its consistent with the behavior I've seen.
The more important point: Siri's real and weird limitations don't seem to have much to do with limitations in local processing. They've said that the speech interpretation all happens on-device. They encrypt practically all of your data that does get shipped to their servers, so a query like "Open the note titled 'Hello World'" is probably also being processed on-device. But: Siri still requires the internet for that query. That doesn't seem like a significant limitation with their ML algorithms or silicon or anything meaningful; it seems like just a case of dumb coding, which is certainly something Apple is no stranger to.
interesting - latest iOS in airplane mode - "hey siri what time is it?" - "you need to turn off airplane mode to do that"
but - "hey siri set a time for 5 minutes" - works just fine
Focusing on privacy and on-device learning is great, but when the strength of these models is in consuming all the data they can hoover up your motive is at odds with your philosophy.
And seriously, did you really try to assert the Google whatever it is has outperformed Siri? I’d love to meet someone that uses it so I can verify this statement. But with the advent of Chatgpt Alexa, Siri, and Google whatever all look like freshman projects. The next generation of assistants have just begun, and I assure you, the state of the past has nothing to do with what comes next. OpenAI hit a giant reset button across an awful lot of stuff.
Anything outperforms Siri. Google Assistant is head and shoulders above Alexa and Siri.
We only see the tip of the iceberg so who knows what else it is capable of.
“Sorry, I’m having trouble connecting to the network”
“Now playing Eminem Love the way you lie featuring Rihanna”
Which TV? Bedroom or Living Room or Everywhere?
(Only 1 of the 2 TVs is ever on)
I won’t spam this thread anymore, but I would be pleasantly surprised if it improved.
- Hey Siri, <do something that involves AppleTV in any capacity>.
- I'm sorry, one of your devices is off
Siri: I'm sorry, that contact is not in your list
Me: Siri, what is the number for <my city> Toyota.
Siri: <my city> Toyota's phone number is 123-456-7890 [said too fast to remember or write down in one go]
Me: Siri, call <my city> Toyota
Siri: I'm sorry, you do not have that contact number.
Me: &$@@&/&&/&!!!
I find the worst errors to be when I ask it for information, or to send a text, and it instead places a phone call to someone. It will even call people that I have never called on my phone (a fact it should know), without asking first.
Starting to think all these Siri complaints are either made up or really outdated.
I should be able to say hey siri, start playing <name of media> on <name of Apple TV>, and it should be able to start the TV and start playing it.
This isn't in response to a prompt either. He'll unpause the connected TV and it'll just start blasting U2.
In case someone doesn’t know they gave that album to all iTunes users as part of a promotional thing when it came out.
On my iPhone if I connect the Bluetooth headset and accidentally push the call/cancel button out comes U2…
For me, Siri is like an 80s text adventure game, except I was better at those.
Edit: just now I was able to call by interacting a second time with Siri, and using the physical button. Using the button never would have occurred to me while driving.
done
Perhaps the verb "text" is unclear to Siri?
"Hey Siri, what is today's date?"
"That's sweet, but I think of you as a friend"
"What? The date, Siri, what is the date?"
"You're so sweet"
"Sigh"
Just swipe down on any iPhone.
Apple has access to SUBSTANTIALLY more fine-grained personalized data than Google does. Full stop. That's a weird thing to assert in the tech crowd, but think about it deeply and you'll realize its true.
Google has islands of services that they've done an extremely good job of building bridges between. They have some islands that Apple has nothing like (YouTube is the biggest one) (Search is a huge island, but only indefensible if the interfaces customers use to access it don't leak the queries and results; no one opens a browser and types "www.google.com" on mobile and even if they did, Apple controls the browser renderer, they could lift everything if they wanted to, they won't, I'm simply illustrating how much Apple should fuckin scare Google). But the waters between those islands are patrolled by Samsung, Xiaomi, Oppo, and Vivo, and there's sharks in there as well (uh, the metaphor is falling apart, the sharks are "Google's tumultuous history with privacy and user blowback if they overstep").
Apple is a continent. Many of their cities are ports that connect to some of Google's islands, but Google's product still has to be checked by customs.
You're right that even given that power, hoovering it up is at-odds with Apple's philosophy. Or, is it? I think "hoovering it unencrypted to the cloud" definitely is; but that's the point OP was making: if we're extremely close, as a species, to solving "make AI work", one of the challenges for the next five years is inevitably going to be "make it more personal". Its awesome that I can ask ChatGPT to write a date comparison function. It'll also be awesome if I could ask Siri "when did sarah and I talk about getting a cat" or something.
That requires personalized data to set the context. If Apple can swing at the fences and say "everything is on-device, nothing leaves, we can't see it, and now you can ask Siri that, and by the way there's new functionality built-in to iOS that apps can leverage to integrate with Siri's new LLM capabilities just like ChatGPT plugins"; that's an extremely compelling product. Extremely. And I know they would ingest that data, that they would do that, because they already do! Go ask Siri to call Sarah, or if you have any meetings tomorrow (assuming you're using Apple Calendar), and it will respond. I don't know where you're getting this take that they don't "hoover data"; they ingest everything from all their first party apps into Siri's `DATA MATRIX`. Its just, you know, 2010s era querying and data crunching.
Don't get me wrong, Google's gonna do a lot of this too. But Google is playing from the position of "we have a model, we have the data, we just need to pay hundreds of millions of dollars a year maintaining all these servers". Apple is playing from the position "our customers are paying us for the silicon to run this, we have the data, we just have to figure out the model". If its not obvious at this point: the models/algorithms/etc are not a moat. They're going to be commoditized with time. Publicly accessible data isn't a moat. Personal data is a moat, because people care about privacy. And cost effective training and inference silicon is a moat, because its Physical, and like literally One Company on the planet makes it, and they're in a big time situationship with Tim Cook.
You dont spawn a network request on every local search on an apple device.
They may not be storing these queries or using them for further analytics, training, etc. But they could.
I got Nvidia on my paper. Did I do the math wrong?
Apple has purchased the entirety of TSMC's 3nm production for the next 1 to 2 years. They can do that, and can continue to do that, because they have a ton of money. They have a ton of money for buying chips because there isn't some crazy business middle-logic justifying the cost of these chips' performance; they buy chips, they sell chips to consumers. In comparison, literally zero other customers of TSMC derive the majority of their hardware revenue from selling to consumers. Companies like Nvidia make some money this way, but most of their money goes to data center sales, which has their own business justification for buying them which fluctuates (ChatGPT subscriptions? training models? is the past model good enough? etc).
Do they? I can completely fathom, given my own anecdotal experiences, how garbage Siri quality is today. I haven't built anything against Siri APIs, but I've used Siri and various integrations and every single time I give her a shot, she disappoints me. It can do cookie-cutter, super well-traveled code paths that were engineered together for demos, basically just one off tricks without cohesion, but the whole infrastructure around Siri is not open and has not been improved in any way in many OS versions in any way that I've discerned... But I am not an expert with these systems or APIs.
The Neural Engine architecture especially considering M1/M2 hardware leaps (which are absolutely mind-bogglingly impressive) are both really superb technologies from my layman view as a mere software engineer, with improved battery life as well as performance, a real improvement over Intel x86 architectures - I just don't see Apple as a serious AI player right now, despite head start in AI w/ Siri acq, despite these hardware leaps that may make it easier to do cool stuff in the future like this post. The OCR stuff in Photos is cool and useful and I use it every day, usually to look up my Known Traveler Number.
Ramblings over, just wanted to ask you to opine as to Apple LLM related information if you have any additional context!
I think this can be a scenario of converging incentives: on one side large models will incentivized hardware manufacturers to increase the memory available on the devices, while on the other sides model developers will be incentivized to trim the fat on the models and devise compression mechanisms that don't compromise quality too much.
It's not unthinkable to imagine a hand held device able to run full inference locally a few device generations in the future.
Note I'm not suggesting you can pack the full knowledgebase of humanity into those 2GB of RAM, but the key feature of an edge AI is simply to understand instructions, something Siri and Ok Google struggle with at best..
a few notes on this
- siri is incredibly underinvested. They have another team that's building some sort of search and natural language processing engine, that has slowly sapped away some key headcount from the siri team.
- apple doesn't get the full advantage of tons of user data from the wild. this is both a bug and a feature
- the siri api model is clearly generations old, and i suspect that apple has been marshaling its resources into a big leap forward - alongside their hand tracking, ar, and hardware
- apple has shipped everything required for you to point at a light /in your house/ to turn it on or off (and optionally flick). this includes software - the individual components are built and ready - the only thing missing is gluing it together
There's something happening, and I think that the rumored glasses are the hardware totem.
You might be right but there’s no real evidence for it yet.
For Siri itself, I think it runs locally now.
that's a mischaracterization of what i said. what i said was that they have been clearly hiring key positions and cannibalizing the siri team for something new. that something new will likely be a major release. i also believe that the ar headset is the unifying product under which they're rallying.
The first proper integration of a Whisper & a GPT-class LM will be a big step change, a Siri-like AI that you can actually talk to somewhat sensibly and expect it to "understand" more than pre-set phrases... Google could be in a position to release something for Android, and their are Android SoCs with neural accelerators as well (sure nothing as impressive as Apple's chips but Samsung have demonstrated a cut-down SD model on their SoC).
Could you expand on this bit? I’m pretty deep in the Apple ecosystem, but I’m not sure what you’re referencing here
Alternatively whole-home mapping via AR also solves this. No triangulation needed.
Everything required for apple to know not only where you are, but which way your hand is pointing, as well as where "smart devices" are in your house is already being sold and rolled out en masse to the majority of people in the ecosystem.
To turn a single lamp on/off by pointing at it, something that nobody ever wants to do (alright, once for the cool factor). If someone wanted to overpay for useless features, they can already go for a Philips Hue.
I can turn off my lamps from anywhere in the world using a $10 Tuta ZigBee bridge and a $8 LIDL light.
I've put zigbee switches into every wall switch in my house and now when I stay in hotels, I forget to turn off the lights before getting into bed.
>something that nobody ever wants to do
I think that this gesture, if it works well, is something that will be a killer app for smarthomes. The other point of friction, which is interop, has been more or less solved by thread/matter. Homepods are $100, btw. Not to mention what happens if apple integrates u1 into a smart switch/bulb, which is something that the thread protocol allows for.
So my prediction would be that apple releases a ZigBee bridge and light for $50 in about 3 years and calls it a fancy name. People will tell you how absolutely essential and innovative it is and how you could never get that without an iDevice.
I remember seeing AirPods in somebody's ears for the first time. "What a dork" I thought back then. Now they're cool and everywhere.
One of these things is true.
AirPods are cool by that measure as well.
I won't buy Meta products but I keep waiting for someone to do VR workspace right - where I can truly work and travel from anywhere with a decent chair and a desk for keyboard (no need to lug around monitors, no need for huge table).
Having said that, for me to spend money on it it would (a) also be able to realistically replace my set up at home, which currently consists of 2 HiDPI screens and (b) not be that much more expensive than what that cost me.
Do you think OpenAI did? Perhaps Apple doesn't have the vast amount of data that Google has, but if OpenAI managed to have like 200 wikipedia-sized corpuses of different textual data for their English GPT models, that's certainly not out of reach of Apple.
Point at it? You mean with a finger? What do you mean they've shipped everything required?
Alexa is DIMENSIONS better than Siri. Siri can create a timer ... okay even an alarm. That’s it. It is comically bad. Their text to speech is excellent, but the rest is unbelievable bad.
If Apple has some kind of silver bullet, it’s time to put it out or be left behind.
We’re circling back around to local compute being the default as hardware performance of next gen phones and tablets reaches a “good enough” point for most users.
There will be scientific problems that will require modern server clusters but most consumer facing AI needs will be done on hardware within a decade.
I’m not saying AGI in a decade, I’m saying cutting edge logic embedded in software now will be the basis for logic in chips in the years to come.
I think it's quite likely that Apple realized the same quite a bit earlier and eased up on Siri versus focusing on other things.
I think that the only thing that changed is adoption of fingerprint reader. Both in iPhone 8 and in mac.
AI might be that thing that could tremendously change my usage pattern. But it should be as smart as human assistant. GPT4-level intelligence might have the necessary power for that.
> Do they? I can completely fathom, given my own anecdotal experiences, how garbage Siri quality is today.
Siri is a dead end. When Jobs bought Siri (what, 10 years ago?) he explicitly junked almost all the AI back end, mainly buying the speech recognition engine. I didn't understand why and still don't (but strangely he didn't ask me :-).
John Giannandrea has run Apple's AI effort for the past five or six years. He is the reason Google has a big AI effort (he consolidated a bunch of AI projects and bought Deep Mind, etc) before he decamped for Apple. For all I know the Siri team isn't even part of his remit.
You can never look into Apple (even if you work there) so one can only speculate based on what visible signs appear. But Siri isn't one of them.
Blue Origin still has not reached orbit.
I'm sure he's great but this made me laugh
(And yes he’s a great guy)
If you do that, you can do inference on mobile devices, which is a huge privacy win; it plus into their general privacy positioning in a big way. If you open up that SDK to developers it would be the iOS app gold rush all over again.
Note, Google is also moving in the direction of on-device inference with Coral, but they obviously will want to transmit the personalized model weights back to the mothership.
Ultimately, I'm very reassured that it's Apple and Microsoft leading the way for personal and business AI assistants respectively. As Stratechery has emphasised, the latter not only already has all the data from most firms outside Silicon Valley, but more importantly has always adhered to and designed their software according to a computer-as-a-tool philosophy versus Google's "ML/AI will do everything for you behind the scenes"; the former has not only made a show of data security but invested massively in the hardware architectures necessary to making on-device AI a possibility. Without at all being a "fanboy", Apple is quite literally the only company I would trust my personal data to for use in a GPT-based assistant.
What they have been invested heavily in is the Apple Neural Engine ("ANE"), special silicon right on the SoC to handle ML / AI code. Optimize on a server, then run the model on your iPhone or probably soon, your Apple Watch.
WWDC this year is going to be very, very important.
[1] https://www.cnbc.com/2023/01/06/amazon-fully-committed-to-al...
I dominated a 13" mbp m1 for like two solid years on hardcore startup software engineering. The base model. It was almost entirely perfect until the last few months when I upgraded to the new 16".
These are alien devices. They're just not possible yet here I am using one.
Apple has decent hw but no LLM software. Unfortunately, in LLM space software changes are the ones driving performance for now, since the space hasn't stabilized yet. Since their competitors control the software, they get to adapt it to their hardware. That is, Google and Microsoft are going to adapt GPT/Bard to Qualcomm ARM etc. while Apple is being ignored. Unless Apple gets in with their own LLM (quite possible), their hw advantage will end up not mattering one bit.
Examples would be the fact that Samsung almost always beats them to the punch on camera technology and other whiz-bang features, but Apple eventually adopts to much mass market consumer acclaim and groans from Android techies. Another example is they’re just now considering touch screen on laptops.
A counter example is the TouchBar - which was “innovative” but many didn’t like.
The bull mindset is fun to watch unfold (especially here on HN) but I think people should temper their expectations.
It’s only the being but the general process to build this stuff now exists and the value proposition is crystal clear.
Almost 5 years ago TalkToTransformer did 80% of ChatGPT's job with 0% of the hype. Once people realize just how glacial all this stuff moves, I think the honeymoon phase will be over.
But what incentive does Big Tech have? Especially considering that presumably they could monetize the cloud services easier. It's not like most consumers care either.
I expect smaller companies like Stable Diffusion and grant/govt-funded research to bring cheap, local inference. There's definitely a lot of demand. But is there more demand than cloud-based services, so that it's economically viable for the biggest companies (Apple)?
one hypothetical is that you need an icloud subscription, a token is retrieved from apple, the token "unlocks" your AI module, on the phone and allows you to do the inference.
in this way apple could charge monthly for this and claim that the inference happens locally. sadly in this was it was similar to the whole csam debacle
apple is also one of the few companies that could realistically get manufacturers to massively produce a hypothetical "AI model chip" that has the model on device, at a quantity that would make it realistic to pay for in a hypothetical iPhone 19 Pro model.
I can see it - "Siri+", pay $5 a month for fine tuned model upgrades straight to your device and remote fallback. local inference available for Pro devices only.
In addition to the trainers, you have producers, graphic designers, artist spotlights, celebrity interviews, etc.
Yes, some costs are fixed so with more subs you get higher profitability but that’s really different than charging for something that doesn’t have a variable cost.
I can imagine similar for iPhone on the edge, when someone manages to decrypt the model it will be free to grab for anyone, unless there is going to be some proprietary thing going on only available on Apple Silicon that is undocumented.
Apple is really getting a niche here for machines to run models locally. That’s pretty powerful.
I... don't agree. They have an acceleration API and a large install-base, but all of these models have run just fine on traditional hardware. GPT-Neo, Stable Diffusion and LLaMa can all be used and accelerated without Apple Silicon.
Powerful, maybe. But not really unique, just putting up table stakes.
I think the usual current model would be the 24GB Nvidia RTX 4090.
at the very least, even if that's not the case, inference will be drastically less gpu heavy by then I suspect.
"We find that current large language models are significantly under-trained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant."
It's difficult to evaluate a LLM's performance as it's all qualitative, but Meta's LLaMA has been doing quite well, at even 13B parameters.
I think what we have access to is a fair bit slower.
If you focus on english only, this can easily reduce the paramters 5fold
LLMs seem to be comfortable with hundreds of programming languages, DSLs and application specific syntaxes so how does supporting a couple more natural languages become so expensive?
I see how more training data would be needed, but I don't understand how that maps to a greater parameter count.
The devs working on llama.cpp have been discussing ways to further reduce the memory requirements by mmapping the large weights files (I thought LLMs mutated the weights as they run inference, but they clearly know more than me about the internals), bringing it within reach of phone memory.
So, iPhones are not as far off the computational capacity to run these models as you'd think. Memory (and to a greater extent, battery and cooling) are the limiting factors. iPads even less so, given they run M1 chips and have much larger batteries & much more RAM
My other motivation is making sure I understand what offline LLMs can do... while I use GPT-3 and 4 extensively, I don't want to send something over the wire if I don't have to (e.g. if I can summarise e-mails locally, I'd rather do that than send them to OpenAI).
It's also surprisingly good at defining things if I'm somewhere with no internet connectivity and want to look something up (although obviously that's not really what it's good at & hallucination risks abound)
It's certainly nowhere close to the quality of OpenAI summarisation, just better than what I previously had locally (e.g. in summarising a family history project with transcripts of old letters, gpt-3.5-turbo was able to accurately read between the lines summarising an original poem which I found amazing).
I half wonder if the change in spelling from US -> UK makes a difference...
I'd run a test on that but I've just broken my alpaca setup for longer prompts (switched to use mainline llama.cpp, which required a model conversion & some code changes, and it's no longer allocating enough memory)
Alpaca seems like it could be significantly improved with better training (some of the old training data was truncated), so I think there's a decent amount of improvement to be had at the current model size.
In the future though... what would really be a meaningful change would be a larger context size - the 8k tokens of GPT-4 was a big improvement for my uses... I would guess a future local llm with larger context would exceed 32GB, but that's speculation beyond my expertise, I don't know how context size and network size scale.
If it was a PC I'd say go for 64GB, but hard to recommend that given how much Apple charge for RAM upgrades. On my next upgrade (2+ years time, hopefully) I'll likely opt for 64GB+ though
I'm curious, which configuration of the M1 MBP do you have?
The only drawback I've found with the M1 Max model is the added weight from the bigger heatsink just makes it a hair heavier than I'd like when picking it up at the front with one hand when open... and that in the winter time the case is cold no matter what you're running, I used to love that my Intel MBP acted as a mini leg warmer :-)
RAM hasn't been increasing as steeply as it could, but if there's a strong use-case for it, it may happen. Also consider that Apple is in control of the whole chipset and software, so they could implement things like turning the extra RAM on only during ML computation.
It's not expected to. The consensus seems to be ~2025 https://arxiv.org/abs/1511.05956
I think people are very prematurely counting AMD and even Intel out of this race considering the pace of change.
Look at AMD vs NVIDIA RTX implementations. Nah, NVDA is a solid buy for long term. They also almost acquired ARM (Which shows where their heads are at at the very least.)
Owning ARM, the thing that is inside ya know, everyone's phones, routers, etc
Chips intended for launch in 5-6 years are in planning stages right now. Apple, nVidia, and Intel could bring serious hurt to software companies in the next 5-10.
Open source purists will weep but really most people do not care, and tech should not merely serve the dedicated.
What does that mean? What software companies?
Btw transformers are really simple and the optimisations making them run fast on CPUs have come a long way. I don't know the M1/M2 benchmarks for this, but many CPUs for edge devices have NN accelerators on the silicon that can run this, it's the pairing of the accelerator and gigs of RAM that is the key.
OpenAI is 49% owned by Microsoft. OpenAI is big tech.
A 256GB Mac Pro can dedicate almost all of that to AI.
Their GPUs are the weak spot. The Nvidia 4080 is 2.5x faster than the M2, and the A100 is 15 times faster.
Not a single benchmark in the world has supported Apple's claim that the GPU in the M2 is that powerful. It's just yet another cute embedded GPU that does the job, but nothing more. It's made to push out 8K frames really fast, which it does because of UMA, but want demanding task will have it be eaten alive by any real GPU.
But in any event, yes, that was my point. UMA is a huge advantage, the GPU itself is too weak to be serious.
But it’s a lot easier to drop a dramatically beefier GPU into a new design than it is to update the entire platform for UMA. Apple has a huge opportunity here… whether tbey pursue it or not remains to be seen.
Pursue what though?
UMA is cool, but kinda meaningless if the majority of Macbooks are min-spec. That leaves you with 4-5gb of VRAM, assuming you've left nothing open. What is Apple going to do with that UMA that other manufacturers cannot?
It's certainly nice that 128gb Macs exist for models that might be too big to otherwise load into memory. It's useless for production inferencing though, and I struggle to imagine the "opportunities" they're missing out on here.
A Mac variant that trades CPU cores for GPU/ML cores while having 192GB+ of UMA memory.
> I struggle to imagine the “opportunities”
Two of them: 1) academic / R&D compute, where people could have at least A6000 class GPU on the desktop, and 2) cloud inference servers, probably for Apple’s own services.
I’m not saying they will or should do those things, just that the apple silicon arch is well positioned if they choose to. Bolting on exponentially better GPU is not especially difficult, and they’ve got an OS that would bring existing apps and libraries right over.
Look at it this way: is there a path to UMA on Windows / Linux? If not, those systems will always duplicate RAM and require users to decide in advance whether to allocate RAM budget to OS or ML.
Whichever way you look at it though, neither of those are really opportunities. Apple boxed themselves into the consumer market, and now has to compete with professionally-priced products.
*I used to be one in a Sales and Trading role and fully believe a lot of rumors are to keep the finance analysts happy.
Lets hope Apple doesn’t screw this up.
While also overestimating how many parameters anybody really needs after fine tuning.
The market will find the sweet spot. Right now everyone’s tinkering with the 7B parameter LLM and then going to move up to the 65B one once they've refined the process. I think its fiction that anybody really needs a 10 trillion parameter LLM at all. It will be completely niche.
Basically the curtain has been drawn and shows that current LLM’s are just very inefficient and will be optimized in weeks. Whatever improvements Apple’s Neural chip offers just needs more RAM closer to it, which will likely come at the next hardware fresh where this year its probably too late, while whatever is released at the end of 2024 will be good enough.
Imagine Google Sonar on Pixel phone with Pixel ear buds and more sensors and ChatGPT onboard.
OpenAI promotes thin clients - run everything in the data center.
Apple has been a proponent of thick clients - run everything as possible locally.
So, it make sense that Apple will start to promote local LLM support. That's why they pushed Stable Diffusion support so quickly:
https://machinelearning.apple.com/research/stable-diffusion-...
It may be so that scaling down a language model like GPT 4 is not possible on hardware systems orders of magnitude smaller than the one used by OpenAI.
I'm not saying it's impossible but it's fallacious to just assume outright that it's inevitable because market forces.
For all we know it could turn out that the only way to get a gpt in a pocket format is through some form of analog chips that must be individually trained that Hinton aptly described as an era of mortal computing.
Same goes for cancer cures. I'm not sure if there are physical limitations preventing a room temp superconductor.
Not really. We just know that it’s very uncommon at best. The limit is material dependent and we already have superconductors with more than one order of magnitude difference in their critical temperature (e.g. ~4 K vs ~40 K; YBCO, which is widely studied, is at 90 K). It is not inconceivable that we could come up with some fancy material with a critical temperature three times as high again. There are several laboratories with good money working on it.
OpenAI won’t want to; open source competitors already are, and they will keep getting better. The more of a lead OpenAI has over commercial competitors, the more incentive those commercial competitors will have to back open source options.
Siri is still extremely basic and barely usable for anything besides starting a kitchen timer.
M1 chips?
Leaps in performance across all metrics in an existing thing is groundbreaking, especially when it’s just the beginning.
On a practical level, it feels pretty groundbreaking to me when I go back to use my previously top of the line 16” MBP from a year before. I suspect you haven’t had the pleasure of using an M-series computer.
Local has a lot of advantages as well (latency, privacy, etc).
Maybe Apple has the ultimate
Crazy bet, what makes you think that? You cannot optimize infinitely. Raytracing probably had decades of ppl trying to make it run fast and yet even today you need strong hardware
The reason openai have had to rush out plugins is due to software like langchain coming in at meteoric speed.
Things are moving on a day to day basis in the ai sphere at the moment.
Their technology (LLMs), or the secret sauce, can easily be stolen just by the process of putting that tech out there.
Have a look at Alpaca, FB made it, someone leaked the weights and now there's a dataset of training it for only a few hundred dollars that can beat openai at its best.
Not everyone needs to employ a PhD for doing customer service, in the same way not everyone needs GPT5 for answering support queries.
Their business model is leaking away from them.
I'm running a totally usable 13b parameters llama model in my macbook air, which seems to give outputs equivalent to what I was getting from GPT3 in June 2022.
How much more hardware would it really be needed for GPT-4 level outputs natively? Perhaps software optimizations alone could do most of the trick.
if you work at OpenAI and you’re optimizing anything you’re not doing your job right.
LLMs will absolutely be able to run locally, but whether Apple will be able to stop worrying and love the model remains to be seen.
They’ll have to do something about Siri soon though. Even my 5 year old daughter told me Siri is ‘a bit thick’. And that’s just compared to Alexa never I’d ChatGPT level
...it's shifting...
It's fascinating to me that at least one of two things is true: either (a) Apple has lost its ability to coordinate "hype" between its teams, or (b) the difference between comparable levels e.g. the M1 Max vs. the M2 Max are so negligible that they don't look good in an announcement like this.
Has anyone run inference for LLMs or other transformer models on comparable M1 and M2 Macs? Are there good benchmarks for this specific workload?
And this was put out in Aug 2022. It's very propable that the team worked on it and tested it just on M1, and M2 was kept under wraps in different teams working on it until the announcement. So they just wrote the annoucement on the CPUs they worked on - and since it's not for a commercial product, Apple didn't care to optimize marketing anyway.
The M2 added more GPU cores and an advanced media decoder.
M2 NPU is supposed to be 44% better than M1
https://www.cpu-monkey.com/en/article/apple_m2_vs_apple_m1__...
1 - https://findthatmeme.com/blog/2023/01/08/image-stacks-and-ip...
Stable Diffusion sits at about 2GB for fp16, Whisper Medium at 1.53GB, LLAMA is 120GB.
Sure, Apple can ship an optimized model (<2-4GB) as part of the OS, but what if a capable app maker wants to ship one? Users will not be happy with an app sized at >1GB.
I don't think you can pirate e.g. the Facebook app binary and make it your own.
The problem with OpenAI's business model is that it's actually quite expensive for them to maintain centralised processing. With Apple, there are billions of very powerful computers deployed to users and these computers mostly stay idle apart from occasionally running some bloated JS to show a button ar something. If Apple manages to run a good enough model on device with acceptable performance and energy impact, then suddenly OpenAI and Microsoft will be just burning away money with no expectation of recouping if they provide the service for free, if they make it paid they will be making money in a niche.
There really isn’t that much difference between iPhone 12 and 14
If a nee one comes out with LLM Siri + hardware that makes it possible that would be a massive upgrade cycle.
Yeah, but if every individual is running decent chat-capable LLM, and businesses are running their own on their own devices, and those can communicated with each other, who needs to rely on Skynet?
They can announce an iPhone Pro Ultra model that comes with higher RAM and storage capacity along with a souped up Neural Engine, similar to how they do today with how there are differences between the iPhone and the iPhone Pro screen and camera.
Even better, they could bundle the base model for LLAMA with iOS and ship incremental model updates to those iPhone Ultra users (possibly on a monthly subscription).
Most interactions probably don’t require full power
https://twitter.com/LinusEkenstam/status/1638999208911949845...
1. Is this a new LLM from Apple?
2. Is this a way to optimize running LLMs like Llama locally on M1 macs?
3. Something else altogether?
> [T]he device spec for this reference implementation is M1 or newer chips for the Mac and A14 and newer chips for the iPhone and iPad
I am just a little better informed. As I understand it, their code improves model performance and memory consumption using PyTorch and Huggingface libraries.
Their examples compared the A* CPU performance and the repo includes Swift only code samples. But they’ve also made it possible to use them with traditional tooling (torch, huggingface).
Hope that helps explain it.
PyTorch is supported for example, it’s a machine learning library with GPU acceleration that’s been around for 6 years now. It’s used in a few commercial projects, including Tesla Autopilot. It can be used for natural language processing, image manipulation, and possibly to build an LLM I suppose, but as a low level library it just gives the base tech to build such systems from.
Local inference is huge for anything that requires even a little bit of privacy.
This. Nobody really cares about local processing for privacy.
Even those that claim to often don't mean it. Remember the total freakout over Apple's proposed local, privacy-preserving processing to detect CSAM before uploading it to iCloud? The consensus seemed to be that secret, opaque, and un-auditable cloud-based scanning was much preferable.
After all, it's hard to care when you don't have power as an end user too act on those feelings.
"It was pretty open"
Not really. The scanning code wasn't open source. Apple didn't explain how their proprietary hashing algorithm worked, let alone providing any source code. It was just the promise of Apple about how the whole thing would work. If that's sufficient, then iCloud scanning is as open as local scanning.
Auditability of DB updates doesn't mean much either. Even if a third party organization detected "new additions" to the CSAM DB, what would be the next step? There would be no way to verify if those hashes actually correspond to CSAM. A dictatorship, for example, would just say "yes it's CSAM, trust us", and you'd have no option but to trust their word. Even in the US, there is no way to verify if CSAM DB fully corresponds to actual CSAM. It's just NCMEC's word. We simply don't know.
I would love to see some links to any where it has been used?
So far the most promising codebases for running LLMs on Mac have been the 'cpp' reimplementations, which ditch Pytorch and run on the CPU, using other tricks to fit the model into available RAM.
Though this is painted by my personal beliefs, in order to maximise innovation, I believe it's the role of government to implement regulations that support a competitive commercial environment. Without this, monopolies will form around hard to obtain resources, innovation will stagnate and consumers will be subject to exploitation.
Currently, data is mostly acquired without user consent and is accessible retroactively. Companies own that data, they can trade it and they can use it however they want to. You as a consumer have no say in this and it's virtually impossible to live a normal life without being the subject of analysis.
While it's incredible that companies can produce undeniably valuable products like Copilot - ultimately - they will profit from these products. The irony is they built them from data sourced from you, likely from something you paid for (MS Word, etc).
The key ingredient in these products is training data. If you wanted to compete with them, no matter how capable you are as an engineer, you could never make Copilot without the same scale of data Microsoft has gathered.
I don't know what kind of regulation would even out the playing field, but I wouldn't mind being compensated for my role in creating these highly profitable products.
TL;DR: execution of pytorch models on apple's neural engine and standard data-oriented optimisations (changing matrix layout, chunking to optimise temporal cache locality, and minimising redundant memory copies)
(This here only applies to inference, right?)
Latest commit seems to be "on Aug 9, 2022"