I'm curious: is anyone seriously using apple hardware to train Ai models at the moment? Obviously not the big players, but I imagine it might be a viable option for Ai engineers in smaller, less ambitious companies.
I'm curious: is anyone seriously using apple hardware to train Ai models at the moment? Obviously not the big players, but I imagine it might be a viable option for Ai engineers in smaller, less ambitious companies.
"No, we should be use cheap commodity abundantly available cpus and orchestrate then behind cloud magic to write our nl translation apps"
or maybe "no we should build purpose built high performance computing hardware to write our nl translation apps"
Or perhaps in the early 70s "is anyone seriously considering personal computer hardware to ...". "no, we should just buy IBM mainframes ..."
I don't know. Im probably super biased. I like the idea of all this training work breaking the shackles of cloud/mainframe/servers/off-end-user-device and migrating to run on peoples devices. It feels "democratic".
For example, running a local model and access to the features of a larger more capable/cloud model are two completely different features therefore there is no "no we should do x instead".
I'd imagine that a dumber local model runs and defers to cloud model when it needs to/if user has allowed it to go to cloud. Apple could not compete on "our models run locally privacy is a bankable feature" alone imo, TikTok install base has shown us enough that users prefer content/features over privacy, they'll definitely still need SoA cloud based models to compete.
An older example is the “Hey Siri” model, which is fine tuned to your specific voice.
But with regards to on device training, I don’t think anyone is seriously looking at training a model from scratch on device, that doesn’t make much sense. But taking models and fine tuning them to specific users makes a whole ton of sense, and an obvious approach to producing “personal” AI assistants.
A few years ago it wasn’t shared between devices so each device had to do it themselves. I don’t know if it’s shared at this point.
I agree you’re not going to be training an LLM or anything. But smaller tasks limited and scope may prove a good fit.
That said, inference on apple products is a different story. There's definitely interest in inference on the edge. So far though, nearly everyone is still opting for inference in the cloud for two reasons:
1. There's a lot of extra work involved in getting ML/AI models ready for mobile inference. And this work is different for iOS vs. Android 2. You're limited on which exact device models will run the thing optimally. Most of your customers won't necessarily have that. So you need some kind of fallback. 3. You're limited on what kind of models you can actually run. You have way more flexibility running inference in the cloud.
Honestly, you can train basic models just fine on M-Series Max MacBook Pros.
I simply want to point out that these folks don't really care about that. They want a Mac for more reasons than "performance per watt/dollar" and if it's "good enough", they'll pay that Apple tax.
Yes, yes, I know, it's frustrating and they could get better Linux + GPU goodness with an nVidia PC running Ubuntu/Arch/Debian, but macOS is painless for the average science AI/ML training person to set up and work with. There are also known enterprise OS management solutions that business folks will happily sign off on.
Also, $7000 is chump change in the land of "can I get this AI/ML dev to just get to work on my GPT model I'm using to convince some VC's to give me $25-500 million?"
tldr; they're gonna buy a Mac cause it's a Mac and they want a Mac and their business uses Mac's. No amount of "but my nVidia GPU = better" is ever going to convince them otherwise as long as there is a "sort of" reasonable price point inside Apple's ecosystem.
Do you also compare cars by looking at only the super expensive limited editions, with every single option box ticked?
I'd also point out that said 3 year old $1999 Mac Studio that I'm typing this on already runs ML models usefully, maybe 40-50% of the old 3000-series Nvidia machine it replaces, while using literally less than 10% of the power and making a tiny tiny fraction of the noise.
Oh, and it was cheaper. And not running Windows.
For training the Macs do have some interesting advantages due to the unified memory. The GPU cores have access to all of system RAM (and also the system RAM is ridiculously fast - 400GB/sec when DDR4 is barely 30GB/sec, which has a lot of little fringe benefits of it's own, part of why the Studio feels like an even more powerful machine than it actually is. It's just super snappy and responsive, even under heavy load.)
The largest consumer NVidia card has 22GB of useable RAM.
The $1999 Mac has 32GB, and for $400 more you get 64GB.
$3200 gets you 96GB, and more GPU cores. You can hit the system max of 192GB for $5500 on an Ultra, albeit it with the lessor GPU.
Even the recently announced 6000-series AI-oriented NVidia cards max out at 48GB.
My understanding is a that a lot of enthusiasts are using Macs for training because for certain things having more RAM is just enabling.
I keep hearing this meme of Mac's being a big deal in LLM training, but I have seen zero evidence of it, and I am deeply immersed in the world of LLM training, including training from scratch.
Stop trying to meme apple M chips as AI accelerators. I'll believe it when unsloth starts to support a single non-nvidia chip.
But for individual developers it’s an interesting proposition.
And a bigger question is: what if you already have (or were going to buy) a Mac? You prefer them or maybe are developing for Apple platforms.
Upping the chip or memory could easily be cheaper than getting a PC rig that’s faster for training. That may be worth it to you.
Not everyone is starting from zero or wants the fastest possible performance money can buy ignoring all other factors.
It's just more efficient to offload training to cloud Nvidia GPUs
And while the Mac Studio has a lot of memory bandwidth compared to most desktops CPUs, it isn't comparable to consumer GPUs (the 3090 has a bandwidth of ~936GBps) let alone those with HBM.
I really don't hear about anyone training on anything besides NVIDIA GPUs. There are too many useful features like mixed-precision training, and don't even get me started on software issues.
You can't lug around a desktop workstation.
I am honestly shocked Nvidia has been allowed to maintain their moat with cuda. It seems like AMD would have a ton to gain just spending a couple million a year to implement all the relevant ML libraries with a non-cuda back-end.
Intel seems like they could have some interesting stuff in the annoyingly named “OneAPI” suite but I ran it on my iGPU so I have no idea if it is actually good. It was easy to use, though!
https://www.technologyreview.com/2019/12/11/131629/apple-ai-...
But you can get an M2 Ultra with 192GB of UMA for $6k or so. It's very hard to get that much GPU memory at all, let alone at that price. Of course the GPU processing power is anemic compared to a DGX Station 100 cluster, but the mac is $143,000 less.
You want your developers to be able to do training locally and they already use Macs? Maybe an upgrade would make business sense. Even if you have beefy servers or the cloud for large jobs.