Are the New M1 Macbooks Any Good for Deep Learning?
betterdatascience.com
betterdatascience.com
Also, for me, the answer is no: they're not any good. I use PyTorch. I use custom CUDA code. The reality is brutal but simple: if it doesn't run CUDA, no serious ML research will use it for training anything else than toy models.
They will use it to edit code, deploy and remotely run jobs. Serious deep learning on laptops is a non-starter, on any laptop, on account of heat dissipation.
Being able to do light stuff on one's own consumer hardware without having to buy something new is still incredibly helpful to students and other people trying to learn, as well as hobbyists though.
Training Tesla's FSD neural net on a laptop? A student training some models for courses or self study?
Serious schools have computing clusters. ML researchers might be interested to know which undergrad is actually training models complex enough to benefit from better hardware.
> these still aren’t machines made for deep learning. Don’t get me wrong, you can use the MBP for any basic deep learning tasks, but there are better machines in the same price range if you’ll do deep learning daily.
The article has 4 short sentences in the 'conclusion' section which can be found, as expected, at the end of the article.
it really isn't "buried"
“But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.”
I agree, but from what I have seen they also perform poorly against collab. A 1660 Ti 4gb RAM laptop gpu, runs about as fast as Collab CPU does. [1]
I have not been able to find how the RTX 2080 Super (Laptop) and upcoming RTX 3080 (Laptop) compares to Collab. The 3080 has over 3x as many cores as the 1660 Ti.
[1] https://towardsdatascience.com/google-colab-how-does-it-comp...
According to the chart on that page you linked to the Lenovo Legion (with the 1660ti) is about 30% faster than the Colab GPU. The Lenovo 480s (CPU only, no GPU) is about the same speed as the Colab CPU. So if you can train your model on the on 4GB of VRAM, a laptop with a discrete GPU (if you exclude the really basic business laptop GPUs) may be useful.
It can come with a RTX 2060 (Max Q / low voltage) which supports CUDA; as well as Ryzen 9 4900HS, which was the fastest mobile CPU until the 5000 series mobile Ryzen came out.
The 2060 + 4900HS has gone on sale at Best Buy for $1199 multiple times late last year (just check slickdeals).
I think it’s absolute steal of a machine for that price. Not to mention that it’s beautifully built, is lightweight, has an amazing matte color accurate screen, etc.
There’s also a fan sub for this particular laptop with over 13k subscribers (as of today) on Reddit: https://www.reddit.com/r/ZephyrusG14/
However, if your goal is just to have a machine on which you can actually test your CUDA code, then the Zephyris G14 laptop is perfect. It's portable, has a nice screen, etc.
It's "optional" in the sense that things still calculate correctly on CPU without it, but at a 1000x performance penalty. Or you could skip it if you had 64GB of GPU RAM, which you cannot buy (yet).
So if you actually want to work with this on GPUs that are commercially available, you need it.
Are there any examples where custom cuda code implements some op that can't be written in Pytorch/TF/Jax/etc? That would have provided a better support to your claim that M1 needs to be able to run cuda.
Note the cuda kernels in the original repo were added in August 2017. It might have been the case at the time they needed them, but again, if you need to do something like that today, you're probably an outlier. Modern DL libraries have a pretty vast assortment of ops. There have been a few cases in the last couple of years when I thought I'd need write a custom op in cuda (e.g. np.unpackbits) but every time I found a way to implement it with native Pytorch ops which was fast enough for my purposes.
If you're doing DL/CV research, can you give an example from your own work where you really need to run custom cuda code today?
(downvoted for asking a honest question. Ahhh the things you see on Apple related posts... Emotion driven bunch)
Researchers want to get their work done. They don't want to fight against their tools.
There's also the issue that by relying on proprietary frameworks that work now, you might be painting yourself into a corner if Nvidia changes something in the future and then you have to adapt to them because you have no choice.
When nvidia sees a need, they can change CUDA over night to address it, and they pay people to do that.
When you need to do the same in Vulkan, that’s a multi year process till your extension is “open”. A teaser here whose job is to get something done with ML has better things to do than going through that.
And getting email to the internet at large was no mean feat: I remember having to do a slew of UUCP addressing to get an email from AT&T to the internet (something along the lines of "astevens@redhill3!ihnp4@mit.edu). It was the wild west.
While developing a small competitor to Tensorflow back then (Leaf), we were one of the few frameworks that also tried to support OpenCL, but the additional dev work made it unfeasible.
GPGPU has been a thing for 20 years now, and I still can't easily write code that works on Nvidia and AMD and ship it to consumers on Windows. From what I've seen OpenCL seems to be dying, AMD doesn't care about compute on Windows or on their Radeon cards, and Cuda continues to be the only real option year after year in this growing segment. Why would anyone buy anything else than Nvidia if they're using Photoshop, Blender, DaVinci Resolve, or other compute heavy consumer software? Maybe it's unrealistic to hope that any library can fix this and just rename GPGPU to Nvidia compute and be done with it.
If you're using Blender it can absolutely make a ton of sense to use AMD hardware. For most of the time where Blender supported GPGPU AMD was the best choice, and I set up rendering servers with AMD hardware for that express purpose.
I feel like a big part of this attitude is from not actually having tried it. Because SYCL works fine on AMD. In fact, you have more backend options for AMD than NVidia.
If you control your own hardware and software stack maybe an AMD CDNA card is fine, but if you want to ship software to end users it seems to be difficult to even know what will work. So you use cross platform code for a worse experience on Nvidia and spotty support on AMD, or only Cuda and accept that it's Nvidia only but will give you a better experience.
I haven't done a lot of GPGPU programming, but I've tried to look at it from time to time, and I've been disheartened by it every time. Nvidia's handling of OpenCL, AMD's disregard for SPIR. This is what an AMD representative had to say in 2019 [2]:
"For intermediate language, we are currently focusing on direct-to-ISA compilation w/o an intervening IR - it's just LLVMIR to GCN ISA. [...] Future work could include SPIRV support if we address other markets but not currently in the plans."
[0] https://github.com/RadeonOpenCompute/ROCm/issues/1180#issuec...
[1] https://community.amd.com/t5/opencl/spir-support-in-new-driv...
[2] https://github.com/RadeonOpenCompute/ROCm-OpenCL-Runtime/iss...
oneAPI DPC++ Features Included in SYCL 2020 Final Spec [https://newsroom.intel.com/articles/oneapi-dpc-features-2020...]
Arstechnia's write up on OneAPI provides a good overview [https://arstechnica.com/gadgets/2020/09/intel-heidelberg-uni...]
[edit: added arstechnia reference & link]
Its like Apple sell you the iPhone, but also the iOS APIs and application model so you can run iOS apps. Once you run iOS apps, you are in Apple's ecosystem, both the ISVs and the user are hard to leave. Its like saying when will Apple officially make iOS APIs and libraries run on Android phones. Both will never happen.
Now don't get me wrong, I absolutely adore my M1 MacBook Air. It's so good that it I've kept my i7+2080+32GB RAM desktop turned off for weeks now (outside of gaming). My 8GB M1 MacBook Air is now my main work computer, and it's spectacular.
That said, the basic MNIST timing benchmarks are decent, but don't get your hopes up for anything more anytime soon. There's simply no way in hell today's M1 SoC can train much, much bigger & complex models than hardcore systems with beefy GPUs with 12+ GB RAM on each card.
So no, today, the M1 MacBook Airs/Pros are not really good for DL. And honestly, that's fine. Maybe Apple will compete with NVIDIA down the road, and I'd be happy to see that. But I'm glad the author came to the (correct) conclusion that M1s just aren't anywhere near there yet for anything more than basic DL.
Edit: that all said, there might be an interesting use case for M1 Mac Minis as edge inferencing nodes.
Considering the computationally intensive nature of ML, does it make more sense to train on specialized cloud-based processors?
They have a paid version now that is supposed to be faster and more lenient about disconnecting you, but I recently read a comparison that said it wasnt worth it.
- Much more limited IDE experience if you use any graphical IDE. I prefer Pycharm because it is far superior to pretty much anything else out there when working in Python + Pytorch. You need to use a remote desktop solution like VNC in that situation, which is not remotely close to being on-par with local code prototyping.
Jupyter notebooks I abhor because of how poor they are in relation to any decent IDE. VsCode is the only tool that has a great remote dev workflow but it just isn't near the functionality of Pycharm when you use the latter daily. You also don't directly get the ability to plot and visualize data in something like matplotlib when working remotely which is again an issue. I'll still use it if I have to when VNC isn't snappy enough
- If data security is a concern, everything needs to occur through a VPN which is another intermediate step in your workflow to get started every time you open up your laptop.
- If you have spotty internet access or are traveling, the remote dev workflow suffers immensely.
The other thing I should mention is that outside of the standard data exploration + model training/inference workflow, there are other use-cases where being able to prototype locally is very advantageous. I write a lot of our tooling and internal libraries (for example: a keras-like model training framework in Pytorch) and that involves a lot of pytorch code that references GPUs but testing and prototyping that could easily be done with small models on a laptop. Not being able to do that at all is really annoying and having a native IDE experience when writing a library (especially when you rely on several internally developed libraries) is very critical in my experience.
Again, all of these are not individual deal-breakers, but I feel the pain almost every day.
Not everything needs to be gpt3 scale.
There are some cool packages out there, that detect your emotions and attentiveness while driving.
Aside from the Jetson, I have an 2080 Super on my laptop.
Edit: actually Quattro A6000 cards ship with 48GB vram each, so you only need 22 of those to have 1TB total.
> Radeon Pro SSG (2016) https://www.amd.com/en/products/professional-graphics/radeon...
I guess the lesson to be learned here is that if you want to use an implausibly huge amount of RAM to make a point, a TB is not safe anymore. Go for Exabytes instead, that should be unambiguous for a couple of years.
Not sure what the confusion was.
The M1 GPU is about 1/3 the speed of a 1080TI card when looking at OpenCL score in Geekbench, but may perform faster than that in some cases, due to the shared memory architecture of the M1.
As I understand it, the ANE has very low precision, which makes it unsuited for training in Tensorflow.
installation was a non-issue, The training kept the Mac mini completely silent, having 40% of colab gpu speed is very satisfying for small tests.
pytorch on the other hand is not ready yet (there are installation instructions but they didn't work)
Which seems to align just fine with Apple's ongoing efforts to get rid of all the kernel mode drivers and replace them with user mode drivers over time.
I don't think you should read too much into their entry level chip not supporting external GPUs in the first generation.
Though if you're required to use tensorflow, will you even be able to switch to the bleeding edge version? N=1 but all projects I worked on where customers required tensorflow had to be on version <2
As evident from the benchmarks, the result is very dependant on your network. Some getting huge acceleration boost from the neural processor, while others can't be accelerated. I would suggest try and see approach.
I'm currently using - as many do - a 2080TI for ML training. With this my training is <30min.
I would not use a laptop for training as I no longer use "developer laptops" that are more expensive with lower performance than my Linux desktop, especially now with homeoffice I no longer need a laptop - I understand others have a need for laptops YMMV.
But the benchmarks could give a glimpse for a desktop M1.
What I want to know, is it 0.5x or 2x the speed of a 2080TI.
These Macbooks will be superseded anyway, so I'd rather wait until the software I'm using is fully supported and optimised than to jump into the first generation of M1 Macs with unoptimised / unsupported software running in rosetta.
But right now, in general? No. they are not good for deep learning.
> In general, they’re fantastic.
The unoptimised software and the missing developer tools says otherwise. Especially for users of deep learning, if it lacks the tools they will not use it at all for this use case.
> The question was "Are they good for deep learning" and the answer was no. No one asked about them in general.
Don't you think the answer is to skip the M1 altogether and in general for developers and deep-learning users? At this point, there is no reason on getting an M1 Mac at all since the software required is not even ready for M1 and the hardware will almost certainly be obsolete this year for M2.
I wouldn't want to be an early adopter on a system that has unoptimised software on it and would be running on Rosetta. The answer to the question above lies in whether if the hardware in the newer generation Mac products is powerful enough for deep-learning. In this case, it is not. So just get a desktop with a RTX 3080 instead or wait for an M4 / M5 Mac.
At a certain point, you need just raw throughput, which requires power, which means a bigger chassis with louder fans, which ruins all the aesthetics of macs.
Any comparison with M, other laptops and desktop with various GPU card price points?
What are some current solutions when doing local development without breaking the wallet too much?
A laptop that has the performance of a low-end discrete GPU, but is the size and temperature of a regular laptop, would be a very nice thing to have. Hoping software support for the M1 continues to improve.
I'm not an Apple zealot, but I won't ding them for not being able to support a closed source walled garden API.
- There shouldn't be any comparison to a bare CPU because M1 includes a TPU, a typical CPU doesn't
- Comparing it to Colab's GPU is good because latter includes a TPU but OP should have stated which GPU he got; Google Colab allocates different GPU models
Betteridge's law of headlines again.
They do have an interesting cpu, but the apple target market isn't engineers/scientists/gamers so they traditionally have not done anything interesting with respect to ml and gpu kinds of things. The closest is video.
I think things might get interesting with an arm mac pro if they can add a lot of cores. I wonder if they will add PCIe slots so they can collaborate with other hardware manufacturers, or if they will navel-gaze some more and close it off.
The machine literally has a component in it called a neural engine which Apple advertises as:
> In fact, with a powerful 8‑core GPU, machine learning accelerators, and the Neural Engine, the entire M1 chip is designed to excel at machine learning.
(https://www.apple.com/mac/m1/)
So asking "how well does that really perform" seems like precisely the right question, and to me it'd seem very clear Apple wants a piece of that market.
[1] https://machinelearning.apple.com/updates/ml-compute-trainin...
As for general targeting of engineers the history of MBP pretty shows that they don't really care. Things like virtual escape key on the touchbar, shitty keyboards with sticking keys, hardware designed to not be reparable, and more recently, releasing the M1 models with a spotty backwards compatibility of software out the door is really not something that would be done if you were targeting tech minded people.
Also these laptops have 8 CPU cores (4 high performance, 4 high efficiency) and 8 entirely separate GPU cores. It's a powerhouse of multiprocessing, especially when you consider it's in a laptop that gets 20 hours of real battery life.
It's also fully bidirectional, allowing for the same speed in the other direction at the same time. Thunderbolt is capped at 40Gb/s total bandwidth.