Accelerating TensorFlow Performance on Mac
blog.tensorflow.org
blog.tensorflow.org
That's amazing.
If Apple can work with...Google to get their framework changed for the M1, then there is absolutely no excuse for Intel/AMD. They had a decade to fix this.
They deserve their fate.
AMD had been working with Google to get their framework changed for AMD's GPU for more than two years [1], and all their work are upstreamed. Oh, and AMD's ROCm/OpenCL support is really for general computing, i.e. CUDA alternative, unlike the ML Compute here. ML Compute is something Apple created specifically for running neural networks, nothing more, and roughly equivalent to TensorRT / Android NN if you want to compare with other platforms. And it was here because the wall of Apple's walled garden is too high that nobody other than them can effectively optimize NN inference/training on their chip.
Are they getting competitive results?
I have been party to get AMD/Intel's CUDA alternative out the door on some of the ML libraries - which one is it now ? OpenCL...SYCL ...ROCm...PlaidML ? I cant remember.
All this time I was pissed at nvidia - surely they were playing subversive politics to kill all of this. With so many initiatives, surely AMD/Intel had their heart in the right place.
Apple and Google are cutthroat rivals. And they worked together for a just-released chip to get fully working acceleration support.
Here's where it gets sadder for me - Tensorflow has included a GPU accelerated version of Numpy ( https://twitter.com/fchollet/status/1292893864986984448?lang...) . Numpy itself is only accelerated using BLAS/LAPACK which cant leverage GPU all that well.
https://www.tensorflow.org/api_docs/python/tf/experimental/n...
At this point, it is basically a sealed deal - if you're even remotely dabbling in data science, you better be working on a Mac.
Which makes me hate my XPS all the more :(
> At this point, it is basically a sealed deal - if you're even remotely dabbling in data science, you better be working on a Mac.
This is hilarious. You might have missed that the benchmark was missing Nvidia (or even AMD) graphics cards; I can't think of a lower bar for comparing ML performance than against Intel GPUs - perhaps Intel CPUs? While Apple has brilliant engineering, the M1 cannot possibly outperform the obscene number of transistors Nvidia & AMD throw at the task, even in older, mid-range cards. Not to mention power dissipation.
If you're dabbling, you're better off with Google's Colab[1] which has (free) hardware acceleration which is roughly on par with my 3-year-old RX580 for my Tensorflow projects. Colab will work on anything that can run a browser.
0. https://blog.tensorflow.org/2018/08/amd-rocm-gpu-support-for...
The post does include a benchmark for an AMD GPU (Radeon Pro Vega II Duo) on the Mac Pro. Comparing the Mac Pro GPU vs. MBP M1 results, the GPU clearly wins, although in some cases the margin isn't as large as you might expect.
NVidia cared to make CUDA into a polyglot GPGPU programming model, with nice debugging tools where you can do everything like CPU graphical debuggers, then thanks to PTX it was a matter to just add a new backend to your compiler.
Hence why, even though it doesn't get that much press, you can even use flavours of Java and .NET on CUDA.
Meanwhile Khronos kept driving their C only agenda, and when they realized the mistake, came up with SPIR (then SPIR-V after Vulkan was introduced), tried to also cater to the C++ devs (with printf like debugging tools).
All this effort was largely ignored by OEMs, with their lousy tools, thus ending with OpenCL 3.0 being effectively OpenCL 1.2 renamed to sound cool, and the C++ efforts (SYSCL) are now focusing on compute agnostic backends.
The problem wasn't NVidia, rather Intel and AMD did not deliver and all their alternatives to OpenCL are even worse, half backed attempts that always loose steam half way through.
It will be interesting how long the other deep learning frameworks will need to support the M1. Pytorch has not yet achieved comparable performance on a TPU compared to tensorflow.
The neural engine on the Apple A11 wasn't exposed to apps at least at launch, but that's no longer a thing on A12 onwards.
For one, the M1 engine has 16 cores, vs over 1,500 for a 1660 series NVIDA card. I know it may not be apples to apples, but I have a very hard time believing it would be able to keep up with even the most marginal card for training.
This will only be an issue for the first few months of the transition to Apple Silicon. Upcoming updates to the rest of the lineup will undoubtedly have more unified memory
I don’t even want to think what the 32GB and above will cost.
I know that Apple charges a premium for their upgrades but from everything I've seen so far from the new 13" MBP I'll be happy to pay that premium this time around. I have always found the comparisons of ram/ssd in recent Apple MBP's to generic parts to be slightly disingenuous. Raw speeds don't really matter as much as how well everything works together in my experience and the new AS computers seem to take that to an even higher level.
For comparison, NVIDIA just upgraded from 40GB to 80GB.
Training usually uses large amounts of data to get your system to recognize a pattern. It generally uses huge memory and compute and generates a model.
Inference will use that generated model to recognize the pattern in new data. It uses significantly fewer resources and can run either very fast or on a smaller system.
I believe the speed of inference may be affected by the resources available during training, where more speed/memory for training can produce better models.
Let's wait and see what Apple Silicon has in store for the iMac Pro and Mac Pro.
To be clear, you would work with a very small data sample or synthetic data as your objective isn’t to train a model for production use.
Edit: clarity and grammar
You don't have to throw the biggest model you can at any problem you have.
In the case of the Mac Pro, I'm guessing ML Compute is using the GPU.
In the case of the Intel MacBook Pro 13", which as far as I can tell from Apple's site can't be purchased with a discrete GPU, that will be either the Intel Iris GPU, or the CPU.
In the case of the M1 MacBook Pro 13", I'm assuming ML Compute prioritizes the Neural Engine over the GPU (and CPU), but don't know if there are use cases where the GPU would be preferable.
https://developer.apple.com/documentation/mlcompute/mlcdevic...
Key line:
Until now, TensorFlow has only utilized the CPU for training on Mac. The new tensorflow_macos fork of TensorFlow 2.4 leverages ML Compute to enable machine learning libraries to take full advantage of not only the CPU, but also the GPU in both M1- and Intel-powered Macs for dramatically faster training performance.
So, looks like it's faster on both Intel & M1, but the M1 MBP has a much faster GPU than the Intel MBPI don’t know enough about that hardware to hazard a guess about how easy it would be to get that part of the chip involved.
"There is an optional mlcompute.set_mlc_device(device_name=’any') API for ML Compute device selection. The default value for device_name is 'any’, which means ML Compute will select the best available device on your system, including multiple GPUs on multi-GPU configurations."
Hope this is not a stupid question: Does this TF version just work with M1 GPU-Cores / Intels iGPU or also with AMD and supported Nvidia GPUs?
Desktop wise, you can buy a 1080 for <400 on eBay and the latest gen nvidia are obviously much faster for still less cost than the mac pro.
To put it another way, when you're developing ML models (99.9% of the time), it's not like people are writing if statements in CUDA.
It sounds like the performance gains are because they are now using GPU and CPU instead of just CPU.
https://machinelearning.apple.com/updates/ml-compute-trainin...
A workday at Starbucks would cause me head and shoulder pain for days.
Comparing the dedicated ML hardware to the generic just-decode-video-and-animate-windows-at-decent-speeds Intel Iris doesn't really prove any performance gain. There's no apples to apples comparison to be made here. If Apple and Nvidia would finally get over themselves, you might be able to make a fair comparison between the M1 and Intel + CUDA. The Intel "pro" laptop they're comparing against only comes with a 1.7GHz Intel CPU and only Intel's mediocre integrated graphics, so I don't see why you would use that version for anything professional related to ML anyway; at best you'd use it to test your pipelines.
All you can conclude from this is that the new M1 chips perform better at TensorFlow than the Intel chips + GPUs in the previous models of Macbook after Apple made a Mac-optimised version that better leverages GPU power.
Or because you like their approach?
Also, in my experience an optimized XLA export is usually faster than Pytorch.
I agree that in TF 1.x the ergonomics are/were bad, but things improved considerably in TF 2.x. In the case you don't like that either, the tf.keras package offer another more friendlier option that integrates well with the other tf packages (e.g. tf.data, tf.estimators, ...).
Finally, I believe the decision also depends on your use case: If you are just experimenting, maybe PyTorch gets to your results faster. However, I also I think that being able to use TFX (https://www.tensorflow.org/tfx) seamlessly saves you a lot of time when you need to put your models in production.