Running PyTorch on the M1 GPU
sebastianraschka.com
sebastianraschka.com
https://news.ycombinator.com/item?id=31424048
Accelerated PyTorch Training on M1 Mac (pytorch.org)
436 points by tgymnich 2 days ago 143 comments
Also it'll be cool to compare Tensorflow vs PyTorch on metal device :)
Also, even for ops that do not crash, it often returns garbage. I built PyTorch from git this morning:
>>> torch.arange(10, device="mps")
tensor([0, 0, 0, 0, 0, 0, 0, 0, 0, 0], device='mps:0')
>>> torch.ones(10, device="mps").type(torch.int32)
tensor([1065353216, 1065353216, 1065353216, 1065353216, 1065353216, 1065353216,
1065353216, 1065353216, 1065353216, 1065353216], device='mps:0',
dtype=torch.int32)
Many bugs will probably be squashed over the coming weeks, so I am still very exited about MPS support. I have done some preliminary benchmarks with a spaCy transformer model and the speedup was 2.55x on an M1 Pro. Which is quite nice because transformers were already very fast on M1 Macs, thanks to the AMX units (PyTorch links against Accelerate, so the AMX units are used for matrix multiplication). If they can squeeze more performance out of M1 GPUs in the future, it will be a very nice speedup.Did you get only 2.55x on BERT/other transformer vs CPU version?
As for BERT inference performance, I am not sure what kind of speedups you are expecting. The M1 Pro gives me 2.6 TFLOPs in single precision matrix multiplication of 768x768 matrices. M1 Pro GPU performance is supposed to be 5.3 TFLOPS (not sure, I haven’t benchmarked it).
I was under the impression that outside toy problems, you can't really get decent results unless you have a TPU or fleet of GPU's for training. If you just train on a laptop you're shooting yourself in the foot compared to someone who gets the same job done 10x faster in the cloud (usually with free credits).
Obviously inference is still possible on a laptop for many models... but not many 'hacker' tasks involve only inference.
Further, in general, I’m really charmed by the potential of having unified memory; the idea you can test some batch training iterations with a batch size that fills almost all 128MB is a unique capability!
Last, a large amount of unified memory allows to do inference and “prompt engineering” with very large models, locally. E.g. Using GPT-J like models (6B parameters).
(1) development - quickly train smaller networks locally to validate the code. Then offload to a big GPU machine. (2) Smaller convolutional networks and RNNs may be below SOTA, but can still provide a good performance-accuracy trade-off. You could already train these networks pretty fast on an M1 thanks to AMX, but his makes training such networks even faster.
This is most certainly not the case. It's impossible to achieve SOTA on most benchmarks but there's a lot more ML and data analysis done on smaller datasets where you don't particularly need a fleet of GPUs.
https://softology.pro/tutorials/tensorflow/tensorflow.htm
It handles (most) of the installation issues and gives you a GUI. You can however bypass the GUI. Every time you run a command the console shows you what the command line looks like so copy/paste into a terminal and away you go.
I don't know how much you've done with running ml jobs on cloud boxes but it can be very annoying to work with them when you are actively iterating. Especially if you are doing a lot of work outside the cloud box environment. Even turn key consumer products like Colab can fall over if you accidentally press "Back" in your browser... I once lost many hours of work because I inadvertently swiped left :)
Apple will continue improving their performance in this space and once jobs that run for hours can be run overnight... I suspect people will often choose to run overnight simply for the ergonomics benefit.
So absolutely not mature, but having PyTorch working better also on macs is only positive imho.
make it at least 1/10
The fact the you want Apple to support other platforms doesn't mean that they're wrong for not doing so. They same arguments to require Apple to support your preferred API also applies to MS, NVIDIA, AMD, etc and the same arguments to not do that apply them all as well.
So GL and CL had to make a compelling argument for why apple should continue to invest in the technology, and even now they still aren’t able to show a single compelling argument for them being better than any alternative.
Personally I would have rather a single open API, but the open API did not appear to be super interested in competing with the other platforms. The end result is the open platforms turn into a pile of expense, licensing, and engineering for behind the curve technology and API.
But native OpenGL requires apple writing OpenGL implementations for their hardware. It requires apple spending engineering time implementing misfeatures from OpenGL and opencl in order to support APIs that are only every considered the backup API.
And again, anyone could have made a Vulkan emulation layer on metal.
(a great start especially considering the missing nvidia GPUs on intel macbooks)
SYCL: due to the above, not on the GPU. Just on the CPU.
* x86_64 apps also get to see the OpenCL CPU backend in addition to the GPU one. However, arm64 apps only see the GPU.
Apple is #2 world wide for phones at 18% [2]
Apple is #1 for phones in the US with 51% share [3]
Apple is #2 for phones in the EU at 23% share [4]
Apple's laptop seem to wipe the floor in power usage, while not getting anyway more in gaming land than any prior generation [5]
Apple's revenue is the 3rd highest in the world which also doesn't seem particularly irrelevant.
You can hate Apple all you want, but going "I don't like apple, therefore the company is irrelevant" is a dumbass approach. Hell what will android get without finding out what's in iOS?
[1] https://9to5mac.com/2022/04/11/mac-market-bucks-trend-with-c...
[2] https://www.canalys.com/newsroom/global-smartphone-market-Q1...
[3] https://www.imore.com/apple-dominates-us-smartphone-market-5... (This is iMore so take with a grain of salt)
[4] https://www.appleworld.today/2022/02/28/apples-iphone-has-23...
[5] https://www.pcmag.com/news/intel-core-i9-vs-apple-m1-max-whi...
Apple had (runtime and ABI) stable openCL built into macOS and iOS by default and on desktop devs kept requiring you install NVIDIAs or AMDs CUDA systems, and on iOS everyone found that openCL was just as clunky to use as openGL.
Apple made proprietary APIs that were more usable and better interacted with the hardware, and everyone just used those instead, until Apple deprecated openCL and openGL as they merely represent a pile of engineering work and support that is essentially unused.