Hardware-Accelerated TensorFlow and TensorFlow Addons for macOS 11.0
github.com
github.com
INSTALLER_PATH=https://github.com/apple/tensorflow_macos/releases/download/v0.1alpha0/tensorflow_macos-0.1alpha0.tar.gz
So, see: https://github.com/apple/tensorflow_macos/releases/If you mean the "source code" zip in the latter link, that's just github zipping the repo, which contains no useful source code.
https://github.com/apple/tensorflow_macos/archive/v0.1alpha0...
The archive downloaded by the installer script does contain some source code, but it's mostly "generic" TensorFlow code, with some Python stubs that call off to native libraries (as you'd expect). It seems like all of the ML Compute stuff is contained within pre-compiled libraries (with some header files provided), but no source code.
I could be wrong here, and it might be that the intention is to open source the ML Compute components, but I don't think that's been done yet.
It's version "0.1alpha0" indicates it's at PoC state so things may change.
If there's someone who has an educated guess of the performance difference between...
M1's TPU vs 3090's TPU re training
...please let us know.
3090 memory 24GB bandwidth 936.2 GB/s
M1 memory 16GB max bandwidth 68GB/s (shared with other parts of the system)
M1 GPU isn't really that powerful, it's comparable to nVidia 760 (from 2013). The M1's Neural Engine does have more kick of course, but the GPU otherwise is nothing superb (other than marketing).
A more modern comparison with mobility GPU would be GTX 1050 Ti mobility, which is around 10-20% faster than 760: https://gpu.userbenchmark.com/Compare/Nvidia-GTX-760-vs-Nvid...
Even still, 1050 ti uses up around 75w of power (2016) and 760 has a TDP of 170W (2013) while M1 GPU is much less than 10W (at full load, it peaks at 16w in Mac mini for the entire SoC, not just GPU).
It would be interesting to see what Apple does when it scales it up to 75w or more for their own custom desktop GPUs which is rumored in development. However, separate desktop GPU does lose the benefits of UMA that makes M1 fast.
The nice thing about the desktop is that it can just train for days and I don't need to work about using it for other things, losing time moving locations, etc. It's also still probably cheaper than the m1 laptop, even with a nicer gpu than what you can find in the trash.
From Geekbench it also looks like the m1 gpu is about 1/4-1/3 as powerful as a 1080.
The m1 may benefit from faster ram and shared memory though.
Apple states that the neural engine is able to do about 11 trillion operations per second (but oddly enough, they don’t report tflops).
They have comparable TFLOPs range.
1080 mobile is way faster than 1060 (60-90% faster [^1]) and definitely ahead of M1 by a large factor.
[1]: https://gpu.userbenchmark.com/Compare/Nvidia-GTX-1080-Mobile...
Apple's dedicated ML hardware is probably quite good, but we don't have any way to know how good without doing math on die size + power draw and running benchmarks.
So I'd guess it's slower then an 1080 non-Ti
I currently use a 2080tioc and wish for something faster to not wait for runs that long as they take you out of the zone.
Not to mention a Linux based workstation in the limit will have fewer headaches than mac these days. Package management doesn't require homebrew or dockerized everything, selinux is surprisingly easier to configure than the Mac security subsystems, etc.
It doesn't really prove much beyond that you can probably get enough speed on an M1 to debug your training loop. That's impressive but we need to see more to see if the "only 4x slower than a colab GPU (ie at least a K80)" numbers hold up.
Edit: I kept that tab open for too long...
(The TensorFlow implementation has the same limitation, but using graph execution was traditionally more popular in TensorFlow, since it didn't initially have an eager mode.)
https://pytorch.org/tutorials/beginner/hybrid_frontend/learn...
https://pytorch.org/docs/stable/jit.html
But people normally use PyTorch in eager mode.
It might be that no one has written the code to make it work in eager mode, but I'm trying and failing to think a hardware reason this could be the case.
Please note that in eager mode, ML Compute will use the CPU.
However I find that quite disappointing. I hope there’s significant upgrades coming down the pipeline for when they start releasing more powerful Apple silicon based devices.
However, you can just barely finetune any for the base pretrained transformer models (e.g. BERT base or XLM-R base) with 8GB VRAM and need 12GB or 16GB VRAM to finetune larger models. Given that M1 Macs are currently limited to 16GB of shared RAM, I think training competitive models is currently very limited with the memory limitations.
I guess the real fun only starts when Apple releases higher-end machines with 32 or 64GB of RAM.
Also see the benchmarks in the marketing PR:
https://blog.tensorflow.org/2020/11/accelerating-tensorflow-...
The M1 blows away Intel CPUs with integrated GPUs (and modern NVIDIA GPUs will probably blow away the M1 results, otherwise they'd show the competition ;)).
on the other front, I'm hearing great things about Rosetta 2 compatibility so that's not nearly a deal breaker!
Looking good, still all the more reason to wait for additional product offerings (faster processor, more RAM, larger form factor)
It’s going beyond that is tricky.
This lead me to believe that it supported hardware acceleration for the architecture (because it lists Intel too), which made it seem ambiguous if it supports the specific accelerator abilities of the M1.
What "native hardware acceleration" is available on Intel-based macs?
https://developer.apple.com/documentation/mlcompute
I don't think Apple have stated whether MLCompute also utilises the M1's "Neural Engine", but geohot did manage to target it in tinygrad:
There's some MLComputeANE code in the framework, but afaik there's no way to use it yet. Also, from looking at it, I believe the neural engine only supports float16.
Core ML on the other hand is designed to use trained models and runs automatically on CPU, GPU or ANE depending on what fits the currently executed model layer best: https://developer.apple.com/documentation/coreml
Neural engine support is not stated.
Maybe college kids?
Maybe I'm in my own bubble but does TF have small scale uses where you would run it on a laptop?
>Maybe I'm in my own bubble but does TF have small scale uses where you would run it on a laptop?
That, plus huge scale uses where it doesn't make sense to run on a beefy desktop even, so you just use whatever to deploy in a remote cloud/cluster/etc. A laptop means you can do it from wherever, and a laptop with good specs/battery means you can do all other stuff, for longer, with it...