A new GPU backend for the TVM stack
tvmlang.org
tvmlang.org
Since NVIDIA's volta consumer card is not out yet, I used Titan Xp as the reference card. I grabbed prices from wikipedia, and assume TVM reaches 64% peak perf on Vega and 90% peak perf on Titan Xp:
Radeon RX Vega 64: 12.6TFLOPS * 65% / $499 = 0.01638 TFLOPS/$ Pascal Titan Xp: 12TFLOPS * 90% / $1200 = 0.009 TFLOPS/$
So Vega outperforms a lot here.
A better comparison is probably Vega Frontier with 16GB of RAM for $1000. If you're doing heavy compute, you're probably gonna need a ton of RAM to go with it.
https://www.pcper.com/news/General-Tech/AMDs-HBCC-you-and-me
Out of Nvidia cards only Tesla and Volta have 2x (or more) speed mixed precision ops (in Voltas case actually much faster because of the Tensor Cores) but they are in a very different price bracket.
[1]: https://www.anandtech.com/show/11717/the-amd-radeon-rx-vega-... [2]: https://github.com/plaidml/plaidml/issues/29
The $700 1080 Ti has almost the same performance as the Titan Xp, so why not compare those?
It changes conclusions considerably...
The 1080 Ti has less RAM, and has NO support for 16-bit packed arithmetic. Ultimately, the 1080 TI is a graphics card designed to dominate video games, and NVidia cuts out other features that gamers don't care about.
With regards to the "Compute" sector, its Vega Frontier ($1000) vs Titan XP ($1200) at the low-end at least. NVidia Tesla chips ($4000 to $7000) constitute the higher-end.
Specifically: substantially better drivers for compute.
"The current support on ROCm focuses on the functionality coverage. We have already seen promising performance results by simply adopting existing TVM schedules for CUDA backend. For example, you can try running the gemm test script in the TVM repository and see the result. For two types of cards we tested, the current gemm recipe for square matrix multiplication (not yet specifically optimized for AMD GPUs) already achieves 60% to 65% of peak performance. This is already a promising start, as it is very hard to optimize performance to get to peak and we did not yet apply AMD GPU specific optimizations. We are starting to look at performance optimization and we expect more improvement to come"
You'd either need to dedicate an overwhelming team to shock and awe this (IE support and performance for AMD chips is best in class for all frameworks, period). This is 50 people, at least, who all work really well together, if you want to deliver it in the next year.
Or, you can bide your time, dedicate those resources to better and better hardware, and when the frameworks field shakes out a bit, have a better shot.
None of these people care about running on nvidia gpus (and in fact, the vendors pushing them don't want to be locked in either), so your main concern there is hand tuning and cuda kernel integration they do. The switching cost is something but not huge, and isn't increasing that much over time (unlike the x86 switching cost, for example). So waiting doesn't lose you a lot.
So in the meantime, you make two bets: 1. You try to take over the intermediate IR of frameworks, and make it good enough that people stop writing hand tuned nvidia kernels. This is unlikely to work out, but worth a shot.
2. You slowly decide what customers you want, look at what they are writing hand-tuned kernels for, and try to tackle making the frameworks they use good enough to not need it.
(in case #1 doesn't work out).
You’re right that doing 10 things at once is a recipe for failure, but the reality is, majority of frameworks don’t really matter all that much, and if a couple of solid integrations existed, they could just reuse that work on their own.
AMD has a number of initiatives that haven't panned out as well as NVidia's AI investment. AMD's "HSA" technology is actually quite interesting, although unpopular. AMD's push for HSA has made it faster at Video Rendering tasks (see Blender for instance) and random tasks like GPGPU-accelerated WinRAR decompression or LibreOffice Spreadsheets.
IIRC, AMD Vega beats NVidia Titan XP in Creo and Solidworks benchmarks (CAD programs). So its not like AMD is sleeping on its laurels here, they're just focusing on other, still profitable, corners of the market.
Of course, the current is behind AI, Tensors, and NVidia at the moment. AMD can't afford to fall further behind. The current trend is AI and Machine Learning, and it seems reasonable for AMD to at least get PyTorch running on AMD cards (if not "beating" NVidia, but at least they can play along).
Like literally, NVidia will just pay 3 people to do nothing but argue with those 3 people on mailing lists and keep them from landing patches at a reasonable rate.
It would cost nothing compared to the benefit of slowing your only competitor down in a space worth this much.
(and i've literally seen not-great companies do this to people's open source projects, so ...)
the real preoccupation is that these building blocks need to be as fast or faster than Nvidia's
AMD's management just doesn't seem to be that interested in ML.
If what they have works for them, precisely none of these people will bother.
they have a huge incentive to go so
But then again: AMD Vega 56 / 64 have HBM2 and are under $1000. IIRC, the Vega Frontier Edition is $999 and 16GB of HBM2 at 480GB/s theoretical bandwidth.
NVidia also has an offering with high-speed HBM2 RAM: The Tesla P100, but its way more expensive: $7000 each.
I dunno if there are major benefits of HBM2 over GDDR5x however. Just listing off numbers here. The Titan Xp apparently has more bandwidth from the GDDR5x RAM for example, although the Titan XP is still more expensive than the Vega 64.
------------
If there is some problem that is global memory-bandwidth constrained, then it might be better to run it on AMD Vega 64. After all, you can pretty much afford 7x AMD Vega Frontier editions than the NVidia P100.
Obviously, this very much depends on your workload.
Vega 64 can in theory do 25 TFLOPs half precision.
But as you say there's a large price difference too.
For a market segment that needs 1-8 GPU rigs for ML on a low budget AMD could kill it if they invested in software support and kernel optimisation.
For servers and large scale training, unless AMD has some ML specialised cores in the pipeline, Nvidia Volta and Google TPUs have a serious lead.
But NVidia markets Tensor cores as:
> New Tensor Cores designed specifically for deep learning deliver up to 12x higher peak TFLOP/ss for training, and 6x higher peak TFLOP/s for inference
https://devblogs.nvidia.com/parallelforall/cuda-9-features-r...
I wouldn't be surprised if they were inflating the numbers slightly, as is common in a lot of marketing material.
https://devblogs.nvidia.com/parallelforall/programming-tenso...
Google's TPUs also do low precision training (with some special version of TF). https://cloud.google.com/tpu/