Correcting Intel's Deep Learning Benchmark Mistakes
blogs.nvidia.com
blogs.nvidia.com
Once we start seeing press releases of products that are not available yet being shown to kill the existing competition we'll know that the hype train has officially gone super-product-sonic (that is a hype wave travelling faster than the product releases can support it).
Essentially, anything that needed dynamic parallelism (launching kernels from within kernels, i.e. tasks where you don't know where the difficult/interesting needles are within a haystack), advanced/concurrent scheduling capabilities, or FP64 is going to be much, much better off with Kepler until users can get their hands on GP100 cards. Maxwell is good at neither of those things - it's actually only good at specifically deep learning/neural nets. Which is not every task within the GPGPU space.
... as does nearly every public cloud provider. I agree with most of the article, but you can't fault Intel for benchmarking the hardware that cloud providers are actually offering.
I'm not sure what exactly NVIDIA is doing with their Tesla product line but whatever it is, it's really restricting the availability of recent GPU hardware. Even Azure's GPU instances released this month are using the Kepler architecture from 2012. It's fully two generations out of date now, and that's sad.
I think this blog post is fair game.
[1] http://semiaccurate.com/2016/08/01/nvidia-finally-shows-off-...
The benchmark Intel presented here is as disingenuous as their infamous white paper from 2010: http://pcl.intel-research.net/publications/isca319-lee.pdf
In comparison, a single Knights Landing Xeon Phi will be ~$7K. I know where I put my money. Caveat Emptor.
But Xeon Phi and I go way back here. They've been trying to beat my AMBER GPU code since 2013 or so. Many man years later I believe that a Knight's Corner is now ~35% faster than 2 Xeon CPUs with 1M atoms or more (source: http://adsabs.harvard.edu/abs/2016CoPhC.201...95N)
Meanwhile, the CUDA code has continued to scale with the GPU roadmap and a Titan XP is arguably 9-10x faster than 2 Xeon CPUs. No data is supplied at the low-end for Xeon Phi and I think we can safely assume it's because performance there sucks. (source: http://ambermd.org/gpus/benchmarks.htm)
Xeon Phi? IMO avoid avoid avoid until they start winning head to head 3rd party benchmarking fights like Soumith Chintala's fantastic convnet benchmark data: https://github.com/soumith/convnet-benchmarks
With a single 4x GPU server costing around $7k in total (row 4), you get nearly double the performance you get from spending $28k on four Xeon Phi servers (row 2).
And that's assuming you've spent the time and disk replicating your data on all four of those Xeon Phi servers, or went to a likely relatively large amount of engineering effort to ensure that network IO doesn't bottleneck training.
I'm genuinely interested here because I can't find this anywhere. I don't think it exists personally.
If I manage to access the KNL here, I'll probably run cp2k and gromacs, though single node performance is of limited interest, and ELPA doesn't currently have AVX512-specific support.
http://www.prace-ri.eu/IMG/pdf/wp120.pdf
Even so, right now, little would please me more technologically than a competitive Xeon Phi offering, but while KNL is better than KNC, my inside info says it sucks too (it would have been a lot more interesting, just like Altera's Stratix 10, if it had shipped before GP100 and GP102).
Right now, I have more confidence in AMD GPUs right now than I have in Xeon Phi. This 3rd party benchmark is particularly interesting (and it doesn't look like anyone at NVIDIA is paying any attention to it):
https://techaltar.com/amd-rx-480-gpu-review/2/
Sure, NVIDIA is still in the lead, but not with the ~10x margins they used to have over AMD.
Finally, I figuratively feel like punching the next person who makes the BS scaling argument over raw performance. GPUs scale too if they're coded correctly. And cloud datacenters are the worst place for that given their craptastic ~10 Gb/s interconnect subject to arbitrary network weather effects.
Or butchering Seymour Cray: Your life depends on winning a race, would you bet your life on a 1,350 HP Venom GT or on 20 179 HP Scion FRSs? I mean collectively that's almost 3600 HP, right? Except it's even worse because for GPUs vs CPUs, it's like they priced the Scion FRS like a Venom GT and vice versa.
I wish you luck finding Xeon Phi winning anything but synthetic tests against yesterday's news:
https://www.xcelerit.com/computing-benchmarks/libor/intel-xe...
@modeless the new Azure instances have M60's or you can purchase a 1080 or new TitanX which are both available (although stock has been tight).
https://azure.microsoft.com/en-us/blog/azure-n-series-previe...
Sure I can buy a Titan X for myself (already have), but I can't rent a hundred, or even one, on EC2 or Azure or GCE. And I can't get a P100 yet at all. I don't want to hear NVIDIA claiming unfair benchmarking and citing P100 numbers until P100s are actually available either to buy (and ship immediately) or in the cloud.
A docker container that runs their performance suite would be ideal.
Corporate smacktalk...