AMD isn't even mentioned in this article, and I recently read that nobody uses AMD for machine learning. Is this only down to the proprietary Cuda having more momentum than OpenCL? I've used both, albeit not very much, and my impression was that Cuda was a bit easier to write but OpenCL didn't have any deal breakers. I know AMD cards were the most popular for crypto currencies, so it's not like they're always the slowest. Granted, I only have experience with each company's consumer cards, and I think Nvidia holds their GeForce cards back more than AMD does their Radeon cards concerning compute performance.
Why is Nvidia this dominant, and why are so many researchers using a proprietary API like Cuda when there's open alternatives?
Related to this, are anybody using Intel's Xeon Phi? Tianhe-2 [0], the worlds fastest supercomputer since June 2013 use them as co-processors, but I rarely if ever hear about them in other projects.
[0] https://en.wikipedia.org/wiki/Tianhe-2