Intel Announces Knights Mill: A Xeon Phi for Deep Learning
anandtech.com
anandtech.com
There is a lot of software infrastructure being built atop these frameworks, and switching costs are getting higher by the day. No one wants to use some kind of 'non-standard' fork of [name your DL framework of choice] customized for Intel hardware, because such a fork can quickly get stale in comparison to the upstream project.
Intel needs to be both better/faster and drop-in compatible with the popular frameworks.
Granted, you do have to bother to make a fast x86 / AVX-512 port of your code. But because the shape of GPUs is so different than CPUs -- GPUs have a more complicated memory hierarchy for one thing -- I kind of doubt that "just run your CUDA code in the Phi" is going to work well for nontrivial examples.
(Disclaimer, I work at Google on CUDA support in clang. Which is awesome, you should try it out. :) Google "cuda clang" for instructions.)
I believe Caffe and Theano uses the same model, but I didn't study it. There are also some similarity in the model to what OCaml does with the incremental library, though it is not for machine learning.
That is exactly the problem: no one has bothered to make a fast x86 / AVX-512 port for any of the most popular frameworks (at the upstream level, not in some fork), and no one has an incentive to bother, other than Intel. For example, as far as I know, none of the popular frameworks take advantage of Intel's MKL out of the box.
Right now, if you want out-of-the-box high performance, Nvidia hardware is your only practical choice.
The problem is that as far as I can see, OpenCL is in no way that. Basically, OpenCL gives me the impression that the oceans of boiler plate required both make development hard and effectively locks you into a specific vendor also since the boiler-plate is going to be setting things up for one's specific vendor.
OpenCL falls down in terms of standard libraries such as cu{dnn,sparse,blas} but if you're writing everything from scratch it's fine.
I can see simple, comprehensible 20-50 sample code for cuda that does most simple tasks. With OpenCL, I get references to version, boiler-plate, mode with nothing that boils down to simple code.
If you have a simple sample, you should post it here or blog about it.
Later, it's straightforward to port to c or c++ if that's your thing, though I find having numpy et al handy even in production code
Whereas CUDA supported C++ and Fortran from day 1, with the PTX support added a few versions later.
Also the debugging tools, from the presentations I have seen, are much more developer friendly on CUDA.
Of course developers rather use APIs that offer more modern experiences than ones still stuck in pure C, with a compiler at the driver level, forcing each programmer to writer the boilerplate to compile and link.
Now it might already be too late for OpenCL in spite of the latest improvements.
In that context, x86 compatibility is less of an issue than the quality of the compilers. As long as you are easier to program than a GPU, you are good.
Thankfully this sort of situation is becoming less common over time, but we're not all the way done yet.
http://wccftech.com/amd-cuda-compilercompatibility-layer-ann...
LLVM also has an AMD GPU backend, and says this thing is built on clang/llvm.
So i suspect it's based on that support :)
But yeah this isn't production ready
So I honestly don't get the Google clang CUDA compiler right now. It's really really cool work, but I don't get why they didn't just lobby NVDA heavily to improve nvcc. With the number of GPUs they buy, I suspect they could have anything they want from the CUDA software teams.
However, if it could compile CUDA for other architectures, sign me up, you'd be my heroes.
For I'd love to see CUDA on Xeon Phi and on AMD GPUs (I know, they're trying). And if Intel poured the same amount of passion and budgeting into building that as they are pouring into fake^H^H^H^Hdeceptive benchmark data and magical powerpoint processors we won't see for at least a year or two (and which IMO will probably disappoint just like the first two), they'd be quite the competitor to NVIDIA, no?
That said, the Intel marketing machine seems to have succeeded in punching NVDA stock in the nose the past few days and in grabbing coverage in Forbes (http://www.forbes.com/sites/aarontilley/2016/08/17/intel-tak...) so maybe they know a thing or two I don't.
Are we talking startups? A lot of startups know python so that would make sense...I'd love to see some actual stories though.
CUDA seems to me like the only simple SIMD-type computing system that's fairly straightforward to program and understand at this point.
Drop-in CUDA compatibility seems like a good thing.
The post I am replying to.
Why didn't Larrabee fail? https://news.ycombinator.com/item?id=12293308
If anyone would like to know more about Intel, I think this AMA is much better
https://www.reddit.com/r/IAmA/comments/15iaet/iama_cpu_archi...
Also more than hardware how do Intel's libraries compare with CuDNN?
At the end of the day ease of use and software support matter along with the hardware.
Nervana has its own silicon, but I doubt they will tape out.
I would like to play with these things.
If you want to get your 'feet wet', then why bother?