Microsoft releases CNTK, its open source deep learning toolkit, on GitHub
blogs.microsoft.com
blogs.microsoft.com
I'm looking forward to a review of the major systems available by someone who has a clue (i.e. not me). The investment to learn one of these is great enough that it will be worth some time invested to understand what each can do.
Sigh...
./configure
Defaulting to --with-buildtype=release
Cannot find a CPU math library.
Please specify --with-acml or --with-mkl with a path.
"ACML End of Life Notice: We have transitioned our math libraries from a proprietary, closed source codebase (ACML) to open source solutions"[1]
[1] http://developer.amd.com/tools-and-sdks/archive/amd-core-mat...
[1] https://mran.revolutionanalytics.com/documents/rro/installat...
Compare and contrast with ~1 TFLOP ~$7,000 Xeon CPUs.
I have brought this up with multiple Intel engineers and for the most part they nod and agree. Then they tell me that there's no way Intel would ever start doing things like NVIDIA does here. And then I nod and tell them why I continue to bet on NVIDIA for the immediate future, sigh...
I think this recent open source push by microsoft is depending on the truth of that statement. I think they're working under the assumption that the recent push of developers using Linux has more to do with those developers wanting to use a superior open source environment than preferring Unix/Linux as an operating system.
All of their open source stuff is really easy to use, IFF you're also using all their other open source software and Visual Studio.
If you're the type of developer who is only using Linux because it has the least path of resistance to using open source libraries and software, they're making good progress towards getting you back into a microsoft ecosystem.
If you're the type of developer who likes the free software philosophy, they're not trying to grab you, because they feel that's not a sizable portion of the people using Linux.
I think they're probably right.
[1]: http://systemml.apache.org/
[2]: https://singa.incubator.apache.org/
Yann LeCun states on Nov 2015 (29:00 min mark) GPU's short lived in Deep Learning / CNN / NN https://www.youtube.com/watch?v=R7TUU94ir38
https://www.altera.com/en_US/pdfs/literature/solution-sheets...
"...Worry about scaling; worry about vectorization; worry about data locality...." http://www.hpcwire.com/2016/01/21/conversation-james-reinder...
Nvidia Chief Scientist for Deep Learning was poached from Intel https://il.linkedin.com/in/boris-ginsburg-2249545?trk=pub-pb...
http://www.cs.tau.ac.il/~wolf/deeplearningmeeting/speaker.ht...
AVX-512 instructions https://software.intel.com/en-us/blogs/2013/avx-512-instruct...
Is there anyone from the Microsoft team here that can explain this decision?
--
[1] See examples on https://github.com/Microsoft/CNTK/wiki/CNTK-usage-overview
[2] See examples on https://www.tensorflow.org/versions/0.6.0/get_started/index....
Quoth: "Models can be described and modified with
• C++ code
• Network definition language (NDL) and model editing language (MEL)
• Brain Script (beta)
• Python and C# (planned)
"
[1]http://research.microsoft.com/en-us/um/people/dongyu/CNTK-Tu... thanks to @sharms
Hopefully this is a viable alternative, I would love to see a online course in machine learning leveraging this. I found http://research.microsoft.com/en-us/um/people/dongyu/CNTK-Tu... on the homepage which looks well put together to start off
I think this is now fixed upstream
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1494095
No idea what the catch is :)
If so, while very cool, that's not a general solution. Scaling batch sizes of 256 or lower would be the breakthrough. I suspect they get away with this because speech recognition has very sparse output targets (words/phonemes).
Too bad the code below isn't open-source because they got g2 instances with ~2.5 Gb/s interconnect to scale:
http://www.nikkostrom.com/publications/interspeech2015/strom...
Training data throughput isn't the right metric to compare -- look at time to convergence, or e.g. time to some target accuracy level on held-out data.
There are many :)
http://icri-ci.technion.ac.il/events-2/presentation-files-ic...
http://icri-ci.technion.ac.il/files/2015/05/00-Boris-Ginzbur...
Nvidia Chief Scientist for Deep Learning was poached from Intel ICRI-CI Group https://il.linkedin.com/in/boris-ginsburg-2249545?trk=pub-pb...
http://www.cs.tau.ac.il/~wolf/deeplearningmeeting/speaker.ht... Look for quote: "...In a very interesting admission, LeCun told The Next Platform ..." http://www.nextplatform.com/2015/08/25/a-glimpse-into-the-fu...
Yann LeCun states on Nov 2015 (29:00 min mark) GPU's short lived in Deep Learning / CNN / NN https://www.youtube.com/watch?v=R7TUU94ir38
https://www.altera.com/en_US/pdfs/literature/solution-sheets...
I think that is all that has been revealed. I suspect that this will be announced within a couple of months and this cntk release has something to do with that
"...Worry about scaling; worry about vectorization; worry about data locality...." http://www.hpcwire.com/2016/01/21/conversation-james-reinder.... http://www.hpcwire.com/2016/01/21/conversation-james-reinder...
Nvidia Chief Scientist for Deep Learning was poached from Intel https://il.linkedin.com/in/boris-ginsburg-2249545?trk=pub-pb...
http://www.cs.tau.ac.il/~wolf/deeplearningmeeting/speaker.ht...
AVX-512 instructions https://software.intel.com/en-us/blogs/2013/avx-512-instruct...
Look for quote: "...In a very interesting admission, LeCun told The Next Platform ..." http://www.nextplatform.com/2015/08/25/a-glimpse-into-the-fu...
Yann LeCun states on Nov 2015 (29:00 min mark) GPU's short lived in Deep Learning / CNN / NN https://www.youtube.com/watch?v=R7TUU94ir38
https://www.altera.com/en_US/pdfs/literature/solution-sheets...
http://icri-ci.technion.ac.il/events-2/presentation-files-ic...
http://icri-ci.technion.ac.il/files/2015/05/00-Boris-Ginzbur...