A Convolutional Neural Network for Modelling Sentences (2014) [pdf]
arxiv.org
arxiv.org
There's some pretty good stuff there. I really liked Crypto-Nets: Neural Networks over Encrypted Data[1]:
The problem we address is the following: how can a user employ a predictive model that is held by a third party, without compromising private information. For example, a hospital may wish to use a cloud service to predict the readmission risk of a patient. However, due to regulations, the patient's medical files cannot be revealed. The goal is to make an inference using the model, without jeopardizing the accuracy of the prediction or the privacy of the data.
http://deeplearning.net/tutorial/
And this book (work in progress):
[1]: Wiki with code, exercises and explanation
[2]: Video lecture one with a recap on back-propagation
[3]: Video lecture two on Sparse Auto Encoders
[4]: Handouts
[1]: http://ufldl.stanford.edu/wiki/index.php/UFLDL_Tutorial
[2]: http://www.stanford.edu/class/cs294a/video1.html
I got frustrated by a paper yesterday which contained function definitions, summations-of-summations, products of sequences, convolutions, set theory, switching back-and-forth between unary-functions/vectors and binary-functions/matrices, converting back-and-forth between {0, 1}, {-1, 1} and {true, false}, weighting elements of a set by 0/1 instead of taking a sub-set, linear programming, etc.
What was their result? To speed up pair-wise comparisons of structured data, only do N% of the comparisons and it will only take N% of the time. To decide which comparisons to discard, see what works well on a small sample of inputs.
A cold-splash-in-the-face intro for laypeople can be found in the TEDx talk of Jeremy Howard (founder of Kaggle):
The wonderful and terrifying implications of computers that can learn – https://www.youtube.com/watch?v=xx310zM3tLs
For cutting-edge research, it seems the "NIPS" conference each December is where many of the new results appear:
My hunch (as a total outsider) is that anything Google publishes is about 2 years behind their current best practices.
obviously writing it up does mean the specified results will be a bit behind what they can actually do in their labs, but that's true of everywhere
As an example model, imagine all three are years ahead of published material, and each has policies that encourage publishing only the minimum necessary to "take the crown" (because any more would risk diluting proprietary advantages).
That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marquee conferences – it wouldn't necessarily force an acceleration of the "published state-of-the-art" up to the level of all of their "internal states-of-the-art". (Perhaps, for a group in trailing/catch-up position, their published results will be very close to their best. But the leader could be arbitrarily further ahead.)
By analogy to English auction bidding: outsiders only learn the second-highest reserve price just before the end, and never learn the winner's true reserve price. But here there's no end, and there's always potential competitive reasons for the top-N pack to hide some of their leading-edge practices.
This doesn't require explicit coordination… but there could be explicit coordination, too! All the programs are staffed by former students/colleagues/coworkers of each other.
If they aren't coordinating very carefully, a single defector would result in the entire slack being used up by a single announcement; and the bigger the slack used up, the more a PR win it is...
If these techniques are commercially valuable and in current use – and I believe they are! – then no leading company, or even member of the leading pack, would want to do a current-best reveal. They all have competitive reasons to do only carefully vetted, incremental reveals of somewhat-older work.
For example, submitted to HN sometime ago was the blog of a dude trying to make Deepmind's 'neural turing machine' work; he was having a hella time because the published paper seems to be unclear or skip over a number of crucial points. Or more historically, the German chemical giants made an art of filing patents on all their key techniques, to gain IP protection, but leaving out enough crucial details that when the USA gleefully seized their IP rights during WWI, the American companies discovered they couldn't make the processes work.
Given that intent, there's no way to deduce from the pattern of new claimed results whether what's being revealed is a merely a few months, or many years, behind.