1,344 karma · joined August 27, 2010
Co-organiser of : https://www.meetup.com/Machine-Learning-Singapore/
At first, it's about a robot and another robot doing stuff. Then they're on a journey. Then the desire for friendship. And hospital. Then (gradually) the idea of different motivations comes in. Perhaps some ecology. Then conflict, etc.
All the while, my daughter is delightfully more interested in robots and space travel than standard 'pink princess' fare.
Of course, YMMV.
Also 'findable' : The Clangers (charming 1970s show from the UK); Mr Men; Shaun the Sheep; Mr Benn.
All of these are 'gentle' but fun, and not really educational per-se. Even though you'll end up watching all of these again-and-again, Shaun the Sheep remains entertaining for parents too.
* https://en.wikipedia.org/wiki/Kipper_the_Dog
* https://en.wikipedia.org/wiki/Clangers
* https://en.wikipedia.org/wiki/Mr._Men
The NNSE paper has associated code already, but I found setting the sparseness preference parameter was very hit-and-miss, which is why I preferred the explicit sparse-by-percentage measure in my work.
These techniques are impressive, and yhat is demonstrating that they are very capable. It's just that I feel a little sad that the 'AI pitch' is being turned on, when the 'really good tech' is a much more valid way to understand what they're doing.
One infers from a quick read ~"Algorithms are now like people, and can learn about anything." But careful parsing of the commas shows that the sentence is true, but in the precise sense that "People can learn about anything. Now, algorithms can also learn about anything." - and the extent of learning/understanding is not being compared.
Perhaps I'm nit-picking, but this statement appears to have been constructed to support an AI pitch, and is literally true, but no 'actual AI' is involved (and no-one is actually claiming it is... unless you /want to believe/).
btw : Is this "Correct-by-Construction" seminar [1] relevant?
If so, then (naively) could one pack ~10 on that single FPGA? Or does the 'packing overhead' become a big problem? Or does the design use more (say) multiply units pro-rata, so that they become the limiting factor?
In addition, as a EP-holding business-owner, I was 'asked' to explain my plans to employ Singaporeans/PRs in the coming year.
OTOH, I don't see this as entirely unreasonable : After all, the USA does the same (for E-2 investor visas, for instance). Moreover, the Singapore visa turn-around time is only a few days (after an online application), whereas the US visa process makes one long for the friendliness of the DMV...
[1] http://www.mom.gov.sg/employment-practices/fair-consideratio...
Developer == Essential (e.g. : CoFounder)
in Singapore :
Developer == programmer == glorified typist == someone whose job can be outsourced.
This gets in the way of attracting people to programming in the first place.
The environment is also disappointingly oriented towards local graduates aspiring to become Trainee Managers at Multi National Companies (MNCs). It's not a money thing IMHO - more of the status associated with working at a recognisable brand-name firm.
Clearly, there are exceptional people around, but there's an additional (surprising) hurdle for local tech talent here.
What's neat is that the technique is an almost comically simple way to add extra layers to a network. It's commonly accepted that deeper networks can learn better, but they get very unwieldy/difficult to train as they get deeper.
Roughly speaking (and please correct me if I'm off-base), the paper's technique is to slot in additional layers that that are initially 'identity+', where the new layer then gets trained to hone in on the differences from 'identity'. This training on residuals alone is more stable, since answers near each '~0' starting point are simply as good as the original network - any improvement is a pure win.
So... their winning network has a breathtaking 152 layers (and then ensembles a few of them together).
Is there a paper/presentation that embodies the current best practices/thinking that you could recommend? I'm not trying to be lazy, it's just that there is clearly a lot of retro thinking among the top search results on the net, and it's difficult to separate the wheat from the chaff...
Unfortunately, the focus at the Montreal lab that has a huge influence on its development seems to be (a) 'blocks' for a high-level DNN environment (which is very cool) and (b) CUDA-to-the-max (which is understandable, given Nvidia actively seeds research labs with freeby cards, and -- as evidenced by the article -- is putting a lot of effort into supporting deep learning).
rant start:
It's a shame that OpenCL doesn't get more love. Just the other day there was a cool Clojure GPU project (based on OpenCL) announced on HN. One of the comments was 'will you be building this for CUDA too?'. Rather than pressure open source writers to support closed systems, it would be better to pressure Nvidia to provide up-to-date OpenCL drivers. Newer Nvidia cards are at OpenCL 1.2. And the (somewhat old) OpenCL drivers are always there in an Nvidia install. But does Nvidia ever talk about that : No. It's entirely in Nvidia's interest to encourage everyone to talk CUDA-only. But on a GFLOPs/$ basis, and for the cause of Free, CUDA isn't the right way to go.
rant end.
In addition to the 'word vectors' as inputs, the RNNs illustrated are also iterating over an internal state (flowing from left to right though the same network for each new word) - and this internal state is also an embedding of some kind. But it's going to be very difficult to decipher what each dimension here represents, as it's being built purely as a function of the input word vectors, its own previous state and a NN with initially random weights.
Now, although actual 'brain experiments' have shown that individual neuron (or local clusters) apparently light up when particular thoughts are had (alternatively, cause thoughts to be had), each cluster seems likely to be just one aspect of (say) 'dogginess'. So, one area will correspond to the smell of dogs, others to wet noses, others to being outdoors (i.e. all aspects of the overall 'dogginess' concept) - but these things will all overlap in multiple ways with other concept 'vectors'. Which is how huge spaces of ideas are searched in parallel, rather than sequentially (using, say, an is_doggy_quality symbol).
There are also parallels here with the Numenta Sparse Distributed Representations [2].
Overall, this presentation seems to be probing at the frontier of what works, and how to leverage that up into something that's more about 'general thinking' rather than pattern matching. It also appears to be a thought-piece, rather than a conference presentation (though, of course, Hinton deserves to be heard on just about anything in NNs, IMHO).
[1] http://colah.github.io/posts/2014-07-NLP-RNNs-Representation... [2] https://github.com/numenta/nupic/wiki/Sparse-Distributed-Rep...
"Hinton's internal presentation on AI and Deep Learning" might attract the attention it deserves.
(I'm not really advocating a title change/resubmission, but if someone's reading the comments before going to Google Drive, it may be more of a hint about value/bandwidth)
IMHO, the constraints here (as an armchair engineer) are more that the 'big rocket' isn't very responsive to requested changes in thrust, whereas the 'little nose rockets' may be responsive, but very weak compared to the mass of the thing they're trying to control.