HNHacker News
TopNewBestAskShowJobs

jimfleming

477 karma · joined August 22, 2011

<Something new> Prev. Direct, Machine Learning @ Attentive Prev. Founder, Fomoro Research

twitter.com/@jimmfleming

submissionscomments
jimfleming··on Artificial Neural Nets Grow Brainlike Navigation Cells
To draw too many parallels here would be like comparing stick figures to still life paintings and proclaiming "They're both flowers!" While it might be true you won't learn much about still life paintings from stick figures.

At best this research says something about the task of navigation and optimal representations for that task rather than anything profound about neural networks other than they can both optimize for some task—which should surprise no one.

jimfleming··on Lime: Explaining the predictions of any machine learning classifier
SHAP[0] is another model-agnostic method for interpreting predictions. It's a bit newer and builds on LIME, Shapely, and a few other works. There's also an associated tool[1].

[0] https://arxiv.org/abs/1705.07874

[1] https://github.com/slundberg/shap

jimfleming··on PointCNN – A simple and general framework for feature learning from point cloud
If I'm understanding correctly (I haven't read it in depth yet):

Points are processed as local clusters. Each cluster of points is ordered according to a transformation matrix X. This produces a "canonical" ordering so the processing of points in these local clusters do not need to be invariant to the order since the order now has consistency and meaning.

It's kind of like placing each point into a regular grid like an image before running the convolution. The trick is which points to put in which grid cell which is determined by the transformation matrix X. This takes advantage of locality which PointNet does not, if I recall correctly. By acting locally and stacking many layers of these you can produce a hierarchy of more and more abstract clusters of points, each with an inherent relationship to nearby clusters of points. In addition, the transformation matrix also appears to act as an attention over the points in the cluster.

jimfleming··on Creating believable crowds in Planet Coaster
If you're interested in this, RVO (reciprocal velocity obstacles) is another multi-agent system for collision avoidance with some good reading[0][1][2]. I've been porting the RVO2 library to Python3 over break and it's incredibly natural the way the particles float around each other.

[0] http://gamma.cs.unc.edu/RVO/ [1] http://gamma.cs.unc.edu/RVO2/ [2] http://gamma.cs.unc.edu/HRVO/

jimfleming··on Why Deep Learning surprises me
Adversarial examples don't really support the claim that deep models are just memorizing examples. If they were, they wouldn't generalize to unseen examples at all. However, the human brain is also susceptible to adversarial examples (e.g. optical illusions). Yet human brains still generalize quite well. Likewise, deep learning can both suffer from adversarial examples and generalize well.

Generalization is a multi-axis scale, not a switch: you can have more or less generalization in many different dimensions. Being terrible at adversarial examples just means that axis is weak.

jimfleming··on Why Deep Learning surprises me
It's more like giving you a bag of random house numbers and instructing you to place the numbers on the correct houses in an area you've never been to before. An instructor teaches you where some of the numbers go, and you can memorize those examples, but when the instructor leaves you to finish the job on your own you have no way of knowing how to assign the remaining numbers.

Memorization is pretty easy. Generalizing from past examples requires that there be a relationship not just between one person and their phone number but between all people and their phone numbers.

jimfleming··on Why Deep Learning surprises me
That is not a conclusion that can be drawn from the findings in the paper. While the models they evaluate can achieve zero training error on random labels, the test error is obviously not zero: it doesn't generalize at all. However, training on real labels often finds solutions which can generalize quite well.

A better way to summarize the central question of this paper would be: "Why is it that a large-parameter model trained with gradient descent on real data _could_ just memorize all of the training data (it has the capacity) yet finds solutions which generalize well to an unseen test set?"

To say that deep learning is _just_ memorizing its training data would be incorrect. We have empirical evidence to the contrary and this paper is part of that evidence.

jimfleming··on Self-Normalizing Neural Networks
Regarding your edit, the authors of the paper in question focus on FNNs and note the reason in the paper:

> Both RNNs and CNNs can stabilize learning via weight sharing, therefore they are less prone to these perturbations. In contrast, FNNs trained with normalization techniques suffer from these perturbations and have high variance in the training error (see Figure 1).

Essentially FNNs stand to benefit more from this work than CNNs or RNNs.

jimfleming··on How the TensorFlow team handles open source support
Here imperative style means you can use language constructs in e.g. Python, such as if-statements and while-loops, to directly construct models. TensorFlow, Theano, and other graph-based frameworks typically require creating branches and loops as nodes in the graph.

Graphs have advantages but they can be unfamiliar and sometimes difficult:

Some advantages: easier to serialize the whole graph and distribute computation; optimization can also be performed across the graph nodes (e.g. see XLA in TensorFlow).

Some disadvantages: it can be more difficult to write and reason about, particularly for recurrent neural networks which can utilize loops a lot; also interop with reinforcement learning environments where much of the computation is performed in an environment outside of the graph.

jimfleming··on MuGo: A minimalist Go engine modeled after AlphaGo
The sparsity encouraged by ReLUs is one reason they are used but not to the exclusion of all other activations. ReLU variants can indeed outperform standard ReLU[0] but sometimes sparsity is a desired property. For example, in generative models.

[0] https://arxiv.org/abs/1505.00853

jimfleming··on MuGo: A minimalist Go engine modeled after AlphaGo
Nature uses something like STDP (spike timing dependent plasticity), where the change in a synapse strength is proportional to the sign and delta time of the last spike at each end of the synapse. You can approximate something like backprop by using symmetric STDP (discarding the sign).

"Towards a Biologically Plausible Backprop" from Benjamin Scellier and Yoshua Bengio (2016) would be a recent paper on the topic.

jimfleming··on MuGo: A minimalist Go engine modeled after AlphaGo
The reasons to use ReLUs are sparsity and improved gradient flow. ReLUs encourage sparsity because when the input to the ReLU is less than 0 as the activation becomes 0. This means some fraction of activations in a given layer will be omitted which can encourage better representations. They also have improved gradient flow because the gradients are zero or constant and thus don't suffer from vanishing/exploding gradients.

In deep learning, I would _generally_ not look towards biology for the reasons behind why things are done as this is usually an after-the-fact explanation. When in doubt, blame the gradients.

jimfleming··on Parallelizing Word2Vec in Multi-Core and Many-Core Architectures
Titan X is the product line which has multiple generations, with the Pascal architecture being the latest.
jimfleming··on FluidNet – Accelerating Eulerian Fluid Simulation with Convolutional Networks
This looks like a great summary, thank you.

My reading of [2] is that the model only needs to be retrained if the environmental boundary changes. For example, to handle objects that do not fit within the current boundary. Since I imagine scale matters here, and you couldn't simply make everything inside smaller. Internal geometry such as replacing the castle with a car does not appear to require retraining.

jimfleming··on Neural Symbolic Machines: Learning Semantic Parsers with Weak Supervision
A GRU is more like a simplification of the ideas in LSTM, rather than a building block. At a high level, it uses the hidden state as the memory of the cell (rather than a separate cell state) and it uses a single "update" gate, merging the forget and input gates. Overall it performs similarly to LSTM while being more computationally efficient (fewer matrices).
jimfleming··on Numerai – A hedge fund built by a global community of anonymous data scientists
Spent some time experimenting with Numerai. Really fun competition, clean (encrypted) dataset, and Bitcoin payouts. I wrote about my experience and open-sourced all of the models here[0] if you're looking to get started.

[0] https://github.com/jimfleming/numerai

jimfleming··on Don't Read the Comments
Interesting, I read this as encouragement for founders rather than complaining.

The point is that there will always be haters, so being discouraged by them or complaining isn't fruitful.

jimfleming··on People for the Ethical Treatment of Reinforcement Learners
I'm honestly not sure. I think it's serious (they seem pretty cautious in their statements which you wouldn't do in satire). But I did find it interesting regardless of it's seriousness.
jimfleming··on OpenAI technical goals
You should have a look a PETRL[0]: People for the Ethical Treatment of Reinforcement Learners :)

[0] http://petrl.org/

jimfleming··on OpenAI technical goals
It was an off-hand remark. I'm aware of the landscape, though perhaps slightly more optimistic. The first goal is simple enough and largely underway with the Gym. Significant progress has been made on #3 and #4 just in the last year but I agree that "a few years" is a bit brief. I remain doubtful about #2.
jimfleming··on OpenAI technical goals
This is great! The goals seem reasonably ambitious and mostly doable over a few years.

I am surprised by #2: "Build a household robot". It's my understanding that efficient actuation and power are largely unsolved problems outside of the software realm. What's the plan for tackling stairs, variable height targets, manipulator dexterity, power supply, etc. in a general purpose robot with off-the-shelf parts? (Answering these questions may be part of that goal but maybe someone knows more on the subject.)

jimfleming··on Perceptrons – the most basic form of a neural network
Look into computational neuroscience, neural engineering, spiking neurons and spike-timing dependent plasticity (STDP). These are more accurate models of biological neurons and their synaptic interactions.

From the Hodgkin-Huxley (HH) neuron model to integrate and fire (I&F) you can approximate and simulate various levels of realism in artificial neural networks. There's even the Nengo library[0] which can simulate these for you in Python. Unfortunately, they're very hard to work with compared to perceptrons, artificial neural networks, and deep learning (see Neural Engineering[1] for a good framework). Additionally, they're far more computationally intensive with the HH model coming in at over 800 FLOPS to the I&F model (which lacks a lot of the behavior and dynamics of biological neurons) which has 7 FLOPS. Then try representing numbers with spiking neurons... You need several dozen, working as a population[2], to accurately represent a single floating point because they're so noisy. This makes working with data pretty difficult and expensive.

[0] https://github.com/nengo/nengo

[1] http://compneuro.uwaterloo.ca/research/nef.html

[2] https://en.wikipedia.org/wiki/Neural_coding

jimfleming··on AI, Deep Learning, and Machine Learning: A Primer [video]
I agree, but I guess my point (in this comment and others) would be that we should stop thinking of intelligence, consciousness, free-will and other attributes as a hard line but rather gradients or quantities.
jimfleming··on AI, Deep Learning, and Machine Learning: A Primer [video]
> Only at such a high level of abstraction as to be meaningless.

I'm not sure what this means or how the abstractions are meaningless? From Gabor filters to concepts like "dog", the abstractions are quite meaningful (in that they function well), even if not to us.

> They are not. Hundreds of man years worth of engineering time go into each of those systems, and none of those systems generalizes to anything other than the task it was created for. That's nothing like human intelligence.

This isn't strictly true if we look at the ability to generalize as a sliding scale. The level of generalization has actually increased significantly from expert systems to machine learning to deep learning. We have not reached human levels of generalization but we are approaching.

Consider that DL can identify objects, people, animals in unique photos never seen before and that more generally the success of modern machine learning is it's ability generalize from training to test time rather than hand engineering for each new case. Newer work is even able to learn from just a few examples[0] and then generalize beyond that. Or the Atari work from DeepMind that can generalize to dozens/hundreds of games. None of those networks are created specifically for Break Out or Pong.

It's also not entirely fair to discount the hundreds of years of engineering considering most of these systems are trained from scratch (randomness). Humans, however, benefit from the preceding evolution which has a time scale that far exceeds any human engineering effort. :)

[0] https://arxiv.org/abs/1605.06065

jimfleming··on AI, Deep Learning, and Machine Learning: A Primer [video]
For #2 you've touched on the "AI Effect"[0] or moving goal posts.

[0] https://en.wikipedia.org/wiki/AI_effect

jimfleming··on My path to OpenAI
The core learning method in biological neurons is believed to be something like STDP (Spike-Timing-Dependent Plasticity). Basically, the arrival time of a spike at the post-synaptic end of a neuron is compared to the arrival time of a spike at the pre-synaptic end. The sign and scale of the difference of arrival times will cause a change in the respective synaptic strengths.

It has similarities to backprop, and depends on backpropagated signals but backprop is simpler, faster and can exploit knowledge about the loss function (as far as we know STDP cannot). The downside (?) to backprop is that it cannot exploit temporal information like STDP but several recurrent models have found ways around that.

jimfleming··on Real-Time Texture Synthesis with Markovian Generative Adversarial Networks
Games. Good texture synthesis can augment or replace the texture creation process in games which often rely on hand-painting and manual seam-removal in Photoshop[0][1].

[0] https://www.allegorithmic.com/products/substance-painter [1] http://quixel.se/

jimfleming··on Some Starting Points for Deep Learning and RNNs
"Solved" is pretty broad (as is "AI") but deep learning, specifically, performs well (SotA) on a number of benchmarks in speech, language and image recognition. Some challenges immediately come to mind:

1. The pace of published research is pretty fast right now. This makes it difficult to know where the research fits in when solving problems. It'll probably take a few years before we know where to use many of the approaches published last year.

2. Iteration performance (trying many new things quickly) is improving with high-level frameworks but still lots of work to do here. Since we don't know where a lot of research fits it's not always apparent which methods work best (and "best" changes every quarter).

3. We're still missing theoretical foundations for much of deep learning. This is useful, not just for research, but to know what can and cannot work with current approaches.

4. Model architectures are still based largely on trial-and-error, intuition and search.

jimfleming··on XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
For one, more compression during parameter transfer in data parallelism scenarios.
jimfleming··on Deep Learning Is Going to Teach Us All the Lesson of Our Lives
I don't agree with the premise of the article but the 5-0 results refer to Fan Hui's match ("Europe’s top Go player").

The article describes the outcome of Lee Sedol's match correctly:

> Lee went on to lose all but one of their match’s five games.

← PreviousPage 2 of 4Next →