MuGo: A minimalist Go engine modeled after AlphaGo
github.com
github.com
I tend to wonder - beyond "ease of implementation and good 'nuff" reasons - if there are other reasons to use RELU, over other activation functions like TANH or Sigmoid?
I'm beginning to suspect that we may be seeing the "engineering side" of neural networks coming into play; that instead of using the more "biologically accurate" activation of the sigmoid function, we instead use RELU (and other ELU derivatives) because it works well, and is easier to understand?
Much like how things progressed better in heavier-than-air flight once engineers realized that flapping wings weren't absolutely needed, and low-weight engines turning propellers, with fixed wings, worked better for flying than what nature uses...?
https://en.m.wikipedia.org/wiki/Rectifier_%28neural_networks...
Hmm - didn't think about that aspect, but that would be a big plus for backprop as I understand it! Thanks for the insight, and the link.
That said, a lot of the history of neural networks has been brief moments of biological inspiration followed by hacking and engineering that drifts further away from the biology the better it gets. The biggest example is backpropagation; despite how essential it's been to artificial neural networks, it really doesn't exist in the brain, at least not as simply as it does in code. For now, we're all still exploring, some looking towards biology, some towards abstract principles, and it remains to be seen if one provides consistently better results.
Do you have any links to papers or such explaining "at least not as simply as it does in code"?
EDIT: nvm, I followed the links in a wikipedia article on RELU to relevant PDFs...
For a while now this is one area I have been questioning - that we do use backprop, and maybe there is something to be learned from nature that might (?) simplify how a NN is trained (then again, nature might be doing it in such a way that is more complex than can be engineered or practical)...
> For now, we're all still exploring, some looking towards biology, some towards abstract principles, and it remains to be seen if one provides consistently better results.
It might end up being a combination; at least, that seems the direction so far to a point.
I want to thank you for your comments, though. I'm still learning this stuff (I'm working thru the Udacity Self-Driving Car Engineer Nanodegree), and you've given me some stuff to think about and explore further.
Getting inspiration from nature is often useful (convolutional neural networks are loosely inspired by the way information is processed down the optic nerve) but for the specifics of neuron function, the brain is almost certainly doing things that are more complex than is practical to simulate. The brain contains hundreds of billions of neurons, and each one is orders of magnitude more complex than a node in an ANN, interacting with local and distant neurons through all kinds of separate but interlocking electrical, chemical, and long-term genetic actions.
Is there any NN model which tracks a "supply" each neuron has of signalling "molecules", such that a given neuron won't be able to communicate a message of class X if it's out of X-amine, unless/until it receives some from a neighbour?
Or, is there any NN model that allows each neuron to send broadcast messages, ala extracellular ionic neurotransmission, which decay with "physical distance" (requiring each node on the neural graph to have a physical position) and which remain active until "sucked up" by something?
I've always thought these two things—neurons needing to "spend" chemicals for neurotransmission, and neurons being able to observe the local-in-physical-space "loudness" of certain broadcast signals—seemed to have high potential for modelling agentive systems generally, since they seem to replicate other successful systems (markets; pheromonal communication), but I've never heard of either concept being studied in an ML context.
That said, this is only from the perspective of constructing ANN's for computational purposes. There are, of course, many detailed models of various processes in the brain, but this is generally under the heading of Computational Neuroscience[2] (though the field boundaries are blurry). The main software for simulating biologically accurate neural networks is "Neuron"[3], which is used in the Blue Brain Project[4].
[0] https://en.wikipedia.org/wiki/TrueNorth
[1] https://en.wikipedia.org/wiki/SpiNNaker
[2] https://en.wikipedia.org/wiki/Computational_neuroscience
"Towards a Biologically Plausible Backprop" from Benjamin Scellier and Yoshua Bengio (2016) would be a recent paper on the topic.
Thank you for this link; I've found a copy of this paper - and some others - and will be reading them with interest!
In deep learning, I would _generally_ not look towards biology for the reasons behind why things are done as this is usually an after-the-fact explanation. When in doubt, blame the gradients.
https://en.wikipedia.org/wiki/Lasso_(statistics)#Geometric_i...
The derivative of that loss with respect to x would be J'(x) + beta * sgn(x). So using some variant of SGD (which is what basically all neural network training does these days) we would essentially update x as: x = x - alpha * (J'(x) + beta). (The specifics depend on the algorithm, but it doesn't change the result).
So for x to end up as _exactly_ 0, we have to be extremely lucky, which in practice I have never observed. Using L1 regularization definitely leads to small weights, but not to ones that are _exactly_ 0.
You can prove that L1 regularization is equivalent to taking the optimal unregularized parameters, setting to parameters below a threshold to 0 (the threshold depends on the regularization parameter), and penalizing the other parameters.
2) When you use actual numerical optimization techniques with floating point arithmetic, you don't find an exact minimum (global or local). And you don't get exact zeros.
Have you tried this on a real problem? I wouldn't consider MNIST a real problem, but even there you will not get _exact_ zeros. Try it.
This is critical, because producing that one requires implementing the full reinforcement learning. Even if you skip that and use the policy network, you still have the task of playing a few tens of million games.
Learning a value network as big as AlphaGo from public data does not work: you overfit to hell.
Which playout policy is this using? There doesn't seem to be any?
Looks like it's just a neural network player. There's dozens of those already. You don't need to credit AlphaGo if you're only using policy networks for Go: The critical research for that was done at the University of Edinburgh.
IIRC, Monte Carlo Tree Search with "dumber" heuristics than NNs yielded amateur dan level AIs for the first time (somewhere around 2006?). Lately there has also been some AIs that bolt a NN in, and get around 1 stone stronger (which is still miles away from AlphaGo!).
But since this is specifically modelled after AlphaGo, I wonder how it fares against other AIs.