Why Artificial Brains Need Sleep
discovermagazine.com
discovermagazine.com
However, noise injection on its own has been extensively shown to decorrelate signal from noise so that the networks don't fixate on noise in the input data while training. This is widely used when using CNNs on image data and is known to increase performance of network predictions. So I am not sure how much different this article's findings are from this.
No link to arxiv publication makes it harder to make sense of these findings.
Edit: "Sleep" is not relevant for most ANN architectures used today. Synaptic pruning however is, similar to dropout.
Then, gradient descent is a centralized and supervised learning algorithm that has nothing to do with all the decentralized and emergent processes that take place in biological brains. For example, consider all the complex interactions of the various neurotransmitters and neuroreceptors.
Biological brains (let alone human) are actual physical realizations of networks with very complex and specific topologies. New connections grow and are formed in a very real sense, and this process interacts with the environment in a variety of ways, plus it is spatially embedded.
The problem is we don't understand how the brain is architected at anything other than a very coarse level. That architecture is the product of millions of years of evolution so we have some catching up to do.
And more generally, it’s pretty clear that there is a mixture of context-sensitive inhibitory and excitatory signals both deriving from the dendrites (inputs) and from the chemical stew of hormones and so forth they find themselves immersed in.
There are neurobiologists who claim (controversially, I concede) that a single neurone is a fully self-contained ‘processor’[1], and not the equivalent of single logic gate.
Those who somehow confuse our “neural network architectures” with the actual biology in our skulls are really misguided. Based on the evidence, one can at best opine that a TPU-based matrix add/multiply unit is approximating a very clunky and un-nuanced neural network. To assert that they are one and the same is... indefensible, when one goes beyond even the coarsest detail. A neural network as we implement it technologically is nothing but a weighted network of threshold units. That is all.
[1] To clarify: some kind of highly context-aware DSP with both long- and short-term storage facilities that allow it to react not only to the inputs that it is being fed, but also to inputs it was fed previously (and possibly outputs it produced previously as well).
To be fair, many lectures/books/articles/blog posts on ANNs start off by presenting a biological neuron, and saying that ANNs approximate this.
E.g. https://youtu.be/uXt8qF2Zzfo?t=385 (MIT 6.034 Artificial Intelligence, Fall 2010. 12a: Neural Nets)
I started off along much the same path (I bought and read Neural Networks: A Comprehensive Foundation (1994) by Simon Haykin) when I was in high school, read it all and thought I knew everything there was to know on the topic, and then started to read about neurobiology and was stunned to discover I was living a lie. Then I read The Computational Beauty of Nature (1998) by Gary William Flake and decided I was going to dedicate myself to understanding the complexities we brush under the rugs.
I think the Hacker News/Silicon Valley/Machine Learning/Developer crowd really need to look beyond the virtuosismo of their implementations. It’s a very introverted and self-referential mindset.
You might like this article: https://blog.piekniewski.info/2018/08/28/fun-numbers-about-t...
If the purpose of intelligence is to "convert data to information", then perhaps it is akin to doing some work that decreases information theoretic entropy?
For example, in regression, we express input data as parameters of the model and error, so in the heat engine analogy, the input is the incoming heat, the parameters are the work output, and the error is the waste heat. (The error has higher entropy because there is less to explain, and you typically throw it away.)
So perhaps, if somebody would figure out this analogy correctly, we would see that we always need some cycles to continue the operation, just like in heat engines. So then we could conclude that any agent might require some kind of cycles of information creation and destruction in order to have intelligence.
Thought experiment in order to clarify: how would aliens, staring at our planet for aeons through a telescope, reach the conclusion that there is or is not intelligent life here on Earth? Well... they could observe that on average, historically, about one every hundred million years or so there is a major asteroid impact. They would eventually observe that from some point onwards a larger and larger delay would set in from the last impact and the next, far longer than expected. If they zoomed in they’d notice that somehow asteroids on impact trajectories would be deflected as if by a physical force emanating from the planet.
This would be our descendants, using the amazing technology at their disposal, deflecting asteroids to ensure that the last major extinction level event will be the one that killed off the dinosaurs 65 million years ago (and counting). They’d be using their intelligence to keep their options open and avoid being wiped out. They’re exerting a physical force on their environment and clearing the sidereal neighbourhood of dangerous rocks and comets.
Conclusion: intelligence is about perception of data but for the sake of control of the environment. “Data into information”, honestly, is too close to being a truism to be of any utility.
This includes the ability to make accurate predictions about threats and opportunities within the environment, and the likely outcomes of planned actions.
The more intelligent an entity is, the more successfully it can generalise from past experience and anticipate and influence future events.
This is a completely different definition to plain old IQ test intelligence, which seems to be closer to raw mental agility. I suspect you can score high on mental agility - e.g. figure and number series manipulation - and still have poor applied intelligence. The latter needs the ability to synthesise a range of diverse stimuli, which is not the same as being able to manipulate abstracted symbols with no surrounding context.
EDIT: Revised preamble.