596 karma · joined October 27, 2016
We basically have to distrust our memories all the time: https://www.youtube.com/watch?v=dHd_TNHsyVA
> 18.2
Because the negative phase involves drawing samples from the model’s distri- bution, we can think of it as finding points that the model believes in strongly. Because the negative phase acts to reduce the probability of those points, they are generally considered to represent the model’s incorrect beliefs about the world. They are frequently referred to in the literature as “hallucinations” or “fantasy particles.” In fact, the negative phase has been proposed as a possible explanation for dreaming in humans and other animals (Crick and Mitchison, 1983), the idea being that the brain maintains a probabilistic model of the world and follows the gradient of log p ̃ while experiencing real events while awake and follows the negative gradient of log p ̃ to minimize log Z while sleeping and experiencing events sampled from the current model. This view explains much of the language used to describe algorithms with a positive and negative phase, but it has not been proven to be correct with neuroscientific experiments. In machine learning models, it is usually necessary to use the positive and negative phase simultaneously, rather than in separate time periods of wakefulness and REM sleep. As we will see in Sec. 19.5, other machine learning algorithms draw samples from the model distribution for other purposes and such algorithms could also provide an account for the function of dream sleep.
> 19.5.1 Wake-Sleep
One of the main difficulties with training a model to infer h from v is that we do not have a supervised training set with which to train the model. Given a v,we do not know the appropriate h. The mapping from v to h depends on the choice of model family, and evolves throughout the learning process as θ changes. The wake-sleep algorithm (Hinton et al., 1995b; Frey et al., 1996) resolves this problem by drawing samples of both h and v from the model distribution. For example, in a directed model, this can be done cheaply by performing ancestral sampling beginning at h and ending at v. The inference network can then be trained to perform the reverse mapping: predicting which h caused the present v. The main drawback to this approach is that we will only be able to train the inference network on values of v that have high probability under the model. Early in learning, the model distribution will not resemble the data distribution, so the inference network will not have an opportunity to learn on samples that resemble data.
Another possible explanation for biological dreaming is that it is providing samples from p(h,v) which can be used to train an inference network to predict h given v. In some senses, this explanation is more satisfying than the partition function explanation. Monte Carlo algorithms generally do not perform well if they are run using only the positive phase of the gradient for several steps then with only the negative phase of the gradient for several steps. Human beings and animals are usually awake for several consecutive hours then asleep for several consecutive hours. It is not readily apparent how this schedule could support Monte Carlo training of an undirected model. Learning algorithms based on maximizing L can be run with prolonged periods of improving q and prolonged periods of improving θ, however. If the role of biological dreaming is to train networks for predicting q, then this explains how animals are able to remain awake for several hours (the longer they are awake, the greater the gap between L and log p(v), but L will remain a lower bound) and to remain asleep for several hours (the generative model itself is not modified during sleep) without damaging their internal models. Of course, these ideas are purely speculative, and there is no hard evidence to suggest that dreaming accomplishes either of these goals. Dreaming may also serve reinforcement learning rather than probabilistic modeling, by sampling synthetic experiences from the animal’s transition model, on which to train the animal’s policy. Or sleep may serve some other purpose not yet anticipated by the machine learning community.
The idea is basically that junk DNA contains a memory in shape of a distributed representation of the past of the organism and its environment, akin to how neural networks encode information. It basically provides a basis for fast adaptability by introducing noise into the gene expression and morphogenesis process so as to have more versatility and robustness to explore alternatives (very similar to dropout in neural networks).
But it also has more direct causal motivation: If you don't fight for basic human rights, taxes, plurality etc., you will likely erode the stability of your own life unless you are completely emotionally, economically and technologically independent (which I doubt you are). There is always the possibility of an unexpected economic decline which a rational agent should take into account for maximizing their reward signals.
'Learning to reinforcement learn'
https://hn.algolia.com/?query=pdf&sort=byPopularity&prefix=f...
There are many things that are too small, too large, too fast or too slow etc. for us to notice. For example, you miss out all the air eddies you create as you move through the atmosphere. Lots of things are also simply out of sight and our minds construct perhaps 50% of what we perceive by pattern completion. I think what we perceive is mostly right, but in many ways it only corresponds superficially to the computations that are actually going on.
I don't find this distinction very useful either though. Yes, there is an illusion of confidence in the accuracy of what we experience, and a lot of it is tainted by values and false memories, but still, it seems plausible that our experiences correspond very directly to things happening in the universe.
I also don't agree with the claim that mathematics being real in a Platonic sense is an illusion. It all comes down to how you define 'real'. I find a sensible definition is that it is a thought that corresponds to how the world works, to things in the world (future, present or past), but also to how the world hypothetically, but very plausibly works. In the same way as you can plausibly assume that any imaginable thought corresponds to something in a remote region of an infinite universe (or multiverse), you can also assume that our mathematical theorems correspond to some computation somewhere. It makes sense to extend the definition of real by immediacy: The fictitious novel is also real (in some multiverse), but it does not have as much immediate real-ness as mathematics. For mathematics you can find all kinds of correspondences in nature (e.g. 1+1=2), but in the novel you can mostly only find things you've previously taken from nature to write the novel, but if your predictions are very real, then it might at some point be impossible to tell whether it is true or not (just like a computer game that almost looks real).
Reality simply provides a field that vibrates at different wavelengths and it is intelligence and evolution that come up with computationally convenient category boundaries along that spectrum. We are certainly unaware of many things that are part of the universe without the help of tools.
I agree that we can be certain that when e.g. we see a large rock falling, that our mental representations relate with 99.999% certainty to the event actually occurring. There are however many things which are more subtle, and I think this is what he was referring to.
If you are referring to the question from the IRC chat, I think it was not about silencing but rather about genuine interest in further reading material since he didn't yet publish on anything from his three CCC talks. Some parts make quite sophisticated points so it would be really helpful if it was elaborated in written form.
It seems the technique presented in the paper was simultaneously discovered by Sixt et al: https://arxiv.org/abs/1611.01331 (Nov. 4th vs Nov. 15th)
The Size and Shape of “Idea Space” (2011) [pdf]