That blog post also mentions the noisy TV problem.
I am not skilled enough to describe how these approaches differ.
That blog post also mentions the noisy TV problem.
I am not skilled enough to describe how these approaches differ.
> The environment also contains a TV for which the agent has the remote control. There is a limited number of channels (each with a distinct show)... even if the order of shows appearing on the screen is random and unpredictable, all those shows are already in memory!
So in google's "Curiosity and Procrastination in Reinforcement Learning" they could not handle a TV with pure noise (snow) as it could not remember all that noise.
But how do we decide whether the agent is seeing the same thing as an existing memory? Checking for an exact match could be meaningless: in a realistic environment, the agent rarely sees exactly the same thing twice. For example, even if the agent returned to exactly the same room, it would still see this room under a different angle compared to its memories.
Instead of checking for an exact match in memory, we use a deep neural network that is trained to measure how similar two experiences are.
If there is an unlimited number of shows and the agent walks right up to it, I think its still trapped.