Entropy Maximization and intelligent behaviour
paulispace.com
paulispace.com
"Causal Entropic Forcing" is something like an AI's utility function, where the agent attempts to maximize future possibilities. Since this is meaningless (all possible futures are possible), what you actually want to do is make it as easy as possible to get to those futures - aka, their entropic adjacency, hence the name, causal entropic forcing.
However, CEF requires that the agent can actually predict possible future states of the system, which comes with some serious issues. In the original paper, this is covered by access to perfect simulators, but those aren't available in real-world situations.
This post discusses how to (possibly) use recurrent neural networks to make such predictions; how to do so effectively, and with consideration of the NN's confidence in it's predictions.
It's pretty cool!
Nonetheless it's an interesting post; I like the idea of coming up with abstract "goals" that can be applied to any environment (without having to construct a reward function) that yield complex behavior. Even if it doesn't do precisely what you want it's useful for exploration and perhaps a good stepping stone towards the desired behavior.
On a related note, I believe you can learn to predict the entropy of a Markov process using reinforcement learning, so it might be possible to extend it towards control.
I wrote up the basic idea: http://rl.ai/posts/generalized-returns-entropy.html which argues that if you had some sort of state transition model you could construct a reward function from it, and then learning the value of a state is also the "expected entropy" starting from that state. The state transition model can itself be learned, so no simulator is required. The reason I say "I believe" is that this is the product of original research, that is, procrastinating on my thesis. So there's a risk that I've made an error somewhere or missed prior work.
Basically we should be considering the current posted results versus Q learning with an equally accurate pre-trained forward simulator, which I don't think anyone has done.
AFAIK, the claims being gained are that they didn't have to supply the system with any goals, and it "figured out" the basic tests - included tool use and cooperation.
> with an equally accurate pre-trained forward simulator
AFAIK, that's exactly what the OP is discussing: how would this system perform when you replace the perfect simulator with an RNN trained to predict?
[edit]
Okay, I don't really understand what they're doing, so, my guess is that they have a component that predicts future states of the game, and they use something inspired by fractals to determine which future states to sample. Then, they're either using the score as the metric for the "value" of that future state (in which case it's not CEF), or they're ignoring the score and measuring something corresponding to entropy (or future-possibilies-remaining since it's pacman), and then they are using CEF.
It's like Data playing that weird chess game; the CEF aspect doesn't help you play any better, but it gives you a different win condition that turns out to "win" better than directly trying to win.
If, that is, my bad understanding is in any way accurate.
This, contrary to beating deep learning models to death, may hold the key to artificial intelligence.
I have my own little model which I propose for prediction, I call the Predictive Vision Model, more info here: http://blog.piekniewski.info
I'm looking for more people who see the potential of this.
It combines well with Jeremy England's theory that life is entropically inevitable: https://www.scientificamerican.com/article/a-new-physics-the...
..and I wonder sometimes if you could make a religion out of all this; morality and existence based on entropic math. Consider that the goal of CEF is maximized possibilities, smoke a bowl, and think about fractals and holograms.
One of the things I find really interesting about CEF is that it doesn't specifically help with understanding or predicting the world around you; it just gives a very effective way to determine what possible actions you should actually do. Given that (AFAIK) the human brain/mind is itself a combination of many systems, it seems to me to be very elegant that a CEF agent is also a combination of systems, each of which have limitations and issues.
Well, it gives one way to prescribe actions, but no way to prescribe actions we actually care about. Maximizing possible futures rightfully ought be a mere subgoal or consequence of the actual prescriptions we care about.
Personally I like free-energy theory best, but it really still needs some work to distinguish which "predictions" change to accommodate prediction-errors and which drive action. The original equations basically claim they both change at the same time to minimize the free-energy, but by then the generative models and recognition densities themselves approach tautology.
In a certain sense, you could view anything which behaves "teleologically", which self-organizes and moves itself preferentially into some states over others, as engaging in active inference on some generative density. The problem is to describe or prescribe what the generative and recognition densities actually are, lest the theory just be mere philosophy.
p(u|x)=p(x|u)p(u)/p(x)
+ max in log space: max ln p(x|u) + ln p(u)
+ use variational approximation, e.g. Kullback-Leibler: min ln KL(q(u)|p(x|u)) + ln p(u)
+ define free energy: F = ln p(u) - KL, so it can be maximized
Hence, rather than with KL where we minimize over a ratio with conditional p(x|u), we maximize over the joint 1/p(x,y). So we optimize for both likelihood and prior.
Sounds to me as a sloppy Bayesian approach. :-) Normally, the prior is intended to be used as full distribution. Not to get a max. probable value from.
In this approach we maximize for both prior and likelihood. Seems logical that we get all kind of possible trade-offs. Which should we choose? And why is free-energy so perfect?
To play devil's advocate though, it may be the case that our normal prescriptive preferences have the consequence of maximizing possible futures. E.g. 'attempt to survive' -> you will have future possibilities. It's possible that these are equivalent goals.
https://www.amazon.com/Our-Mathematical-Universe-Ultimate-Re...
One red flag for me is that they're simultaneously claiming that there's no training during the OpenAI Gym while also claiming that the optimization approach is relevant. In that case, what is being optimized? It seems like they might be optimizing over previous simulations - there's frequent reference to having access to a "simulator". In that case, that should effectively count as training, right? I was under the impression that the OpenAI Gym was supposed to benchmark untrained approaches so they could be compared by learning time. Hence the gradually increasing training curves in the other approaches.
Here, we do not need training of any kind either, just a monte-carlo simulation of the environment and an approximation of which path has the greatest path entropy. Bsaically given a state, you do
- Compute the path entropy for all states you can move to
- Move into the state with greatest path entropy
The tradeoff here is that all the work occurs in inference - every decision requires a complex simulation. In training based approaches the heavy lifting is done during training, and inference is easy
It's interesting that this merit function works in the absence of a real reward signal, but there's no fair comparison against systems using a reward signal due to this huge alteration to the problem that is providing a perfect simulation.
Having said that there are situations where this will fail completely, e.g. in maze solving, where the goal is not to play to keep playing but to play to reach the end.
I think I can explain what you're confused about, if I can understand what it is you're confused about better :)
The system has access to available actuators (AFAIK, the X or X+Y position of the agent), a perfect simulator (given this action, that position is the result), and an equation to measure the energy of the system (in a physics / entropy sense).
The first example is the inverted pendulum (segway). The agent can move along X, and it takes more energy to go from the down / fallen position to the upright position, than vice versa. Thus, the upright position has better entropy (I never get the +- right with entropy, so I don't know if that means more or less entropy).
Since the system knows the entropy present in all possible future states of the system (via the perfect simulator plus the entropy math), it can make a sort of "map", and plot a path from where it is to the global max.
In simpler terms, it's optimizing how much energy it takes to get from the current state to all possible future states of the system: in simpler terms, it's way easier (literally, takes less energy) to let the segway fall down than to stand it up in the first place.
Does that help?
CEF is an answer to "what works best" that's (theoretically?) applicable to all systems.
I guess I'm missing something, because this seems to negate the entire point... isn't the point that number of future options is a good measure of "more useful options"?
AIXI applies Solomonoff induction to a reinforcement learning (RL) setting: the sequence is split into three parts: "observations" (passive input, e.g. from a camera), "actions" (which are under the agent's control) and "rewards" (which are numbers). AIXI uses Solomonoff induction to calculate what the total future rewards will be, if the sequence so far were followed by action A; or by action B; etc. and then performs whichever of those actions gave the largest predicted reward. This does tell us which action to take (at least, computable approximations do), but it relies on there being a source of reward; all sorts of "AI safety" research (e.g. intelligence.org ) is based around what such a reward should look like, and ways that an AI might achieve high reward whilst subverting our intentions.
This 'causal entropic force' is a sort of implicit reward: the system is rewarded when it is able to efficiently reach other states; so it ends up 'putting itself in a good position', whatever that might mean in a particular situation.
It hand-waves away a few key points: it needs a good predictor (e.g. Solomonoff induction, or something computable), and it also seems to need a world model which tells it what the "states" are. Solomonoff and AIXI don't need to be given a model: they build their own implicitly. They do need their input to be hooked up, e.g. to take pixels from a camera or whatever, but that's a known property of the implementation (e.g. the hardware available on a robot), whereas there's usually a bunch of ways we could model the world, with no "obvious" right answer, and that can directly affect how the system behaves.