I'm not sure I'm right, but this logic works for me...
So, they generalized the method to find a predictable future, but in very unstable states. Later, Der and Ay generalized this towards maximization of predictive information - for that, it is not sufficient to predict to environment, but there also must be something nontrivial to predict. This gives rise to a number of highly interesting behaviours in many of their scenarios.
So, yes, your idea is a good one...
If there wasn't, why would say french soldiers be exchanging their wine and pate de foie gras MREs - especially after having satisfied their curiosity of what other MREs taste like, ie after the first trade?
There is actually a line of research where channel capacity of the actuation/perception channel is being used to model that (it's also named "empowerment", in analogy to the social science term, but it is measured in terms of information theoretical quantities). Here, one does not only model the entropy of the future, but in fact that amount of entropy about the future that the agent can systematically control. Or, more precisely, how well the agent is able to systematically control its future (which allows it maximal future entropy - but only entropy it can actually itself generate!)
See e.g. Klyubin et al. 2005, 2008, Capdepuy et al. 2007 and later, Jung et al. 2011, Salge et al. 2012. Indeed, maximizing the control over potential futures (empowerment) seems to work well for a range of scenarios, including the discovery of points of high centrality in mazes and graphs, as well as pole-balancing, acrobot balancing, bicycle balancing examples, in short, survival scenarios; also for discovering new object manipulation modes, driving sensor evolution models, self-organizing multiagent collective behaviour and a few other cases. It does not work well for systems with prespecified goals, systems with funnels/bottleneck transitions, in short, for cases where you have to relinquish potential futures to commit to a decision. This needs to be treated with a different principle.
The information-theoretic treatment takes into account action routes with controllable dynamics, it will avoid state space regions with more noise and thus less controllability. Importantly: it is the "potential" to do something, it does not enforce to actually carry out the option. Having the option is all that counts. Like in chess, the "threat" is more effective than the "carrying out".
Disclaimer: I am one of the authors of the empowerment work.