Why I’m Remaking OpenAI Universe
blog.aqnichol.com
blog.aqnichol.com
Here's an example of a guy who made a general game playing algorithm that brute forces it's way through any NES game: https://www.youtube.com/watch?v=xOCurBYI_gY This isn't necessarily interesting from an AI perspective - the playing algorithm is just brute force. But it shows what can be done with the platform, easily reloading to previous states and exploring counterfactual futures (which is exactly the sort of thing RL algorithms do.) He also has a cool algorithm for finding the objective function of an arbitrary game, by watching a human play, and seeing what memory addresses increment. Which is a lot more easy to use than writing OCR code to read the score and game over states from the screen.
Great project. We've found that the VNC Universe environments are hard for today's RL algorithms primarily due to the their async nature. We're currently working on a new set of Universe environments without VNC; I'm very happy to see others inspired by the core ideas of Universe as well.
I took a lot of inspiration from Universe and am grateful for OpenAI's work on RL in general :). I probably wouldn't have started on this project if a company like OpenAI hadn't already decided it was a worthy goal.
The way I see it, having hooks into the engines themselves helps with what the article talks about - not needing to go through VNCs or other _glue_ to get realtime data. It could potentially send the framebuffers themselves directly from the game/simulation and tie in the actions back to the game/simulation. And using framebuffers is just one direction, we could instead stream the co-ords/the current payoff/etc.
Also, having such plugins would help with the adoption in both directions - games now have an always updating/learning AI (might need a network connection + cloud backend), and researchers can have training/testing environments.
However I think this approach is bad. Machine vision is a separate problem from reinforcement learning. You shouldn't need to be able to do both well. Machine vision consumes a ton of processing power and researcher time in figuring out the hyperparameters. And all it's doing is figuring out information that's already in memory like the location of various objects and the score. It really limits what can be done. E.g. the famous atari playing AIs by deepmind were limited to no memory and only knowing the last few frames, because backpropagating through thousands of frames was too expensive.
Because of the way NNs work, it's trivial to separate out the machine vision into a separate module. So if you have a good RNN reinforcement learning system, you can easily add a machine vision learning system to it later if you need.
That's a bit premature for a project that was just released less than 7 months ago, isn't it?
https://blog.openai.com/universe/
Edit: that said the project seems to have some interesting and needed improvements (esp time adjustment). Glad to see dialog between muniverse and openai here.
Despite the flaws, the nice thing with VNC is its universality to support any apps on a computer. Using HTML5 in a browser limits the scope of things we could encapsulate as environments, and makes it less "universe".
However, there is a difference between the universality of the tech stack and the exposed interface. In my opinion, the future universe would be rich clusters of RL environments with unified API, each of which implemented using different underlying technology to meet the desired synchronicity and frame performance.
HTML5 could deliver one of such clusters.
https://www.quora.com/What-is-the-difference-between-OpenAIs...
Source: Am an AI research scientist.
What we do know is that current techniques won't get us close to AGI, so something new is needed (or perhaps like backprop, something old will work once we have enough compute power). Personally I'm bullish on AGI because I have strikingly low faith in the ability of evolution to operate very effectively as a tool for algorithm discovery, so I suspect that once we've hit the compute threshold we'll find that many different algorithms can do the trick, and 40 years is probably not out of the question for us to hit that point (or 10, or 100), depending who you talk to about what the compute threshold might be.
I'd caution against putting too much weight in what experts say, though, since with a tiny few set of exceptions anyone working on "AI" today is actually just working on narrow AI, which is, as someone put it, just glorified linear regression. Those tools will almost certainly be part of the solution, but only in the sense that the classical theory of Diophantine equations was part of Weil's proof of Fermat's Last Theorem - they are not the core of the theoretical approach.
Evolution is a slow algorithm, but it had access to an absurd amount of compute (all neuronal organic matter on Earth) and environment simulation (all of physical reality on Earth) when discovering us; so the discovery of the algorithms/architectures/principles in our heads shouldn't be viewed as trivial.
Personally, I'm bearish about AGI because I believe we will eventually realize that the brain is a glorified linear regression too, with a custom wiring to help learn language and vision.
With backprop we didn't just need bigger machines, we needed better algorithms, palliatives for the exploding-gradient problem that made values exceed our numerical representations, and then hardware specifically designed for doing the matrix-ops involved.
If I saw something capable of speeding up probabilistic program inference the way GPUs sped up backprop, I'd start saying we should expect to see powerful AI applications quite soon.
Probabilistic programming isn't going to help general AI much. Things like dropout seem to work well enough, and for the most part AI is severely underfitting rather than overfitting. Our models are far to simple and small to really learn language and do complicated reasoning. Making them bayesian doesn't fix that.
Excuse me while I laugh.[1,2,3,4]
>Things like dropout seem to work well enough, and for the most part AI is severely underfitting rather than overfitting.
For the most part, neural networks can't reason at all. They just induce deterministic functions over high-dimensional Euclidean spaces.
>Our models are far to simple and small to really learn language and do complicated reasoning.
They're also not compositional (new concepts as functions of old concepts), productive (able to draw an unbounded number of inferences from each representation), or unbounded in size of representation (unboundedly many concepts). Neural networks don't even represent causal structure, let alone model how an intervention will affect outcomes!
It is, however, really nice to hear an AI booster admit just how incredibly limited connectionist models actually are.
>Making them bayesian doesn't fix that.
No, changing to a causal, compositional representation that allows for productive and nonparametric (unboundedly large) learning does that. The Bayesian part just makes it extra nice by letting us "put information in" anywhere in the model (at any variable) by conditioning.
[1] -- http://forestdb.org/models/learning-physics.html [2] -- http://forestdb.org/models/word-learning.html [3] -- http://forestdb.org/models/arithmetic.html [4] -- http://forestdb.org/models/politeness.html
The idea that OpenAI could talk him down is pretty impressive, and if true I would significantly positively update my impression of OpenAI. (I thought OpenAI was funded by people on this hype train.)
My advisor shared the following wisdom with me: "When the experts in your field say that saying can be done, they are probably right. When the experts in your field say that something cannot be done, they are not necessarily right."
Generally yes, but they may be significantly off on the timeframe. One famous example is that once alpha-beta search was invented (in the late 1950s), Herb Simon predicted that "within ten years a digital computer will be the world's chess champion". That did eventually happen, using techniques not even all that different from alpha-beta search, but it took 40 years rather than 10. Many of the 1980s neural nets claims turned out to be eventually vindicated too, but it took 30 years, which was quiet a bit longer than the optimistic portion of 1980s "connectionists" expected.
That's the type of skepticism I usually have with claims today too. When people say "there will be fully autonomous self-driving cars on the road by 2020", I don't doubt it'll happen, but whether it'll happen in less than 3 years I have more doubts about. You could argue AI researchers have gotten better at accurately predicting the timeframes of advances than they were in the early days of AI, but I'm not sure there is solid evidence of that (would be interesting if someone has studied it).
I'm guessing you work in image recognition, or mostly hear from people who work in image recognition.
There is more to AI, and not all of it is instantly improved by a convolutional neural net.
What are the most valuable, unsolved problems in the field?
Edit: And seems like you are wrong anyway, see top comment.
https://github.com/namuol/muniverse
If I had more time I'd submit a PR to integrate it...