OpenAI Retro Contest
contest.openai.com
contest.openai.com
You say that like you're talking about two separate things. DQN is just one type of DRL.
Well, it needs to be 2 months so Musk can transfer this learning to Autopilot... /s, I think
If you have the time to work on this contest, I'd recommend trying a method which is either not deep learning or not reinforcement learning or not both. There will be enough submissions from people who will try obvious ideas, and so you're better off trying something unique. Good luck and have fun!
ConvNets were invented by Yann Lecun, a professor of machine learning who had spent his life investigating these kinds of subjects.
What I'm saying is, machine learning is super interesting. You need to really be ready invest yourself though.
From the downloaded CSV of game states, all stages from the 3 games are valid training stages, and the custom stages used for testing are derived from all 3 games. Yikes.
I really hope those custom stages use the same art assets as the original games.
For an example of a system that does learn new rules using a single model check out this post from Vicarious[0].
[0] https://www.vicarious.com/2017/08/07/general-game-playing-wi...
Acknowledge its a hugely complicated subject.
Would be cool to see some kind of adversarial competition. You train to, say, beat a game level but you test to beat someone else’s submission. (Short on the specifics, I know.)
Anything else?
Like BDD for AI.
In this contest, you're not just submitting a trained model, you're submitted a docker environment capable of performing training, which they will run with their secret levels, with specified constraints. So you want to make a highly trained model whose capabilities can be transferred to secret levels.
So yep, the RGB image.
This is in contrast to current models which generally need to be trained up from scratch on each new game, even if the games are mechanically similar.
That provides a rather unique challenge, of making a good general baseline agent, but one that also is flexible for further training.
Check out https://contest.openai.com/details
I'm not sure how different those skills are from an AI development point of view though, why do you think they are?
> The reward your agent receives is proportional to its progress to the predefined horizontal offset within each level, positive for getting closer, negative for getting further away. If you reach the offset, the sum of your rewards will be 9000. In addition there is a time bonus that starts at 1000 and decreases learning to 0 at the end of the time limit, so beating the level as quickly as possible is rewarded.
So mostly time - but past a certain threshold just completion. And if you don't finish "portion completed".
The contest will run from April 5 to June 5 (2 months)
and *winners will receive some pretty cool trophies.*