ConvNetJS Deep Q Learning Demo (2013)
cs.stanford.edu
cs.stanford.edu
P.S. username checks out.
[1] https://github.com/nilq/genesis [2] https://github.com/nilq/plane
I've always had a soft spot for q-learning ever since :)
edit:
In grad school, I implemented q-learning to learn strategies for playing games (like tic-tac-toe, towers of hanoi, etc).
But there are certain applications where it’s the only good way of doing it (for example in games, where you don’t have access to a gradient over the space of how good a certain move is given the current game state, and it’s relatively quick and and efficient to simulate the evolution of the game).
NEAT: https://www.cs.ucf.edu/~kstanley/neat.html
HyperNEAT: http://eplex.cs.ucf.edu/hyperNEATpage/HyperNEAT.html
There are explanations, links to the research papers, and links to implementations.
Here is a working example that runs in your browser using Javascript: https://liquidcarrot.io/example.flappy-bird/
Edit: i think what i'm thinking of are two "global" numbers. A "closer to the goal" number (higher when closer), and a countdown timer, and then it has to maximize those within the above setting
sometimes a big negstive reward can scare it away from ever pressing the scary button ever again
You do need to know the expected outputs - positive rewards and negative.