Thanks, karpathy!
I did submit your JS implementation to HN when I came across it: https://news.ycombinator.com/item?id=9108738
Monte Carlo Tree Search could be the missing link. In other words, use DQN to model the world and map actions to a value function. Then use playouts and backpropogation of action tree results to find tactics. Of course, it does not solve the big question: how to model "memories" and "inferences"? Indeed, very exciting times for AI/ML!