That got solved using Deep-Q learning. I think David Silver did loads of work on it?
Basically, instead of computing every state value like in normal Q learning. You use a Neural Network to estimate the best value.
Basically, instead of computing every state value like in normal Q learning. You use a Neural Network to estimate the best value.