The industry has been focusing hard on 'System Two' approaches to augment our largely 'System One'-style models, where optimal decision policies make Q-learning a natural approach. The recent unlock here might be related to using the neural nets' general intelligence/flexibility to better broadly estimate their own states/rewards/actions (like humans). [EDIT: To be clear, the Q-learning stuff is my own speculation, whereas the 'System Two' stuff is well known.]
Serendipitously, Karpathy broadly discussed these two issues yesterday! Toward the lecture's end: https://www.youtube.com/watch?v=zjkBMFhNj_g