Q learners have always been too brute forcey imo. Every action from any known state gets a value. That will explode on any meaningful task. Its better to use it in combination with some other model that reduces how many states are seen by the q learner.
But it is a good model to understand some principles and motivations behind the general ideas in AI and machine learning.