Can you name two real world problems with practical applications that fit all the criteria? I.e. where (1) an agent needs to act on its world but it would be okay for it to spend many tries exploring and failing; (2) it will be able to quickly and cheaply (i.e. without 24/7 human supervision) get a reward/punishment for its choice, and (3) the set of world states and possible actions is sufficiently small so that Q-learning is tractable?
IMHO Q-learning isn't talked about because it really is not a good fit for the kind of problems people actually want to solve.
What behavior does Hiome 'learn' with Q-learning? From your site it's not obvious what actions on the world would be implied where you can actually get some feedback/reward/punish depending on whether these actions were desired by the smart home inhabitants; and the behavior that your page does show - occupancy sensing - is essentially a classification problem.