I tried to learn deep Q-learning recently using OpenAI gym. I looked at the "leader boards" and tried to learn from their code, but I wasn't getting results nearly as good as the "leaders".
I eventually checked out the leaders code and ran it myself, but removed their carefully selected random seeds, and found that the supposed leading solutions often failed to converge at all without their magically selected random seed.
I left the experience believing deep RL is still very unreliable.