But "getting" the idea of Q learning for a small state space is fundamental and surprisingly approachable.
I wrote this extensive tutorial for teaching deep reinforcement learning, with a focus on getting intuition from code. you will find RL theory is heavy on math despite needing math for very little other than abstractly representing some machine goal and intuition, of which code serves a native programmer already very well.
i spent years failing to learn machine learning and RL until i just started reading source code. books of integrals i never ended up needing.
dont be turned away by the joking nature of my tutorials. there is a real depth in there
I'd suggest getting a good book or other teaching resource and solve a few Gymnasium[0] environments. Unlike supervised machine learning, you don't need someone else's data, you generate your own data.