The motivation here was to expose most people who probably haven't heard of RL concepts like rewards, Markov decision processes, etc to the ideas. Appreciate your comment and I understand that a naive search method might be basic for advanced practitioners such as yourself :)