I will address the supervised vs unsupervised issue in my next post. Here, I believe the analogy would be that when a field is applied to a spin glass, it does not exhibit a glass transition to a non-self-averaging (highly non-convex) ground state.
As to supervised vs reinforcement learning, its not that different. See how Vowpal Wabbit incoporates both the 2 ideas in how the SGD update is formulated.