My intuition on this:
Maximum likelihood training -> faithfully represent training data
Reinforcement learning -> seek out the most preferred answer you can
Maximum likelihood training -> faithfully represent training data
Reinforcement learning -> seek out the most preferred answer you can
No comments yet.