> if one greedily maximize the sum of entropy and value with one-step look-ahead, then one obtains π˜ from π
> if one greedily maximize the sum of entropy and value with one-step look-ahead, then one obtains π˜ from π
If you're just wondering about the meaning of (18), it claims that for a given state, the entropy of the original policy plus its expected Q-value is less than the entropy + expected Q-value of the π-tilde policy. Where they are using Q_{soft}^{\pi} to denote the state-action values (with entropy term) for an arbitrary policy. In (17) they introduce a new policy, \tilde{\pi} defined through the Q_{soft} values for some policy π, with the probability of taking action 'a' in state 's' proportional to the exponentiated Q(s,a).
It works by iterating two steps:
- Keep the policy $\pi$ fixed, and determine the $Q$ values for that policy
- Keep the $Q$ value estimates fixed, and create a policy $\pi'$ such that
-- $\pi'(s) = \arg\max_a Q(s,a)$ for the ordinary greedy algorithm,
-- $\pi'(s,a) \propto \exp Q(s, a)$ if using greedy soft-max.
[1] - Lecture 3 of http://www0.cs.ucl.ac.uk/staff/d.silver/web/Teaching.html