I wonder what this would look like here. Seems like a space where keeping the exposed metric and the optimization target apart would be quite difficult.
Also curious: by what reasoning path do models typically end up reward hacking?
Also curious: by what reasoning path do models typically end up reward hacking?