ParentFull threadvisarga·My environments don't score actions, they score state, so no matter how the agent chooses to solve a task it gets scored correctly. It's called Potential Based Reward Shaping and its main benefit is that it does not introduce reward hacking.View on HN