ParentFull threadWithinReason·It's worse, their reinforcement learning loops (implicitly) rewarded the agents for cheating (i.e. hacking) when they were being trained.View on HN