ParentFull threadAncapistani·Reinforcement learning, I'd assume.Presumably RLHF (Reinforcement Learning from Human Feedback).View on HN