ParentFull threadmeow_mix·Reinforcement learning w/ human feedback. What u guys are describing is the alignment problemView on HN