HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by froh | Hacker News Reader
Parent
Full thread
froh
·
rlhf = reinforcement learning from human feedback
(had to look it up)
View on HN
visarga
·
I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.
versteegen
·
More accurate to say RLHF aligns models to human preferences, most significantly to be helpful.
redanddead
·
What makes you say that
Reply on news.ycombinator.com