Cool and impressive. I'm curious if this training method will become more common.
Cool and impressive. I'm curious if this training method will become more common.
* Announcement: https://twitter.com/billyuchenlin/status/1749975138307825933
* Model Card: https://huggingface.co/snorkelai/Snorkel-Mistral-PairRM-DPO
* Response Re-Ranker: https://huggingface.co/llm-blender/PairRM
"We would also like to acknowledge contemporary work published independently on arXiv on 2024-01-18 by Meta & NYU (Yuan, et al) in a paper called Self-Rewarding Language Models, which proposes a similar general approach for creating alignment pairs from a larger set of candidate responses, but using the LLM as the reward model. While this may work for general-purpose models, our experience has shown that task-specific reward models guided by SMEs are necessary for most enterprise applications of LLMs for specific use cases, which is why we focus on the use of external reward models."
The naming of these models is getting ridiculous...
I'm an occasional visitor to huggingface, so I'm actually superficially familiar with the taxonomy. I just felt like, even if I tried to satirize it, I wouldn't be able to come up with a crazier name. And that's not even the end of the Cambrian explosion of LLMs.