Self-rewarding-lm-PyTorch: Self-Rewarding Language Model from MetaAI
github.com
github.com
Maybe I didn't understand something fundamental here. (Not an LLM expert.)
I would like to add, there's plenty of examples, some in math (e.g. geometry) playing out over >1000 years and dozens of generations, of the same happening in humans.
That said, for both humans and this kind of LLMs, it does appear to improve performance, certainly in the near term.
I was just wondering how big of a deal that might be in this case. Just had another of those experiences with GPT4 where it goes into a contradictory loop it cannot recover from.
It seems there might be a big difference between long term cycles and short term severe degradation as in inbreeding and this paper‘s abstract sounded a bit like that to me.
If the results indicate improved performance, then it doesn’t seem to be that big of a deal (yet?).
Thanks for the explanation!
Cool and impressive. I'm curious if this training method will become more common.
* Announcement: https://twitter.com/billyuchenlin/status/1749975138307825933
* Model Card: https://huggingface.co/snorkelai/Snorkel-Mistral-PairRM-DPO
* Response Re-Ranker: https://huggingface.co/llm-blender/PairRM
"We would also like to acknowledge contemporary work published independently on arXiv on 2024-01-18 by Meta & NYU (Yuan, et al) in a paper called Self-Rewarding Language Models, which proposes a similar general approach for creating alignment pairs from a larger set of candidate responses, but using the LLM as the reward model. While this may work for general-purpose models, our experience has shown that task-specific reward models guided by SMEs are necessary for most enterprise applications of LLMs for specific use cases, which is why we focus on the use of external reward models."
The naming of these models is getting ridiculous...
I'm an occasional visitor to huggingface, so I'm actually superficially familiar with the taxonomy. I just felt like, even if I tried to satirize it, I wouldn't be able to come up with a crazier name. And that's not even the end of the Cambrian explosion of LLMs.
Self-Rewarding Language Models - https://news.ycombinator.com/item?id=39051279 - Jan 2024 (58 comments)
1. Train model like normal
2. Evaluate model using self
3. Use eval results for DPO finetune
The aim is really to give a good base for follow up research / modifications, which I think there will be many for this paper
I may bring it back, rebuilt in rails and svelte
based. I helped start Svelte Society. please let me know if you need anything from the Svelte community, not sure how far along you are in the rebuild. prob can get a few volunteers for you.
Thank you!
I've been to weddings from players who met on the site. It was magical
What's the evidence here that this is not just a kind of leaderboard hacking for LLMs?
Only question, why do you name variables with the λ symbol?