Self-Rewarding Language Models
arxiv.org
arxiv.org
Many researchers believe that using synthetic data and automatic classification/scoring, i.e. using LLMs to improve LLMs, is likely to be one of the most successful lines of research. We've been seeing a lot of success from this in recent months, including OpenAI's DALL-e 3, which used (IIRC) 99% synthetic LLM data for captions.
I'm curious about both this and the emphasis on "high quality" data (e.g. Microsoft's Phi models) ...
1) What is/are the goal(s) of using synthetic data, just a source of more data, or "high quality" data ?
2) What is the definition/measure of "high quality" data - is this about consistency, or coverage, or what ?
Starting with llama 70b makes sense for Meta, but I can't help but wonder what the results would look like if applied to Mixtral. If it replicates and isn't overfitting could we see a performant and open source GPT4 competitor?
I also expect that evaluation ability doesn't grow linearly with prediction ability, so I doubt that a model will be able to fully optimise it's evaluation potential on it's own.
Could be wrong though, will be interesting to see. If I'm right on the former and wrong on the latter we could see maybe models self evaluating to an "optimal" state, for whatever it thinks optimal is via self evaluation is anyway.
Or this could be a somewhat useful case of overfitting that just refines a few bits.
It's hard to say how well this would work on a MoE model, but best case scenario is something decently better than GPT4 that can run on 2x3090/4090 or 1x48gb 5090.
Here's hoping AMD forces them to keep pushing the boundary.
But, I liked the idea that the reward model is not static, and if the user is provided with multiple options, then the extra score might help break the tie.
Hard to say — these things can be difficult to predict. I can see this working but there'll probably be some ratio of training data - self-play that we have a hard time getting past because it's a difficult-to-control form of extrapolation.
Also, while humans are better at judging (look at that fake thing!) than generating (drawing a realistic photo), I think you may find that — despite our abilities — detection is actually quite a bit harder. As an example:
"John Doe is dead."
This was very easy for me to create, but it's quite difficult for you to judge whether or not it is true due to a variety of factors (which John Doe, am I being honest, when was the last time you saw John Doe, do you know anyone who knows John doe and could check, perhaps John Doe had a twin who died, etc.)
The authors tried different LLM-as-a-judge promptings to generate a reward score for each answer. A very particular additive 5-point rewarding prompting is found to be the most effective one. The two-step inferencing pipeline (answering questions+evaluating answers) also generates an extra dataset in the form of <question, winning-answering, losing-answering>.
This AI-generated dataset (called preference pairs) is fed back to the model in a training pipeline (using Direct Preference Optimization).
The inferencing and training pipelines are connected to have a closed-loop, iterative process. Each iteration generates better AI feedback training data and subsequently better model. The evaluation shows very promising results, outperforming Claude 2, Genimini Pro, and GPT-4 in selected benchmarks.
The paper has some zoom for improvements. 1) Figure 1 is not very accurate to reflect the entire workflow. A fixed model is used to generate prompts for example. But it is not shown in the figure. The preference pairs should be a matrix instead of a vector in the diagram. Also, the bootstrapping workflow (using seed instruction following and evaluation datasets) should be reflected.
2) The authors did not explain why a fixed model is used to generate prompts, instead of using the self-rewarding model directly.
3) The authors tried to use another form of AI feedback data (question, best-answer), coupled with supervised fine-tuning. However, it did not result in any performance improvement for the model. It is better to explore why or at least propose it as future work.
4) Fundamentally, the paper does not directly compare (or comment on) self-rewarding vs. independent rewarding. The iterative process can still apply to an independent rewarding model.
edit: typos
https://youtu.be/bZQun8Y4L2A?si=RgD7NWfwDdh0bklK&t=1630
But if this approach holds up, it suggests that the most valuable part of the AlphaGo project that applies to LLM development is in fact reinforcement learning through self-play. Why not create a leaderboard of reward-generating language models that constantly "play" against each other, where model selection frequency is based on ELO, and update on the results? What if the Open LLM leaderboard (https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...) evolved into a constantly improving pool of such models? This also alleviates data scaling issues by providing a diverse and continuously changing distribution of new training inputs.
Harnessing the power of human-annotated data through
Supervised Fine-Tuning (SFT) is pivotal for advancing
Large Language Models (LLMs). In this paper, we
delve into the prospect of growing a strong LLM out of a
weak one without the need for acquiring additional
human annotated data. We propose a new fine-tuning
method called Self-Play fIne-tuNing (SPIN),
which starts from a supervised fine-tuned model. At the
heart of SPIN lies a self-play mechanism,
where the LLM refines its capability by playing against
instances of itself. More specifically, the
LLM generates its own training data from its previous
iterations, refining its policy by discerning
these self-generated responses from those obtained from
human-annotated data. Our method
progressively elevates the LLM from a nascent model to a
formidable one, unlocking the full
potential of human-annotated demonstration data for SFT.
Theoretically, we prove that the global optimum to the
training objective function of our method is achieved
only when the LLM policy aligns with the target data
distribution. Empirically, we evaluate our method on
several benchmark datasets including the HuggingFace
Open LLM Leaderboard, MT-Bench, and datasets from Big-
Bench.
Our results show that SPIN can significantly improve the
LLM’s performance across a variety of benchmarks and
even outperform models trained through direct
preference optimization (DPO) supplemented with extra
GPT-4 preference data. This sheds light on the promise
of self-play, enabling the achievement of human-level
performance in LLMs without the need for expert
opponents.(c.f. instances of itself, self-play)
There are no such unambiguous rules to ground language models. Iterative "improvement" could easily degenerate into nonsense without grounding in the real world. That's why some people think that true AGI will need to be grounded by the laws of physics, via experience interacting with the real physical world using robot bodies.
Of course, grounding from robots is also scarce unless you build a whole lot of robots. Ultimately we need systems that generalize rules from a small amount of grounding data. I guess that describes world models, which large language models currently lack (explicitly, at least).
It only goes so far, but considering how far LLM's got already, it seems promising.
[0] https://en.wikipedia.org/wiki/AlphaGo_Zero
What I am saying is there is no analogous source of manually programmed rules to ground language models during training.
It makes perfect sense that predicting the next token can be improved by going back every so often to reevaluate if the sequence of tokens makes any sense as a whole.
If we're conjecturing anyway I reckon the next major step won't come from changes that merely improve training or predicting, but one that fundamentally makes the model capable of learning. Removing the distinction between 'conversing' and 'learning'. This is similar to what you call 'interacting', but I reckon that mere interaction won't be enough if the distinction between letting the model predict and training the model doesn't disappear.
What an interesting turn of history.
(I'll acknowledge that LLMs might eat a layer off the top of that where people are seeking knowledge .... but it's never even going to come close to replacing the traditional core business)
1.Competing purely on result quality with Google is like competing with Coke based purely on having a subjectively better tasting formula.
If it really was better at search than Google (and scaled, was economical, etc.) it would have to be VASTLY better to get people to switch. Google has the brand, the defaults, the ecosystems, decades of habit forming, etc. It would not be enough to just be "better". Traditional wisdom would say it would have to be 10x better (the idea of quantifying a subjective improvement like this is absurd ofc.)
2. AGI aspires to be human-like. I don't see a human like intelligence as anywhere near Google's level for search results. A traditional vision of an AGI would be like asking your clever mate. Useful, but not for search.
3. Even if something came along which was truly better enough, Google themselves would have to not have access to it/something comparable. You would need to somehow have a lasting, huge & exclusive advantage.
4. Even if there was something that was a clear 10x improvement on Google, which had novel tech which Google could not themselves replicate, then Google would simply acquire them. Almost everything Google has done since the original site has been acquiring a tech, applying operational expertise, network effects, etc. & acting as essentially a distributor.
Their business model is built on serving links, which as time goes on will become more and more obsolete. The links are where the ad money comes from.