Alpaca RLHF-ed to beat ChatGPT
crfm.stanford.edu
crfm.stanford.edu
4 isn’t just marginally better at most tasks I use it for, it’s operating at an entirely different level to the point where I have little (no?) day-to-day use of 3.5 at this point.
I guess this doesn't apply if you use it via the api.
When I hit the limit, I work on the problem myself and wait until 4 resets instead of relying on 3.5. 4 is so much better that I don’t trust 3.5 with my work anymore.
Your coding speed is unlikely to be that fast, requiring 25 code segments in 3 hours. GPT-4 outputs something, you need time to double check, test, additional googling etc. Its still a massive speed boost.
Using it recreationally (Especially chatting) will result in a lot more requests.
I haven't used it a whole lot.
It might seem a small thing, having to space out prompts 25 ever 3 hours when you might not have used more than 100-200 in a day anyway, but the net result is liberating. I experiment, explore the limits, and get whimsical with it to a much greater extent than when I can to consciously think about each prompt as a rationed resource.
It’s going to be very funny if being turned into a next generation Clippy is what makes them lose out to their competitors
Reality:
> With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate.
> Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT.
…for the metric we invented, which measures… the difference between a simulated and human evaluated result.
Or something.
Does anyone have a good idea of what this metric actually means and if it is actually relevant to anything useful?
Beating a weaker player more often is not evidence of being able to beat a stronger player on average though
It improved the simulated win rate vs human win rate?
…but chatgpt had a higher win rate overall? (And gpt4 was much higher)
What is the significance of the difference between simulated and human win rates?
You should try asking what you don’t know in a non judgemental manner
The paper says:
> We find that PPO sim trained in AlpacaFarm only achieves a win-rate of 43%, while PPOGPT-4 sim trained on GPT-4 data achieves a win-rate of 50%. To contextualize these results, the initial SFT model has a win-rate of 44%, PPOhuman has a win-rate of 55%, and the best non-PPO human method has a win-rate of 51% (Best-of-16). Thus, training in simulation can provide good models directly for deployment, though this approach suffers a 5% performance gap relative to collecting real human annotations.
...
> However, we also observe that no single LLM-based annotator captures the heterogeneity of human annotation, and substantial amounts of noise had to be injected in the simulated preference for rankings of methods trained in AlpacaFarm to match those trained with real human feedback.
...and, in summary:
> We showed that AlpacaFarm substantially lowers the cost and iteration time of research on and development of methods for learning with pairwise feedback. AlpacaFarm provides a blueprint for constructing other useful simulators for AI research that requires human supervision, and we view it as an exciting opportunity to expand this simulation approach to support data from other domains as well as methods that learn from alternative forms of human feedback.
Ok.
...but that's no what the blog post said. The blog post said:
> Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT.
The closest the paper got to saying that was:
> The other mismatch is ChatGPT against PPO, where human annotators preferred PPO (55.1% vs 52.9%) unlike the simulator (46.8% vs 61.4%).
That's interesting.
> In both cases, these are not major mistakes, as we do not expect SFT52k to be much worse than SFT10k or for a 7B LLaMA model to substantially outperform ChatGPT.
?? Mistakes?
So.. I mean, yes. I'm judging. When you write a blog saying "outperforms ChatGPT" and then, the paper doesn't say that... well.
It's a bit shit isn't it?
For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!
But it also seems like a strange, incestuous, closed system approach.
Like, unless you are introducing something new into the system, you just have the system churning against itself, probably until it reaches an equilibrium (or else becomes incoherent).
Took me a bit of searching too.
They should try to compare answers with similar length.
RLHF is supervised learning on top of unsupervised learning. Is supervised learning at some point of the process a requirement for all reasonable ML models?