Orca: Progressive Learning from Complex Explanation Traces of GPT-4
arxiv.org
arxiv.org
I worked at a company once, where four teams tried different techniques. The best approach was shipped to users and the other three ended up as papers at a major conference.
Just people playing the capitalism game.
The interesting piece is if we can get LLMs to reason well, that finite reward capture can take LLMs much farther than possibly humans could do with the same signal.
Similarly to when a machine has an efficiency of 1%, optimizing it to 10% does not violate thermodynamics.
There is a limit to what you can learn from limited knowledge. Every bit of information you learn can only divide the space of possible theories by half. However, current LLMs are ludicrously far from that limit.
And if you _do_ let LLMs look at the world, and train on that, they also wouldn't have that limitation.
it may be higher quality, but far from perfection.
After many rinses and repeats dataset will accumulate significant amount of errors.
If you ask for an answer first and then ask for an explanation after, you will most likely get self-justifying BS for an answer. You have to ask for the explanation first.
https://twitter.com/search?q=%23ORC%CE%9B&src=typed_query&f=...