5,102 karma · joined March 27, 2020
I guess the nature of the problem lent itself to the 10k agents. Ie, there isn't something general to take here.
Fwiw, imo, you should assume everything you say (and write) is recorded going forward.
So, it doesn't appear there is any way to avoid the clinical trial process. And it won't speed up, and it won't get cheaper.
Am I wrong?
The reason AI is doing so well in math proof writing is that it can verify every idea it has, quickly.
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
Good going, hope it goes well.
And, the expected return from a state is often estimated rather than an explicit trajectory run out all the way. A value function can estimate expected return without explicitly predicting the future states or individual rewards that make up that return.
Fwiw, I use the phrase "total reward" above and use it as a synonym for "return", which is lazy use of language too.
If you want to understand how this stuff works, there are totally decent books about building them from scratch. It's not that hard, and you'll likely find it interesting. Sebastian Raschka and Nathan Lambert have good books out, and the Allen Institute has available all the code and data they have used for several projects.
They are no longer that thing due to post training. They simply aren't making a prediction, and they aren't even optimized for the next token. If I give a distribution of the heights of the population, I'm not giving a prediction either. Distributions don't imply predictions.
Why the desperation to hang onto the word "prediction"?
The policy is optimized to maximize the the total reward, defined as the sum of the reward at each step, discounted by some factor.