141 karma · joined April 1, 2022
The RL techniques present will only work in domains where you can guarantee an answer is right (multiple choice questions, math, etc.). It doesn't really present any convincing leap forward in terms of advancing the capability of LLMs, just a strategy for compute efficient distillation of what we know already works. The fact this shitty PPO proxy works at all is a testament to the fact that DeepSeek is bootstrapping its capability heavily off of the output of existing larger models which are much more expensive to train. What DeepSeek R1 proves is you can distill a ChatGPT et al. into a smaller model and hack certain benchmarks with RL.
If you could just do RL to predict the best next word in general this would have been done already - but the signal to noise ratio on exploration would be so bad you'd never get anything besides infinite monkeys at a typewriter. It's not a novel/complicated idea to anyone familiar with RL to try and improve probability of things you like, and whoever decided to do RLHF on an LLM surely thought of (and did) regular RL first - and found it didn't work very well with whatever pretrained model and rewards they had. it was like two weeks ago people were going crazy about O3 doing arc-agi by running the exact same kind of traces R1 is doing in "GRPO" at test time rather than train time. Doing this also isn't novel and also only helps on shitty toy problems where you can get a number to tell you good vs bad.
There is no mechanism to compute rewards for general purpose language tasks - and in fact I think people will come to see the gains in math/coding benchmark problems come at a real cost to other capabilities of models which are harder to quantify and impossible to generically assign rewards to at internet scale.
To explore the frontier of capability you will still need a massive amount of compute, in fact even more to do RL than you would need to do standard next token prediction - even if the LLM might have fewer paramters. You also can't afford to do all the optimizations as you try many different complex architectures.
A lot of originally little automation/dev scripts bloat into more complicated things as edge cases are bolted on and bash scripts become abominations in these cases almost immediately.
bash script is "okay" I guess if your "script" is just a series of commands with no control flow.
Maybe at some point maybe we only act as meat-robots which shovel coal into the machine, but a lack of redundancy in GPT# due to it's own human like blind spots means it shuts down. Humans can no longer get it running again because they can't query it properly to help fix the complicated problems. The ability to even do the tasks or design systems required to keep modern world robust to unknown future disasters or breakdowns does not and will not exist in any of the training data. If we get rid of all knowledge work, we can no longer bootstrap things back to a working state should everything go wrong.
Maybe the current instantiation of GPT#/SD etc. pollute the training data with plausible but subtly flawed software, text, images etc. halting improvement around here. Maybe the ability to evaluate if the model improved becomes more noise than signal because it gets too vague what improvement even means. RLHF will already have this problem, as 100 people will have a 100 slightly different biases about what constitutes the "best" next token.
No matter how hard it tries, I think we can say GPT will not solve NP-Hard problems magically, it will not somehow find global optima in non-linear optimizations, It will not break the laws of physics, It will not make inherently serial problems embarrassingly parallel. It will probably not be more energy efficient at attempting to solving these problems, maybe just faster at setting up systems to try solving them.
Another trap, as it becomes more human like in its reasoning and problem solving capabilities, it starts to gain the same blind spots as us too, and also gains stochastic behavior which may cause it to argue with other instances of itself. I'm not convinced an AGI innovates at an unfathomable rate or even supersedes humans in all contexts. I'm especially not convinced a world filled with AGIs that is indistinguishable from a very intelligent human or corporation or what have you through imitation does any better at anything than the 9 billion embodied AGI agents that currently populate the earth.
There's other ways to tease out if someone is "dedicated to reaching a goal" like actually asking them questions about their life, daily routine, accomplishments etc. I think these are much better signals than "this person can implement many sorting algorithms and solve towers of hanoi or the egg drop puzzle without a google search to refresh their memory" with an N=1 sample size used to gauge how reliably they can do that. How many people if asked to do this on a job rather than an interview a year later would then go and implement it from memory without double checking they didn't forget some edge case on stackoverflow?
I know grinding leetcode is fundamentally useless because you will immediately start losing your ability to interview once you get hired. If you don't change jobs within a year you'll need to start studying all that crap again for the next interview.
I think this will introduce unavoidable background noise that will be super hard to fully eliminate in future large scale data sets scraped from the web, there's always going to be more and more photorealistic pictures of "cats" "chairs" etc. in the data that are close to looking real but not quite, and we can never really go back to a world where there's only "real" pictures, or "authentic human art" on the internet.
There are millions(?) of human researchers (and orders of magnitude more computers) doing gradient free optimization through their research in this direction, and the progress is painfully slow, I know because I'm one of them. There are billions of years of optimization (evolution) towards this goal, and a total of (1) species has achieved any kind of notable intelligence. We are collectively giant parallelized optimization.
We already have "AGI" orders of magnitude more capable than any single human in the form of billions of people networked through the internet searching for fulfillment, money, power, fame, etc. for the next big discovery or supporting this effort by providing everything the entire global "machine" needs to run. The idea one little box running the right program can have access to the energy to beat this effort and exponentially improve things seems laughable in comparison.
The global "AGI" formed by all of us, the internet, and computers, is more likely to destroy society in the next 20 years in some catastrophic event than some paperclip machine.