Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
arxiv.org
arxiv.org
> To evaluate comprehensibility quantitatively, we employ an LLM-as-a-Judge framework
This isn’t the worst idea, but it’s still a bit incestuous. Adding an LLM judge to check for hallucinations creates two new kinds of problems: false positives, where your judge hallucinates an incorrect fact, and false negatives, where the judge lets a hallucination slip by.
> We measure reproducibility through knowledge distillation. By fine-tuning a weaker model on the generated CoT traces, we use the downstream performance gain of the student as a proxy.
And my problem here, as a member of the GPU proletariat, is that this just seems incredibly inefficient. In other words, you’re going to generate a bunch of rollouts from your model then wait for the student to train? I guess if you have the compute to train a trillion params then maybe you don’t care.
Empirically, many problems look like they're easier to check than they are to solve. This seems like a reasonable way to bootstrap a little extra performance, with prior art in well-known DeepMind experiments. It's unclear if it works recursively (I imagine not), but the core idea is solid.
I’m specifically reacting to using it to reward the chain of thought during RL training because models love to hack their rewards, learning any possible shortcut rather than the task we want them to.
Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
However, a neuron is much more than a single parameter. The brain is estimated to have from 10^14 to 5x10^14 synapses.
If I had a magic button I would not only pause AI development but set it back 10 years. Sadly I have no influence on events and those who do, don't care about the future of humankind or actively wish us dead.
I'm pretty confident that algorithm/hardware design insights over the next decades will allow us to build human-rivaling cognitive abilities on basically todays consumer tech (analogously to how computer chess improved over time).
huge parameter models with many small but efficient layers can work quickly on low resource hardware
similar to how neurons experience chemical spiking to activate small portions of the brain at once
In 1956 a 5mb hard drive shipped on a large truck and took a team of men to unload. It consumed huge amounts of power, and cost about $3,200/month to run. In today's dollars that would be about $160,000 per month.
Aren't you glad we didnt just give up because it was kind of expensive?
Maybe there’s also a hardware component to it, but there’s very little point in trying to optimize the hardware to work with a poor algorithm. Once we discover an efficient way to train and infer, then it will be worth hyper-engineering the hardware.
Yes, we need better architectures and algorithms. We can point to massive advances in software as well in many spaces, including in LLMs (e.g. compare early GPT versions with current smaller open models), but the hardware comparison came from further up-thread.
There is a lot of low hanging fruit here as models stabilize, and the more ambitious tech that might take a decade or so to land offers efficiencies 10-100x compared even to biological systems. (Most of it requires cryogenic temperatures though). A forward looking tech investment might be inexpensive, tiny, efficient cryogenic cooling systems optimized for desktop or portable use.