Understanding R1-Zero-Like Training: A Critical Perspective
github.com
github.com
[0] benchmark used by all major vendors to "showcase" coding ability, turns out to be <10% properly solved: https://www.youtube.com/watch?v=QnOc_kKKuac
Like Sabine says, if the LLM models, already read all the Math books in the world but are not yet able to do basic math, without calling upon a calculator, how much reasoning is really emerging?
"The Path to AGI is Coming Into View": https://youtu.be/mfbRHhOCgzs?t=219
Having trained quite a few LLMs by now, especially around the “uplift” from a text completion model to an instruct one, I noticed that the instruction following capabilities tend not to be uniform across all the tasks the LLM is able to perform.
In other words, it doesn’t know what is the implication of your request (time on the clock) in terms of change in output (the svg drawing of the clock), and all the abstractions and steps in between, this is why changing the time in the prompt might yield little difference.
Instruct datasets are quite small in retrospect and hardly cover the range of tasks instruct LLMs are able to perform, and mostly rely on extrapolation (interpolation?) capabilities for the LLM from its small training set. And it is this generalization that is not evenly distributed.
Just my two cents from my modest observations, sprinkle [citation needed] everywhere.
I don't deny that performance for certain logic tasks goes up with these models but I don't fully understand what role the thinking tokens take in these cases.
There's been publications with a pause token (https://arxiv.org/abs/2310.02226), backspace token (https://arxiv.org/abs/2306.05426), or a think token (https://arxiv.org/html/2405.08644v1#Ch0.S4), all of them based on the theory that a generic token can sort of act a placeholder for manipulating attention further without further meaningful output.
However, in practice, those approaches haven't been used in the training of a large scale model, i.e. I haven't seen it at all, the most adventurous people have gotten at scale is doing Mamba. (and RL)
* It had a particular technical meaning. The first round of the telephone game was when it came to mean "a 3 spatial dimensions-like space, with N dimensions, in an image diffusion model, that contains all possible image styles, that is navigated by a prompt." We're many iterations afield of it now, I'm afraid. Now, you sort of have to interpret it like you would negative space, defined by what it is around it.
I love this sort of “anti-hype” research. We need more of it.
"DeepSeek-V3-Base already exhibit 'Aha moment'."
I tried to read the screenshot they present as evidence of this, and indeed it does say "Aha!". But both the preceding reasoning and the following conclusion look like gibberish to me. I'm not sure what we're supposed to conclude here and I gave up reading the article after this inauspicious start.
Currently it seems that this shift of cost to inference-time is a necessary tradeoff that we have to live with (for now at least).
It's more of a problem for many people running those models locally because they have constrained hardware that can't handle those long contexts.
It seems that this paper shows not only that their method is cheaper in terms of fine tuning but also significantly reduces inference time cost for CoTs.
If what they say gets confirmed it looks to me like quite significant contribution?