Explainer: What's r1 and everything else?
timkellogg.me
timkellogg.me
It seems like the AI space is moving impossibly fast, and its just ridiculously hard to keep up unless 1) you work in this space, 2) are very comfortable with the technology behind it, so you can jump in at any point and understand it.
R1 or the R1 finetunes? Not the same thing...
HF is busy recreating R1 itself but that seems to be a pretty big endevour not a $30 thing
And while this is true that this experiment shows that you can reproduce the concept of direct reinforcement learning of an existing LLM, in a way that makes it develop reasoning in the same fashion Deepseek-R1 did, this is very far from a re-creation of R1!
Most important, R1 shut down some very complex ideas (like DPO & MCTS) and showed that the path forward is simple, basic RL.
This isn't quite true. R1 used a mix of RL and supervised fine-tuning. The data used for supervised fine-tuning may have been model-generated, but the paper implies it was human-curated: they kept only the 'correct' answers.Does this guy know people were writing verbatim the same thing in like... 2021? Still always incredible to me the same repeated hype over and over rise to the surface. Oh well... old man gonna old man
Given how far gen AIs have improved since 2021, these people were quite spot on.
I'm not asking for the proof. Just the source, even a self-claimed statement. I've read the R1's paper and it doesn't say the number of $5.6M. Is it somewhere in DeepSeek's press release?
> We are living in a timeline where a non-US company is keeping the original mission of OpenAI alive - truly open, frontier research that empowers all. It makes no sense. The most entertaining outcome is the most likely.
> DeepSeek-R1 not only open-sources a barrage of models but also spills all the training secrets. They are perhaps the first OSS project that shows major, sustained growth of an RL flywheel. (…)
With distillation, can a model be made that strips out most of the math and coding stuff?
There's a person keeping track of a few writing prompts and the evolution of the quality of text with each new shiny model. They shared this link somewhere, can't find the source but I had it bookmarked for further reading. Have a look at it and see if it's something you'd like.
https://eqbench.com/results/creative-writing-v2/deepseek-ai_...
R1 is quite large (685B params). I’m wondering if you can make a distilled R1 without the coding and math content. 7B works well for me locally. When I go up to 32B I seem to get worse results - I assume it’s just timing out in its think mode… I haven’t had time to really investigate though.
The R1 sample reads way better than anything else on the leaderboard to me. Quite a jump.
The only measurable flaw I could find was the errant use of an opening quote (‘) in
> He huffed a laugh. "Lucky you." His gaze drifted to the stained-glass window, where rain blurred the world into watercolors. "I bombed my first audition. Hamlet, uni production. Forgot ‘to be or not to be,' panicked, and quoted Toy Story."
It's pretty amazing I can find no fault with the actual text. No grammar errors, I like the writing, it competes with the quality and engagingness of a large swath of written fiction (yikes), I wanna read the next chapter.
The lack of "gpt-isms" is really impressive IMO.
[1] https://eqbench.com/results/creative-writing-v2/meta-llama__... [2] https://eqbench.com/results/creative-writing-v2/gemini-1.5-f... [3] https://eqbench.com/results/creative-writing-v2/claude-3-opu... [4] https://eqbench.com/results/creative-writing-v2/mistralai__M... [5] https://eqbench.com/results/creative-writing-v2/deepseek-ai_... [6] https://eqbench.com/results/creative-writing-v2/gpt-4o-2024-...
Nobody really saw the LLM leap coming
Nobody really saw R1 coming
We don’t know what’s coming
It's funny to watch too. I played with Gemini 2.0 on Google AI Studio and asked it to "come up with your favorite song as you take a long walk to really think this through".
The reasoning can then be shown, and it talked to itself, saying things like "since I'm an AI, I can't take walks, but with a request like this, the user seems to imply that I should choose something that's introspective and meaningful", and went on with how it picked candidates.
In the reasoning process it concludes on: From the brainstormed genres/artists, select a specific song. It's better to be concrete than vague. For this request, "Nuvole Bianche" by Ludovico Einaudi emerges as a strong candidate. Craft the Explanation and Scenario: Now, build the response around "Nuvole Bianche."
Then in the actual answer it proposes: "Holocene" by Bon Iver.
=)
In summary, even though they are called "reasoning" models, they are still based on prediction and pattern matching, not true logical reasoning. The improvement in accuracy is likely due to better leveraging of the model's statistical knowledge, rather than any deeper understanding of the problem's logic. And the reasons you see it output have nothing to do with the actual reasons it used to determine the answer.
In fact, R1.Zero hints that, it might be even better to let the AI follow a chain of thought that doesn't actually make logical sense or is understandable, and that doing so could even further improve its ability to accurately predict solutions to code, math and logic problems.
> ARC-AGI is a benchmark that’s designed to be simple for humans but excruciatingly difficult for AI. In other words, when AI crushes this benchmark, it’s able to do what humans do.
That's a misunderstanding of what ARC-AGI means. Here's what ARC-AGI creator François Chollet has to say: https://bsky.app/profile/fchollet.bsky.social/post/3les3izgd...
> I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means.
> It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations.
> Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at something that isn't pure memorized skill. But it says rather little about how intelligent your system is, or how close to human intelligence it is.
Misunderstanding benchmarks seems to be the first step to claiming human level intelligence.
Additionally:
> > ARC-AGI is a benchmark that’s designed to be simple for humans but excruciatingly difficult for AI. In other words, when AI crushes this benchmark, it’s able to do what humans do.
Doesn’t even make logical sense.
Common non-technical chain of thought after learning this: 'Previously, only humans could play chess. Now, computers can play chess. Therefore, computers can now do other things that previously only humans could do.'
The error is assuming that problems can only be solved via levels of human-style general intelligence.
Obviously, this is false from the way that computers calculate arithmetic, optimize via gradient descent, and innumerable other examples, but it does seem to be a common lay misunderstanding.
Probably why IBM abused it with their Watson marketing.
In reality, for reliable capabilities reasoning, the how matters very much.
It's known as "hallucination" a.k.a. "guessing or making stuff up", and is a major challenge for human intelligence. Attempts to eradicate it have met with limited success. Some say that human intelligence will never reach AGI because of it.
I’m sure such a product would be met with ridicule considering how often humans hallucinate. Especially since, as we all know, the only use for humans is getting responses given some prompt.
That’s a description of the entire service economy.
Or more on topic see the improvements in LLMs since they were invented. At first each release was an order of magnitude better than the last (see GPT 2 vs 3 vs 4), now they’re getting better but at a much slower rate.
Certainly feels like being at the top of an S curve to me, at least until an entirely new architecture is invented to supersede transformers.
There was also a startup selling/renting bitcoin miners that doubled as electrical heaters.
The problem is that computers are fundamentally resistors, so at most you can get 100% of the energy back as heat. But a heat pump can give you 2-4 times the energy back. So your AI work (or bitcoin mining) plus the capital outlay of the expensive computers has to be worth the difference.
But if we can make computers that run at, say, 2000 degrees, without using several times more electricity, then we can capture their waste heat and turn a big portion of it back into electricity to re-feed the computers. It doesn't violate thermodynamics, it's just an alternative possibility to make more computers that use less electricity overall (an alternative to directly trying to reduce the energy usage of silicon logic gates) as long as we're still well above Landauer's limit.
That aside, we would need to see some evidence of AI developments being bootstrapped by the previous SOTA model as key part of building the next model.
For now, it's still human researchers pushing the SOTA models forwards.
When people use the term exponential I feel that what they really mean is 'making something so _good_ that it can be used to make the N+1 iteration _more good_ than the last.
https://www.lesswrong.com/posts/qLe4PPginLZxZg5dP/almost-all...
Where's the falsifiable framework that demonstrates your conclusion? Or are we just supposed to trust your intuition?
I like that the HN crowd wants to believe AI is hype (as do I), but it's starting to look like wishful thinking. What is useful to consider is that once we do get AGI, the entirety of society will be upended. Not just programming jobs or other niches, but everything all at once. As such, it's pointless to resist the reality that AGI is a near term possibility.
It would be wise from a fulfillment perspective to make shorter term plans and make sure to get the most out of each day, rather than make 30-40 year plans by sacrificing your daily tranquility. We could be entering a very dark era for humanity, from which there is no escape. There is also a small chance that we could get the tech utopia our billionaire overlords constantly harp on about, but I wouldn't bet on it.
Mr. Musk's exitement knew no bounds. Like, if they are the ones in control of a near AGI computer system we are so screwed.
Unfortunately, we seem to be on this exact trajectory. If open source AGI does not keep up with the billionaires, we risk sliding into an inescapable hellscape.
Dunno about Zuckerberg. Standing still he has somewhat slided into the saner spectrum of tech lords. Nightmare fuel...
"FOSS"-ish LLMs is like. We need those.
[0]: https://ourworldindata.org/grapher/exponential-growth-of-par...
[1]: https://ourworldindata.org/grapher/exponential-growth-of-dat...
[2]: https://epoch.ai/blog/trends-in-training-dataset-sizes
[3]: https://ourworldindata.org/grapher/exponential-growth-of-com...
Obviously predicting the future is hard, and we won't know where this stops till we get there. But I think a degree of skepticism is warranted.
It certainly will be sigmoid-shaped in the end, but the top of the sigmoid could be way beyond human intelligence.
I'm not a fan of this meme that seems to be very popular on HN. Someone with knowledge in EE and drivers can easily acquire enough programming knowledge in the higher layers of programming, at which point they can fill the gaps and understand the entire stack. The only real barrier is that hardware today is largely proprietary, meaning you need to actually work at the company that makes it to have access to the details.
Things can be complex without being intelligent.
> Exponentially smarter AI meets exponentially more difficult wins.
Another is that it doesn't seem like intelligence is the main/only bottleneck to producing better AIs right now. OpenAI seems to think building a $100-500B data center is necessary to stay ahead*, and it seems like most progress thus far has been from scaling compute (not to trivialize architectures and systems optimizations that make that possible). But if GPT-N decides that GPT-N+1 needs another OOM increase in compute, it seems like progress will mostly be limited by how fast increasingly enormous data centers and power plants can be built.
That said, if smart-human-level AGI is reached, I don't think it needs to be exponentially improving to change almost everything. I think AGI is possibly (probably?) in the near-future, also believing that it won't improve exponentially doesn't ease my anxiety about potential bad outcomes.
*Though admittedly DeepSeek _may_ have proven this wrong. Some people seem to think their stated training budget is misleading and/or that they trained on OpenAI outputs (though I'm not sure how this would work for the o models given that they don't provide their thinking trace). I'd be nervous if it was my money going towards Stargate right now.
Babies are born with a fully functioning image recognition stack complete with a segmentation model, facial recognition, gaze estimator, motion tracker and more. Likewise, most of the language model is pre-trained and language acquisition is in large part a pruning process to coalesce unused phonemes, specialize general syntax rules etc. Compare with other animals that lack such a pre-trained model - no matter how much you fine-tune a dog, it's not going to recite Shakespeare. Several other subsystems come online in the first few years with or without training; one example that humans share with other great apes is universal gesture production and recognition models. You can stretch out your arm towards just about any human or chimpanzee on the planet and motion your hand towards your chest and they will understand that you want them to come over. Babies also ship with a highly sophisticated stereophonic audio source segmentation model that can easily isolate speaking voices from background noise. Even when you limit yourself to just I/O related functions, the list goes on from reflexively blinking in response to rapidly approaching objects to complicated balance sensor fusion.
Can we do better than evolution? Probably; evolution is a fairly brute force search approach and we are pretty clever monkeys. After all, we have made multiple orders of magnitude improvements in the state of the art of computations per watt in just a few decades. Can we do MUCH better than evolution at finding efficient intelligences? Maybe, maybe not.
So six billion bits since two bits can represent four values. Base pairs and bases are effectively the same because (from the link) "the identity of one of the bases in the pair determines the other member of the pair."
And this same RL is also creating improvements in small model performance.
So, more LLMs are about to rise in quality.
If yes, then you get exponential increases very trivially. If no, then something external continues to bottleneck progress.
Take the trajectory of chess. handcrafted rules -> policies based on human game statistics -> self-play bootstrapped from human games -> random-initialized self-play.
And if the improvements it makes are not asymptotically diminishing.
If that is a normal human estimation I would guess in reality it is more likely to be in 6-10 years. Which is still good if we get it in 2030 - 2035.
I want to say that I have all the respect and admiration for these Chinese people, their ingenuity and their way of doing innovation even if they achieve this through technological theft and circumventing embargoes imposed by US (we all know how GPUs find their way into their hands).
We are living a time with a multi-faceted war between the US, China, EU, Russia and others. One of the battlegrounds is AI supremacy. This war (as any war) isn’t about ethics; it’s about survival, and anything goes.
Finally, as someone from Europe, I confess that here is well known that the "US innovates while EU regulates" and that's a shame IMO. I have the impression that EU is doing everything possible to keep us, European citizens, behind, just mere spectators in this tech war. We are already irrelevant, niche players.
The only way to win this war is to deescalate. Everybody wins.
And AI competition is a good thing for Europe especially when it lags behind technologically.
But there is absolutely no way that will happen, so the pragmatic question is which horse to bet on.
One can't help wondering what kinds of classified AI results the US military is getting when running on El Capitan.
> I have the impression that EU is doing everything possible to keep us, European citizens, behind, just mere spectators in this tech war. We are already irrelevant, niche players.
Citizens of the US are just as irrelevant, if not more, since none of the productivity gains trickles down to them. Their real wage growth has stagnated since the 1970s, and each year that goes by, their actual power to purchase more goods or services goes down.
The victories in AI only matter to those who will profit from it.
When it comes to the citizens of a country benefiting from AI or not, being the leader in AI tech is not very important. It is more a matter of if AI benefits them or not. That their country has the leading AI tech can as easily results in them having less jobs, and being paid worse, as it could the opposite, depending on the policies of that country.
But given that, it can very much be better to live in the EU, with second grade open-source models, but where the productivity benefits of AI benefit the general citizens, then to live in the US, where the productivity benefits of AI benefit only the few.
Long version: It's marketing efforts stirring up hype around incremental software updates. If this was software being patched in 2005 we'd call it "ChatGPT V1.115"
>Patch notes: >Added bells. >Added whistles.