I’m excited to see models become open, but given the curve of progress we’ve seen, even being “a little” behind is a gap that grows exponentially every day.
I’m excited to see models become open, but given the curve of progress we’ve seen, even being “a little” behind is a gap that grows exponentially every day.
Most importantly, this is a signal: openAI and META are trying to build a moat using massive hardware investments. Deepseek took the opposite direction and not only does it show that hardware is no moat, it basically makes fool of their multibillion claims. This is massive. If only investors had the brain it takes, we would pop this bubble alread.
I mean, sure, no one is going to have a monopoly, and we're going to see a race to the bottom in prices, but on the other hand, the AI revolution is going to come much sooner than expected, and it's going to be on everyone's pocket this year. Isn't that a bullish signal for the economy?
They do have the best models. Two models made by Google share the first place on Chatbot Arena.
In my experience doing actual work, not side by side comparisons, Claude wins outright as a daily work horse for any and all technical tasks. Chatbot Arena may say Gemini is "better", but my reality of solving actual coding problems says Claude is miles ahead.
I consider it unlikely that the new administration is philosophically different with respect to its prioritization of "national security" concerns.
That's not a crazy thing to say, at all.
Lots of AI researchers think that ASI is less than 5 years away.
> deepseek's performance should call for things to be reviewed.
Their investments, maybe, their predictions of AGI? They should be reviewed to be more optimistic.
If people can replicate 90% of your product in 6 weeks you have competition.
The moat for these big models were always expected to be capital expenditure for training costing billions. It's why these companies like openAI etc, are spending massively on compute - it's building a bigger moat (or trying to at least).
If it can be shown, which seems to have been, that you could use smarts and make use of compute more efficiently and cheaply, but achieve similar (or even better) results, the hardware moat bouyed by capital is no longer.
i'm actually glad tho. An opensourced version of these weights should ideally spur the type of innovation that stable diffusion did when theirs was released.
And this is based on what exactly? OpenAI hides the reasoning steps, so training a model on o1 is very likely much more expensive (and much less useful) than just training it directly on a cheaper model.
R1's biggest contribution IMO, is R1-Zero, I am fully sold with this they don't need o1's output to be as good. But yeah, o1 is still the herald.
Like, this idea always seemed completely obvious to me, and I figured the only reason why it hadn't been done yet is just because (at the time) models weren't good enough. (So it just caused them to get confused, and it didn't improve results.)
Presumably OpenAI were the first to claim this achievement because they had (at the time) the strongest model (+ enough compute). That doesn't mean COT was a revolutionary idea, because imo it really wasn't. (Again, it was just a matter of having a strong enough model, enough context, enough compute for it to actually work. That's not an academic achievement, just a scaling victory.)
This theory has yet to be demonstrated. As yet, it seems open source just stays behind by about 6-10 months consistently.
I thought that too before I used it to do real work.