If they were the same, I would have expected explicit references to o3 in the system card and how o3-mini is distilled or built from o3 - https://cdn.openai.com/o3-mini-system-card.pdf - but there are no references.
Excited at the pace all the same. Excited to dig in. The model naming all around is so confusing. Very difficult to tell what breakthrough innovations occurred.
Competition is good.
It is the closed competition model that’s being left in the dust.
Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently.
OpenAI isn't the only company. The Tech companies being beaten massively by Microsoft in #of H100s purchases are the ones with a moat. Google / Amazon with their custom AI chips are going to have a better performance per cost than others and that will be a moat. If you want to get the same performance per cost then you need to spend the time making your own chips which is years of effort (=moat).
DeepSeek has proven that the latter is possible, which drops a couple of River crossing rocks into the moat.
Google with all its money and smart engineers was not able to build a simple chat application.
It seems they have high quality trainingsdata. And the knowledge to work with it.
... is definitely something I've said before, and recently, but:
> That moat is trivially circumvented by lighting a larger pile of money
If that was true, someone would have done it.
The deepseek paper states that the $5mil number doesn't include development costs, only the final training run. And it doesn't include the estimated $1.4billion cost of the infrastructure/chips Deepseek owns.
Most of OpenAI's billion dollar costs is in inference, not training. It takes a lot of compute to serve so many users.
Dario said recently that Claude was in the tens of millions (and that it was a year earlier, so some cost decline is expected), do we have some reason to think OpenAI was so vastly different?
Inference capex costs are not a defensive moat as I can rent gpus and sell inference with linear scaling costs. A hypothetical 10 billion dollar training run on proprietary data was a massive moat.
https://www.itpro.com/technology/artificial-intelligence/dol...
I find huge value in these models as an augmentation of my intelligence and as a kind of cybernetic partner.
I can't think of anything that can actually be automated though in terms of white collar jobs.
The white collar model test case I have in mind is a bank analyst under a bank operations manger. I have done both in the past but there is something really lacking with the idea of the operations manager replacing the analyst with a reasoning model even though DeepSeek annihilates every bank analyst reasoning I ever worked with right now.
If you can't even arbitrage the average bank analyst there might be these really non-intuitive no AI arbitrage conditions with white color work.