It might be the 200M user base of OpenAI that provided the necessary guidance for advanced CoT, implicitly. Every user chat session is also an opportunity for the model to get feedback and elicit experience from the user.
Does Midjourney output look like an average human drawing?
Obviously, OpenAI knows how to train a classifier...
No, perhaps because it's heavily trained on photos.
A LLM can therefore have an higher IQ because it can combine all fields.
Also parameters and architecture might or might not be a limiting factor to us humans or a LLM. But LLM and parameter size, optimizations etc. are just at the beginning.
If we now have a good reasoning llm, we can build more test data automatically. Basically using the original content + creating new ones which can then lead to new knowledge = research.
If they were just doing prompt engineering and multiple inferences they'd definitely want to keep that a competitive secret and send all the open source devs off in random directions, or keep them guessing, rather than telling them which way to go to replicate Q-Star.
This is still clearly CoT, with all its limitations and caveats as expected. That's an improvement, sure, but definitely not a qualitative leap like OAI is trying to present it. (in a really shady manner)
Saying it's just CoT is kind of meaningless. Even just looking at the examples on Open AI's blog and you quickly see no other model today can generate or utilize CoT to anywhere near that quality through prompting or naive fine-tuning.
Starting from the strawberry example: it counts 3 "r"s in "strawbery", because the training makes it ignore grammatical errors if they're not relevant to the conversation (which makes sense in an instruction-tuned model) and their CoT doesn't catch it because it's not specialized enough. Will this scale with more compute thrown at it? I'm not sure I believe their scaling numbers. The proof should be in the pudding.
I've also had mixed results with coding in Python, it's barely better than Sonnet in my experience, but wastes a lot more tokens, which are a lot more expensive.
They might have improved things and made a SotA CoT that works in tandem with their training method, but that is definitely not what they were originally hyping (some architecture-level mechanism, or at least something qualitatively different). It also pretty obviously still has limited compute time per token and has to fit into the context (which is also still suffering from the lost-in-the-middle issue by the way). This puts the hard limit on the expressivity and task complexity.
That's an incredibly cynical choice of phrasing.
Of course they don't want to help the competition, that's what a competition is. The competition isn't helping OpenAI either.
I'm not convinced OpenAI is using one model. Look at the thinking process (UI), which takes time, and then suddenly, you have the output streamed out at high speed.
But even so, people are after results, not really the underlying technology. There is no difference of doing it with one model vs multiple models.
According to OpenAI, the model does it's thinking behind the scenes, then at the end summarizes that thinking for the user. We don't get to see the original chain-of-thought reasoning, just the AI's own summary of that reasoning. That explains the output timing.
If so, I imagine o1 clones could just be fine tunes of llamas initially.
Example prompt for that: "give me three sentences that end in 'is'."