The GPT era is already ending
theatlantic.com
theatlantic.com
Claude is regarded as better at many things, Gemini at some others, Gemini has massive context windows, Flash runs faster and cheaper, Bedrock runs cheaper, Gemma and Llama are open source and doing pretty well for many use-cases.
If it weren't for mindshare, I'm not sure OpenAI's models would be thought of as anything special at this point. Now it does sound like o1 is a step up, but it also takes a long time and costs a lot, making it quite niche. It's not something you drop in and get better results from.
So we are not allowed to? The Hacker News gatekeeping instinct is particularly hilarious to me.
Ten minutes and a teeny bit of mental real estate I will never get back.
Given that the next token is always predicted based on everything that both the user and the model have typed so far, this seems like a false statement.
Practically, I've more than once seen an LLM go "actually, it seems like there's a contradiction in what I just said, let me try again". And has the author even heard about chain of thought reasoning?
It doesn't seem so hard to believe to me that quite interesting results can come out of a simple loop of writing down various statements, evaluating their logical soundness in whatever way (which can be formal derivation rules, statistical approaches etc.), and to repeat that various times.
However, the way in which the author is correct is that there is still functionally one stream, one "thought process", and the introspection is very limited and well defined based on how the model is constructed. If you ask a human a hard question they can take longer thinking about it, and that's not something an LLM can do at all. o1 is one attempt to add further streams of thought as a way to address that limitation, so that the LLM can continue to think and weigh up options through multiple rounds of generation.
This is akin to how real humans almost immediately will formulate an answer but then think before responding.. Except imagine if someone had a probe that looked into your mind and made you commit to the first thing you thought.
Companies like anthropic literally train their model on the introspected thoughts of the earlier model (constitutional AI). The model is capable of 'thought' if the system doesn't force its hand
But "very limited" isn't "not at all", and I think the distinction matters, since the former can often be scaled more easily than the latter. And at the meta-level, one possible output token can always be "let me ramble on this a bit more before I provide my final answer".
In any case, I think "they can't look back at their output", as presented in the article, is just the wrong way to look at it; even "they can (currently) not take longer time to look at their output" is also wrong, since they can just produce more output, which then in turn lets them "think more".
o1 however is often better for coding and for puzzle-solving, which are not the vast majority of uses of LLMs.
o1 is so much more expensive than 4o that it makes zero sense for it to be a general replacement. This will never change because o1 will always use more tokens than 4o.
You are confusing training artifacts for "personality."
> could be better for coding and for puzzle-solving, which are not the vast majority of uses of LLMs.
To see a product fail to evolve and merely stratify itself gives a solid hint as to what it's likely future is going to be.
4o versions have progressively become better at instruction following. I don't think their peak has been reached yet.
> They are merely generating what seems to "come next" logically.
So are you and I: We're predicting what we'll do next.
My guess is to sell it to governments and anyone else willing to pay for it.