If you still think there's a stall despite all evidence pointing to opposite, I don't know what to say..
If you still think there's a stall despite all evidence pointing to opposite, I don't know what to say..
Their best models might be quite good when directed at extremely difficult and very focused problems, but most business use cases are nothing like that.
Edit:
Not trying to imply these new frontier models aren’t also better at other things, just that there’s really no reason for most people to use them when cheaper alternatives exist that get the job done just as well.
Edit:
The "reputable academic" who accused OpenAI had to say this about LLMs and the recent result.
source: https://cims.nyu.edu/~tristanb/statement.pdf
> “the results are not the important thing.”
> “the important thing is instead the significance that a mathematician and an LLM model can now do all this work in a month.”
> “This is a Deep Blue-Kasparov moment.”
> “incredibly important developments.”
> “If indeed an OpenAI model did close the gap to Navier-Stokes, that is a remarkable thing and it should be said loudly, by them, with the history intact.”
Clearly Buckmaster (who is probably one of the most accomplished academics in the field) himself doesn't believe that AI has stalled. What makes you think you are right?
I will note the remarkable goal post shift in your edit - giving direct counter evidence to your own earlier claim of an OpenAI proof - and leave it at that.
Unfortunately, there will be people that refuse to tackle the issue head on and will instead narrow their focus on data that suggests that tomorrow will be like today.
1. It has people in it who are career obsessed and who are willing to ruthlessly go after any opportunity to improve their standing/stock valuation. Using a 2 week old model to snipe a millenium prize for PR is in line with that.
2. There are many researchers and even executives at the company who genuinely think we are speedrunning the end of the world. I don’t know anyone in this field who honestly argues that if we build ASI soon it doesn’t lead to extinction. This group of people can output warnings about the state of research and fears for the future while pushing for regulation out of genuine fear of what they’re building. I tend to agree with them.
You can’t view these companies as a monolith. They’re actions won’t be consistent because it’s built of many people with conflicting beliefs. Please look at arguments regarding AI risk and the current pace of progress and value them as it relates to the argument itself, not who said it. We are in a dangerous place and no one is sure how quickly we’ll get to a bad spot.
Say what? Could lead to extinction, sure. Does? And nobody honestly argues otherwise? Baloney.
I guess people are just going to keep spreading misinformation about this, along the same lines as "Anthropic's C compiler fails on hello world". There is zero indication that they stole anything, unless it counts as "stealing" to spin up a bunch of compute based on rumors that the problem had been solved already.
Are these real step changes - big picture wise, or refinements in RL/agentic orchestration/"taste" and advancements due to bigger models and hardware technology/capacity scaling? If they not, does this tactic - and hardware improvements - continue to scale non-linearly like they need to?
It is clear that whatever does change in each model increment has resulted in meaningfully better end user capabilities (as well as regressions in some areas, honestly), but that doesn't prove anything. I'm not sure what I personally believe, but stating with your full chest that a stall is ridiculous ignores a lot of potential evidence to the contrary.
Stupid example: Astra. Its main improvements are: much much better computer use and 3d modeling capabilities; better subagent orchestration; better and more reliable tool use; slightly worse coding.
This looks to me, from a feature perspective, to be an incremental improvement across several functional areas, plus new features which are unquestionably excellent but are probably the result of RL focus, not magic.
Step change? Ehhhh depends on how you squint. But how many more iterations of this do we have? Are we going to squeeze quintillion parameter transformers into GPUs?