It seems like if they can do it, that there's no reason they can't eventually be trained to do it better up to and beyond human performance. It seems strange to suggest that thinking unlocks some nominal margin of "better" specifically that can't be overcome.
All of that aside, even if they can't outperform the top human programmers...what if they get to within a margin where they're still better than most? Isn't a 95th percentile programmer that can run 24/7 and continuously refine its work still going to ultimately come out on top?
I suspect it largely has to do with how one defines "thinking". It seems like people like to implicitly define it in such a way as to require a human (or animal), but there are many examples of thinking/intelligence in nature that don't require a brain or even neurons.
I'm genuinely curious: without using the word "think" with all of its ambiguity, can you articulate what it is that we're doing that these models are not capable of? Because it's pretty clear (to me, at least) from the research, particularly a lot of the mechanistic interpretability work coming out of Anthropic, that the models are at least doing something akin to what we think of as thinking, even if it appears foreign to us.
Like, I'm not sure how you could read this and not see some spark of seems like thinking: https://www.anthropic.com/research/tracing-thoughts-language...
It's pretty clear that there are significant differences between their intelligence and human intelligence. But that doesn't mean there isn't some sort of intelligence here.
The issue about AI is that it's gobbling so much information that at some point you couldn't tell the difference. Programming specifically is something that inherently documents itself, meaning while human communication and context and memes and culture is something that evolves and exists many times outside of textual mediums, as soon as any new piece of code is born it is now part of the AI's dataset. And it doesn't help that a vast majority of our code is pretty damn repetitive, especially if you insert code written in the span of two decades and more into the future.
Tldr : The better we get at coding, the more code we write, the better AI gets at coding.