Humans can't one-shot non trivial planning tasks either. It's the one problem i have with all the papers that try to evaluate planning for LLMs.
Step away from that approach and they're ok.
We are nowhere near generally intelligent software systems.
There could def be bugs I missed tho.
https://chat.openai.com/share/ef77507e-cb75-4112-97f1-a16cfc...
I'll repeat what I said previously in a different way, LLMs are useful but they are nowhere near what is required to achieve generally intelligent software systems. I'm sure they will continue to improve as engineers and companies learn how to utilize them in their workflows but let's temper the hype a little bit because statistical autocompletion is not enough to achieve general intelligence.
Anyways, these kinds of strawmen always baffle me in regards to AI.
Insert random pseudo trivia that most of the general population wouldn't be able to do, see the AI fail at that specific task. "What did I tell you? The AI definitely isn't generally intelligent yet!"
Everybody is out to prove AI isn't intelligent without first defining what intelligence even is. And when other people rightly point out it can do a lot of stuff, they point to some specific task that it gets 90% of the way there but doesn't get perfect and then triumphantly declare AI isn't intelligent. Crazy.
LLMs are cool toys but calling them intelligent is stretching the definition of "intelligent" way too far. It's important to be clear about what the words actually mean because if people start thinking these software systems can be substituted for their own thinking then we end up with all sorts of unnecessary confusion around what they're actually capable of achieving.
And GPT-4 does do this in a couple iterations.
Adding recursion to neural networks has been tried a few times but no one actually knows how to stabilize their dynamics so the industry has settled on feed forward networks with constrained function blocks which have stable dynamics with respect to back propagation of errors.
> What exactly do you want me to state up front?
Exactly what was in your comment that I told you to state up front.
Most people in a discussion about AI and replacing people are working with a definition of general intelligence that includes humans to a large degree.
> This is because LLMs and all neural networks are simply DAGs of function which do not support recursion or backtracking.
Without adding external state I can't solve a sudoku puzzle.
And how many humans can do this? Those poor exhausted goalposts.
It's taken a few short years to go from "keep a coherent story over more than a sentence" to "write a multithreaded sudoku solver in one shot".
Many programmers would fail to do this. I'm not even sure a randomly selected human would understand the question. And I'm wondering if you've setup any loops letting it write tests, search the internet, inspect and debug? If not what you see as an output is essentially it whiteboarding off the top of its "head". Programmers routinely fail to solve simpler things in interviews, and their skills are much narrower than current llms (you can easily argue deeper, but I feel comfortable saying broader).
I did not say anything about threads but I did hint at the fact that coroutines can be used creatively to solve problems which require backtracking and constraint propagation. The fact that this is still controversial means a lot of people are unaware LLMs do not support backtracking, recursion, and constraint propagation. Sudoku is just an obvious example of a problem that is easy to solve if you know about these concepts and next to impossible if you don't. Such problems can also be expressed as integer programs but even fewer people know how to do that so I don't usually bring it up but it would be another good test for any software system that is claimed to be intelligent by the corporate marketing department.
> Such problems can also be expressed as integer programs but even fewer people know how to do that so I don't usually bring it up but it would be another good test for any software system that is claimed to be intelligent by the corporate marketing department.
It's very telling if the level of testing is "can it do something few people in the field can?".
https://gist.github.com/IanCal/9817f8b21b2ea6d77940966ee399d...