There's a reason why ai xrisk doomers had to come up with the term ASI.
I would seriously suggest that everyone take a look at the wikipedia page for AGI from the month before ChatGPT was released, compare it to the current version, and not come to that conclusion.
https://en.wikipedia.org/w/index.php?title=Artificial_genera...
I have not seen any instance of this frequently-made assertion which is at all justified. It seems to rely on a definition of "understand" which is more about spirituality than actual observable evidence (they clearly can comprehend even complex tasks well enough to execute on them, and if you won't call that "understanding", you're playing word games rather than stating an objective fact).
Likewise, agents can literally come to a greater understanding of a problem through trial and error, and there are plenty of mechanisms to retain that knowledge. If you don't want to call that "learning", you're just making a choice to define it in a way more restrictive than how we use it for humans, and intentionally making communication more difficult.
"Understanding" has enough philosophical leeway in its use to allow at least the possibility of sentience as a prerequisite.
This is where the discussion about LLM capabilities becomes genuinely difficult, and dismissing that difficulty as "word games" or "spirituality vs evidence" is not helpful.
In fact, I'd argue that statements about what "is" and "is not" sentient relies on even more spirituality and word games for anything that isn't a terran tetrapod.
For a meaningful -- "helpful" -- discussion on such things, one has to assume that everyone is choosing a definition which is closer to the median usage and relies on not being totally subjective. Furthermore, given the breadth of options, it should be assumed to be a definition which allows which permits the form of the question to be meaningful, rather than begging the question -- if your definition is tautological enough that non-biological entities can't have understanding, you're just expressing dogma rather than having a discussion.
Anything else is bad faith, or assuming bad faith on the part of the participants.
I do not think it is at all unreasonable or "fringe" to regard understanding as involving intentionality: ie a directedness of thought toward the object-relations being "grasped". That may not be the only possible conception of understanding but it is a mainstream philosophical idea.
In fact, I'd argue that statements about what "is" and "is not" sentient relies on even more spirituality and word games for anything that isn't a terran tetrapod.
Then you seem to be confusing "hard to understand" with "meaningless".
you're just expressing dogma rather than having a discussion.
Anything else is bad faith, or assuming bad faith on the part of the participants
Have a think about that (repeated) tone before responding.
Fwiw I am a long-time believer in consciousness being fully realisable in machines; I think the jury is still out on LLMs.
Various criteria for intelligence have been proposed (most famously the Turing test) but to date, there is no definition that satisfies everyone
https://www.linkedin.com/pulse/announcing-aa-briefcase-bench...
AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.
That latter part is debatable though - have you seen a non-technical person try to figure out something new on a computer?
Our intelligence only seems "general" to us, because we're viewing it through our own eyes. Our "intelligence" is specialized to our survival, and we're terrible at most tasks outside that scope.
AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.
AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.
There have been many leaps forward in the past - tool calling, reasoning, agentic loops etc. 5.6 doesn’t have any of this. More intelligence doesn’t necessarily warrant a major version bump.