What would a frontier API have to be able to do to satisfy you?
What would a frontier API have to be able to do to satisfy you?
They've been "patched" since but all models fail basic tests like "Should I walk or drive to the car wash which is 100 feet away" by recommending you walk.
So you'd just ask questions that require theory of mind, abstract and common sense reasoning, causal inference, learning novel rules, transferring knowledge novel situations, recognizing ambiguity, etc.
I think we should consider slime mold intelligent, and realise that it's a spectrum. Path finding is AI. There are probably forms of intelligence we have yet to discover.
I'm never sure whether this indicates "no reasoning present" or you've just hit an odd behaviour in the AI such that its reasoning fails. For example, you present a problem in a way that's dissimilar to the way problems are presented in its training set. That doesn't mean it's not reasoning, just it can only reason correctly in some circumstances.
I'm pretty sure most people building these models would admit they don't operate as human-like intelligences? It's baffling that anyone thinks they are.
But that doesn’t mean they don’t reason.
These LLM models/agents absolutely do not reason in the sense that humans do, so you're quietly redefining the word.
You can say of course decide to call them an "alien kind of intelligence" that "reasons" but you could just as reasonably say that calculators are an "alien" kind of intelligence that "reasons" about math differently than us.
Do you have a RealReasoningBenchmark, perhaps, that can reliably tell apart that fake mass produced token-flavored AI reasoning from the real, organic, 100% natural human reasoning?
We are now calling text and image generators "intelligent" in the same way a spell checker is intelligent.
Whatever it's become, "AI" research started as a way to study digital neurology, or how to digitize a mind, not just how to generate data.
The Turing Test should have had a caveat, it needs to fool a, "non-stupid" person, and we still have not gotten even close to passing that version.
my coworkers would know almost immediately if i did that.
The same would happen if you were replaced by any random human.
you might say almost no humans can do tht either but some human can but no ai can.