The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.
That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)
Error rate doesn't prove anything. The nature of the errors is what matters.
Are you actually claiming LLMs operate based on human-like intelligence?
Humans keep overestimating just how high the bar of "human-like intelligence" is.
You could have humans calculate 2+2 all day and get a surprisingly high error rate. That reveals a flaw in how humans operate.
LLMs fail for entirely different reasons. Their mistakes don't imply they're human-like at all.
It's not about the error rate.
If you're using the existence of flaws in LLMs to deny the claim of intelligence to them, then why do "generally intelligent" humans exhibit some impressively similar-looking flaws?
And, if we're talking about that conspicuous similarity - do they actually fail "for entirely different reasons"? Or do you just want the reasons to be "entirely different" - and not the same reasons viewed at a different angle?
Because the similarities between humans falling for trick questions or scams, and LLMs falling for adversarial questions or prompt injections don't look coincidental to me at all.
One of the oldest patterns in scamming is overwhelming and confusing the victim. Numerous prompt injection methods seek to overwhelm and confuse an LLM - if an LLM can't keep track of things, can't grasp what's going on, it's far more likely to lose track of what's a prompt and what's data, overlook past instructions or go past its behavioral guardrails.
And humans who fall for trick questions like "1kg of feathers" or "captain's age" due to shallow attention and naive pattern matching? They fail in surprisingly similar ways to how LLMs fail on SimpleBench tasks that are filled with overwhelming adversarial distractors. Many "trick questions" are tricky to humans and LLMs alike - to the point that it's unlikely to be coincidental.
That's not the point at all. It's the fact that they fail in ways completely unlike humans.
You also have the burden of proof reversed. Its on you to prove these LLM agents are human-like intelligences if that's your claim. No one can prove this because it's false.
They hallucinate tool state, drift from the objective while seeming to comply, switch languages randomly (Cyrillic or Japanese characters in output), confuse tasks they've planned for completed ones, and of course follow prompt injections embedded in files or web pages.
(And in particular, switching languages on the fly is normal for people who speak more than one well, it's something you learn not to do for the sake of people less comfortable with the languages involved.)
Who would count that as prompt injection? It's a superficial analogy.
If you were vulnerable to prompt injection, I could order you to do absolutely anything you're capable of doing and you would be helpless to do otherwise.
And at the same time, whole books have been written about how reliably we can induce certain behaviours from humans.
E.g. the Blue-seven phenomenon [1] - I've personally experienced that second hand and it was how I learned about it by searching for it subsequently because I suspected it was a known thing, having read about cold reading before. A co-worker came back from lunch and recited a story about a cold reader that had run a routine on him exploiting the blue-seven phenomenon, and I knew before the story finished that the answer would be "blue" and "seven".
See also Cialdini's book "Influence" which is full of examples of just how predictable peoples reactions are to a whole lot of things.
That there isn't a perfect overlap does not mean there aren't plenty of similar "hacks" that causes us to respond in very predictable ways.
[1] https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon
Yes. Depending on specificity and timescales involved, we call that "reading comprehension" or "social engineering" or "peer pressure" or "motivating literature" or "advertising" or "propaganda" or "religion".
In fact just the other day I commented on it to my fiancee after I randomly switched to French because we were discussing a trip and I mentioned a French location and pronounced it in French, and suddenly I was in "French mode" entirely unintentionally and it took a sentence before I realised.
That you think this is unique to LLM's suggests you simply don't know the diversity of human thought as well as perhaps you think you do. That's fine - none of us have a very complete view of that.
When you suggest that is a "superficial analogy" after you were the one pointing out LLMs switching language as something that sets them apart, you're seriously reaching.
I can often pinpoint afterward what was likely the trigger: E.g. I used a word that is the same in two languages, and continue in the second; I pronounced a word in its native language for whatever reason, and continued in that language; my "context" suddenly included another language because someone else spoke the other languages within earshot of me.
What makes you think this is materially different from an LLM switching language because its probability distribution gives a word in a different language because it fits in context?
In the examples I gave, each even made a word in the language I switched to more probable as a reasonable continuation, just as with an LLM.
I'm not claiming the mechanisms are identical, or even similar, but the behaviour most certainly is more similar than "a superficial analogy" would imply.
But in practice, all the analogies I've seen are in fact superficial, including this one.
The LLM that abruptly switches languages will also likely switch to a wildly unrelated topic. If a human behaved that way, you'd call a doctor.
> The LLM that abruptly switches languages will also likely switch to a wildly unrelated topic. If a human behaved that way, you'd call a doctor.
Haven't met many kids, I see. Or even normie adults talking. I know plenty that tend to jump from topic to topic once they get into a stride talking, and they're not the ones diagnosed with ADHD.
By the way, it's not switching topic. You just pick the concept closest to what you mean from your combined vocabulary. If you're not paying close attention, you might switch language though (until the next concept you need is from the other language again, at which point you switch back)
And you're aware the paper "Attention is all you need" came out of machine translation research at Google, right? You hold an internal semantic representation and map in and out from arbitrary natural languages. I think the (bi-, tri-, multi-)lingual approach is the only proper way to translate, and this is a hill I will fight on!
Google may have gotten more than they bargained for on that particular translation experiment; though they failed to capitalize on it initially, with OpenAI running with the ball.
I'm sure there are other failure modes where it may change topics too, just like humans also regularly digress when triggered by certain words etc.
Your argument "AI needs supervision, therefore it fails in different ways than its operator does" holds.
Your argument "AI fails in different ways than its operator, therefore the AI's intelligence is different in kind" doesn't hold.
Ok, so we've established that it doesn't work like a human being. To paraphrase Dijkstra: The submarine doesn't swim.
But does it exactly sail either? An LLM doesn't exactly work like traditional deterministic software either, does it?
And yet it moves. You can put in data and ask it to process it, and you'll get an answer that's in some ballpark. Closer to quantum or stochastic computing perhaps, but that's not it either, is it? Or SAT-solving? Eh. It's its own computing approach. If you have a problem where the asking is hard but the verification is cheap, it might just be the right tool for the job.
That experiment "proved" that LLMs are statistical text generators without a concept of meanings of words. Which is the same thing as "proving" that there aren't a million tiny humans inside your laptop doing the CPU's work by hand.