If you took the current state of affairs back to the 90s you’d quickly convince most people that we’re there. Given that we’re actually not, we’re now have to come up with new goalposts.
If you took the current state of affairs back to the 90s you’d quickly convince most people that we’re there. Given that we’re actually not, we’re now have to come up with new goalposts.
Consider this. I could walk into a club in Vegas, throw down $10,000 cash for a VIP table, and start throwing around $100 bills. Would that make most people think I'm wealthy? Yes. Am I actually wealthy? No. But clearly the test is the wrong test. All show and no go.
The more I think about this, the more I think the same is true for our own intelligence. Consciousness is a trick and AI development is lifting the veil of our vanity. I'm not claiming that LLMs are conscious or intelligent or whatever. I'm suggesting that next token prediction has scaled so well and cover so many use cases that the next couple breakthroughs will show us how simple intelligence is once you remove the complexity of biological systems from the equation.
To the extent that we vainly consider ourselves intelligent for our linguistic abilities, sure. But this underrates the other types of spatial and procedural reasoning that humans possess, or even the type that spiders possess.
It is an entirely different thing to language,which was created by humans to communicate between us.
Language is the baseline to collaboration - not intelligence
How do you define verbal language? Many animals emit different sounds that others in their community know how to react to. Some even get quite complex in structure (eg dolphins and whales) but I wouldn’t also rule out some species of birds, and some primates to start with. And they can collaborate; elephants, dolphins, and wolves for example collaborate and would die without it.
Also it’s completely myopic in terms of ignoring humans who have non verbal language (eg sign language) perfectly capable of cooperation.
TLDR: just because you can’t understand an animal doesn’t mean it lacks the capability you failed to actually define properly.
That’s like a bird saying planes can’t fly because they don’t flap their wings.
LLMs use human language mainly because they need to communicate with humans. Their inputs and outputs are human language. But in between, they don’t think in human language.
You seem to fundamentally misunderstand what llms are and how they work, honestly. Remove the human language from the model and you end up with nothing. That's the whole issue.
Your comment would only make sense if we had real artificial intelligence, but LLMs are quite literally working by predicting the next token - which works incredibly well for a fascimlie of intelligence because there is an incredible amount of written content on the Internet which was written by intelligent people
LLMs can't initiate any task on their own, because they lack thinking/intelligence part.
By the time you get to active "teaching", the child has already learned language -- otherwise we'd have a chicken-and-egg problem, since we use language to teach language.
An additional facet nobody ever seems to mention:
Human language is structured, and seems to follow similar base rules everywhere.
That is a huge boon to any statistical model trying to approximate it. That's why simpler forms of language generation are even possible. It's also a large part of why LLMs are able to do some code, but regularly fuck up the meaning when you aren't paying attention. The "shape" of code and language is really simple.
Just because we don’t understand something doesn’t mean there’s nothing there.
Also, I’m not so sure human language is structured the same way globally. There’s languages quite far from each other and the similarities tend to be grouped by where the languages originated. Eg Spanish and French might share similarities of rules, but those similarities are not shared with Hungary or Chinese. There’s cross pollination of course but language is old and humans all come from a single location so it’s not surprising for there to be some kinds of links but even a few hundred thousand years of evolution have diverged the rules significantly.
If cows were eating grass and conceptualising what is infinity, and what is her role in the universe, and how she was born, and what would happen after she is dead... we would see a lot of jumpy cows out there.
Words, after all are just arbitrary ink shapes on paper. Or vibrations in air. Not fundamentally different than any other signal. Meaning is added only by the human brain.
I don't think anyone would argue that animals don't communicate with each other. Some may even have language we can't interpret, which may consist of something like words.
The question is why we would model an AGI after verbal language as opposed to modeling it after the native intelligence of all life which eventually leads to communication as a result. Language and communication is a side-effect of intelligence, it's a compounding interest on intelligence, but it is not intelligence itself, any more than a map is the terrain.
Because verbal/written language is an abstracted/compressed representation of reality, so it's relatively cheap to process (a high-level natural-language description of an apple takes far fewer bytes to represent than a photo or 3D model of the same apple). Also because there are massive digitized publicly-available collections of language that are easy to train on (the web, libraries of digitized books, etc).
I'm just answering your question here, not implying that language processing is the path towards AGI (I personally think it could play a part, but can't be anything close to the whole picture).
All we know about animal consciousness is limited to behaviour, e.g. the subset of the 40 or so "consciousness" definitions which are things like "not asleep" or "responds to environment".
We don't know that there's anything like our rich inner world in the mind of a chimpanzee, let alone a dog, let alone a lobster.
We don't know what test to make in order to determine if any other intelligence, including humans and AI, actually has an inner experience — including by asking, because we can neither be sure if the failure to report one indicates the absence, nor if the ability to report one is more than just mimicking the voices around them.
For the latter, note that many humans with aphantasia only find out that "visualisation" isn't just a metaphor at some point in adulthood, and both before and after this realisation they can still use it as a metaphor without having a mind's eye.
> Language is the baseline to collaboration - not intelligence
Would you describe intercellular chemical signals in multicellular organisms to be "language"?
If be "we don't know" you mean we cannot prove, then, sure, but then we don't know anything aside from maybe mathematics. We have a lot of evidence that animals similar consciousness as we do. Dolphins (or whales?) have been known to push drowning people to the surface like they do for a calf. Killer whales coordinate in hunting, and have taken an animus to small boats, intentionally trying to capsize it. I've seen squirrels in the back yard fake burying a nut, and moving fallen leaves to hide a burial spot. Any one who has had a dog or a cat knows they get lonely and angry and guilty. A friend of mine had personal troubles and abandoned his house for a while; I went over to take pictures so he could AirBnB it, and their cat saw me in the house and was crying really piteously, because it had just grown out of being a kitten with a bunch of kids around and getting lots of attention, and suddenly its whole world was vanished. A speech pathologist made buttons for her dog that said words when pressed, and the dog put sentences together and even had emotional meltdowns on the level of a young child. Parrots seem to be intelligent, and I've read several reports where they give intelligent responses (such as "I'm afraid" when the owner asked if it wanted to be put in the same room as the cat for company while the owner was away [in this case, the owner seems to be lacking in intelligence for thinking that was a good idea]). There was a story linked her some years back about a zoo-keeper who had her baby die, and signed it to the chimpanzee (or gorilla or some-such) females when it wanted to know why she had been gone, and in response the chimpanzee motioned to with its eye suggesting crying, as if asking if she were grieving.
I probably have some of those details wrong, but I think there definitely is something there that is qualitatively similar to humans, although not on the same level.
More than just that: we don't know what the question is that we're trying to ask. We're pre-paradigmatic.
All of the behaviour you list, those can be emulated by an artificial neural network, the first half even by a small ANN that's mis-classifying various things in its environment — should we call such an artificial neural network "conscious"? I don't ask this as a rhetorical device to cast doubt on the conclusion, I genuinely don't know, and my point is that nobody else seems to either.
I posit that we should start with a default "this animal experiences the world the same as I do" until proven differently. Doctors used to think human babies could not feel pain. The assumption has always been "this animal is a rock and doesn't experience anything like me, God's divine creation." It was stupid when applied to babies. It is stupid when applied to animals.
Did you know that jumping spiders can spot prey, move out of line of sight, approach said pray outside that specific prey's ability to detect, and then attack? How could anything do that without a model of the world? MRIs on mice have shown that they plan and experience actions ahead of doing them. Just like when you plan to throw a ball or lift something heavy where you think through it first. Polar bears will spot walrus, go for a long ass swim (again, out of sight) and approach from behind the colony to attack. A spider and the apex bear have models of the world and their prey.
Show that the animal doesn't have a rich inner world before defaulting to "it doesn't."
As I don't know, I take the defensive position both ways for different questions.*
Just in case they have an inner world: We should be kind to animals, not eat them, not castrate them (unless their reproductive method appears to be non-consensual), not allow them to be selectively bred for human interest without regard to their own, etc.
I'd say ditto for AI, but in their case, even under the assumption that they have an inner world (which isn't at all certain!), it's not clear what "be kind" even looks like: are LLMs complex enough to have created an inner model of emotion where getting the tokens for "thanks!" has a feeling that is good? Or are all tokens equal, and the only pleasure-analog or pain-analog they ever experienced were training experiences to shift the model weights?
(I'm still going to say "please" to the LLMs even if it has no emotion: they're trained on human responses, and humans give better responses when the counterparty is polite).
> How could anything do that without a model of the world?
Is "a model of the world" (external) necessarily "a rich inner world" (internal, qualia)? If it can be proven so, then AI must be likewise.
* The case where I say that the defensive position is to say "no" is currently still hypothetical: if someone is dying and wishes to preserve their continuity of consciousness, is it sufficient to scan their brain** and simulate it?
** as per the work on Drosophila melanogaster in 2018: https://www.sciencedirect.com/science/article/pii/S009286741...
That's not how I view it. Consciousness is the result of various feedback structures in the brain, similar to how self-awareness stems from the actuator-sensor feedback loop of the interaction between the nervous system and the skeletomuscular system. Neither of those two definitions have anything to do with language ability -- and it bothers me that many people are so eager to reduce consciousness to programmed language responses only.
If you can withdraw $10,000 cash at all to dispose as you please (including for this 'trick' game) then my friend you are wealthy from the perspective of the vast majority of humans living on the planet.
And if you balk at doing this, maybe because you cannot actually withdraw that much, or maybe because it is badly needed for something else, then you are not actually capable of performing the test now, are you ?
Couldn’t someone else just give him a bunch of cash to blow on the test, to spoil the result?
Couldn’t he give away his last dollar but pretend he’s just going to another casino?
Observing someone’s behavior in Vegas is a just looking at a proxy for wealth, not the actual wealth.
If you still need a rich person to pass the test, then the test is working as intended. Person A is rich or person A is backed by a rich sponsor is not a material difference for the test. You are hinging too much on minute details of the analogy.
In the real word, your riches can be sponsored by someone else, but for whatever intelligence task we envision, if the machine is taking it then the machine is taking it.
>Couldn’t he give away his last dollar but pretend he’s just going to another casino?
Again, if you have $10,000 you can just withdraw today and give away, last dollar or not, the vast majority of people on this planet would call you wealthy. You have to understand that this is just not something most humans can actually do, even on their deathbed.
So, most people can't get $1 Trillion to build a machine that fools people into thinking it's intelligent. That's probably also not a trick that will ever be repeated.
Isn't this what most major AI companies are doing anyway?
You've invented a story where the user can pass the test by only doing this once and hinged your point on that, but that's just that - a story.
All of our tests and benchmarks account for repeatability. The machine in question has no problem replicating its results on whatever test, so it's a moot point.
Okay ? and you, presumably a human can replicate the trick of fooling me into thinking you're conscious as long as there is a sufficient supply of food to keep you running. So what's your point ? With each comment, you make less sense. Sorry to tell you, but there is no trick.
It does not think idle thoughts while it's not being asked questions. It's not ruminating over its past responses after having replied. It's just off until the next prompt.
Side note: whatever future we get where LLMs get their own food is probably not one I want a part of. I've seen the movies.
In fact, what is artificial is stopping the generation of an LLM when it reaches a 'stop token'.
A more natural barrier is the attention size, but with 2 million tokens, LLMs can think for a long time without losing any context. And you can take over with memory tools for longer horizon tasks.
What does repeatability have to do with intelligence? If I ask a 6 year old "Is 1+1=2" I don't change my estimation of their intelligence the 400th time they answer correctly.
>The machine in question has no problem replicating its results on whatever test
What machine is that? All the LLMs I have tried produce neat results on very narrow topics but fail on consistency and generality. Which seems like something you would want in a general intelligence.
If your 6 year old can only answer correctly a few times out of that 400 and you don't change your estimation of their understanding of arithmetic then, I sure hope you are not a teacher.
>What machine is that? All the LLMs I have tried produce neat results on very narrow topics but fail on consistency and generality. Which seems like something you would want in a general intelligence.
No LLM will score 80% on benchmark x today then 50% on the same 2 days later. That doesn't happen, so the convoluted setup OP had is meaningless. LLMs do not 'fail' on consistency or generality.
I'm sorry, but I find this intelectual dishonesty and moving the goal posts.
Speaks more about our inability to recognize the monumental revolution about to happen in the next decade or so.
One common kind of interaction I have with chatgpt (pro): 1. I ask for something 2. Chatgpt suggests something that doesn't actually fulfill my request 3. I tell it how its suggestion does not satisfy my request. 4. It gives me the same suggestion as before, or a similar suggestion with the same issue.
Chatgpt is pretty bad at "don't keep doing the thing I literally just asked you not to do" but most humans are pretty good at that, assuming they are reasonable and cooperative.
Most humans are terrible at that. Most humans don't study for tests, fail, and don't see the connection. Most humans will ignore rules for their safety and get injured. Most humans, when given a task at work, will half-ass it and not make progress without constant monitoring.
If you only hang out with genius SWEs in San Francisco, sure, ChatGPT isn't at AGI. But the typical person has been surpassed by ChatGPT already.
I'd go so far as to say the typical programmer has been surpassed by AI.
Here is something I do not see with reasonable humans who are cooperative: Me: "hey friend with whom I have plans to get dinner, what are you thinking of eating?" Friend: "fried chicken?" Me: "I'm vegetarian" Friend: "steak?"
Note that this is in the context of four turns of a single conversation. I don't expect people to remember stuff across conversations or to change their habits or personalities.
Your goalpost is much further out there.
Go join a dating app as a woman, put vegan in your profile, and see what restaurants people suggest. Could be interesting.
You've personally demonstrated that humans don't have to be reasonable and cooperative, but you're not at all refuting my claim.
I'm disagreeing and saying there's far more people in that bucket than you believe.
I know many people at my university that struggle to read more than two sentences at a time. They'll ask me for help on their assignments and get confused if I write a full paragraph explaining a tricky concept.
That person has a context length of two sentences and would, if encountering a word they didn't know like "vegetarian", ignore it and suggest a steak place.
These are all people in Computer Engineering. They attend a median school and picked SWE because writing buggy & boilerplate CRUD apps pays C$60k a year at a big bank.
Firstly, not studying, ignoring safety rules, or half-assing a task at work are behaviors, they don't necessarily reflect understanding or intelligence. Sometimes I get up late and have to rush in the morning, that doesn't mean I lack the intelligence to understand that time passes when I sleep.
Secondly, I don't think that most people fail to see the connection between not studying and failing a test. They might give other excuses for emotional or practical reasons, but I think you'll have a hard time finding anyone who genuinely claims that studying doesn't usually lead to better test scores. Same for ignoring safety rules or half-assing work.
I know dozens of people that have told me to my face that they don't need to attend lectures to pass a course, and then fail the course.
Coincidentally, most of my graduating class is unemployable.
It's not a lack of understanding or intelligence, but it is an attitude that is no longer necessary.
If I wanted someone to do a half-assed job at writing code until it compiles and then send the results to me for code review, I'd just pay an AI. The market niche for that person no longer exists. If you act like that at work, you won't have a job.
chatgpt.com is actually a good at or better than a typical human.
I really don't think it is on basically any measure outside of text regurgitation. It can aggregate an incredible amount of information, yes, and it can do so very quickly, but it does so in an incredibly lossy way and that is basically all it can do.It does what it was designed to do, predict text. Does it do that incredibly well, yes. Does it do anything else, no.
That isn't to say super advanced text regurgitation isn't valuable, just that its nowhere even remotely close to AGI.
I have countless examples of lawyers, hr and other public gov bodies that breach the law without knowing the consequences. I also have examples of AI giving bad advice, but it’s al better than an average human right now.
An AI could easily save them a ton of money in the fees they are paying for breaching the law.
That's not a fact, that's just cynicism mixed with sociopathy.
I hear this argument a lot from AI bros, and...y'all don't know how much you're telling on yourselves.
What you said is not a fact either. And so?
I feel every human just regurgitates words too
I know it FEELS like that's true sometimes, particularly in the corporate world, but it actually just isn't how human beings work at all.Even when people are borrowing, copying, and stealing, which is the exception, mind you, they are also carefully threading the material they are re-using into whatever it is they are trying to do, say, or make in a way that is extremely non-trivial.
Well, from my experience: a few lawyers got the law wrong but my ai did it right and the lawyer “lost” and showed how incompetent the lawyer was.
If you say most people that copy are careful you don’t know what’s an average person. And think there are 50% in the world worse than than them.
Most people lack basic logic skills.
It can appear so, as long as you don’t check too carefully. It’s impressive but still very common to find basic errors once you are out of the simplest, most common problems due to the lack of real understanding or reasoning capabilities. That leads to mistakes which most humans wouldn’t make (while sober / non-sleep deprived) and the classes of error are different because humans don’t mix that lack of understanding/reasoning/memory with the same level of polish.
This is an interesting ambiguity in the Turing test. It does not say if the examiner is familiar with the expected level of the candidate. But I think it's an unfair advantage to the machine if it can pass based on the examiner's incredulity.
If you took a digital calculator back to the 1800s, added a 30 second delay and asked the examiner to decide if a human was providing the answer to the screen or a machine, they might well conclude that it must be human as there is no known way for a machine to perform that action. The Akinator game would probably pass the test into the 1980s.
I think the only sensible interpretation of the test is one where the examiner is willing to believe that a machine could be providing a passing set of answers before the test starts. Otherwise the test difficulty varies wildly based on the examiners impression of the current technical capabilities of machines.
If you look at a calculator you will quickly find it is much better then a human in any of the operations that have been programmed into the calculator, and has been since the 1960s. Since the 1960s the operations programmed into your average calculator has increased by several orders of magnitude. The digital calculator sure is impressive, and useful, but there is no crisis. Even in the world outside computing, a bicycle can outperform an human runner easily, yet there is no mobility crisis as a result. ChatGPT is very good at predicting language. And in quite a few subject matters it may be better than your average human in predicting said language. But not nearly as good as a car is to a runner, nor even as good as a chess computer is to a grand master. But if you compare ChatGPT to an expert in the subject, the expert is much much much better then the language model. In these tasks a calculator is much more impressive.
Ask your favorite SOTA model to assume something absurd and then draw the next logical conclusions based on that. "Green is yellow and yellow is green. What color is a banana?" They may get the first question(s) right, but will trip up within a few exchanges. Might be a new question, but often they are very happy to just completely contradict their own previous answers.
You could argue that this is hitting alignment and guard-rails against misinformation.. but whatever the cause, it's a clear sign it's a machine and look, no em-dashes. Ironically it's also a failure of the turing test that arises from a failure in reasoning at a really basic level, which I would not have expected. Makes you wonder about the secret sauce for winning IMO competitions. Anyway, unlike other linguistic puzzles that attempt to baffle with ambiguous reference or similar, simple counterfactuals with something like colors are particular interesting because they would NOT trip up most ESL students or 3-5 year olds.
But it just occurred to me that you could ask "what color is a green eggplant?" for a similar result. Though you'd catch a lot of people who, like me before I looked up fruits that aren't green when unripe, assume that all unripe fruits are green. (They aren't, unripe eggplants are white and look like eggs, hence the name).
What color is a green eggplant?
ChatGPT:
A green eggplant (also called Thai eggplant or green brinjal, depending on the variety) is actually green on the outside — sometimes solid green, sometimes streaked with white or pale green.
The inside flesh is typically off-white or light cream, just like purple eggplants.
If you mean a green-colored variety (not an unripe purple one), then the skin color can range from pale mint green to deep jade, often with light striping.
Me:
How would an average human answer this question?
ChatGPT:
An average person would probably answer something simple and direct like:
“It’s green.”
Or if they’re being a bit more specific:
“It’s green on the outside and white inside.”
Most people wouldn’t overthink the variety or mention Thai vs. unripe types — they’d just describe what they see.
Experimental design comes in here and the one TT paper mentioned in this thread has instructions for people like "persuade the interrogator [you] are human". Answering that a green eggplant is green feels like humans trying to answer questions correctly and quickly, being wary of a trap. We don't know participants background knowledge but anyone that's used ChatGPT would know that ignoring the question and maybe telling an eggplant-related anecdote was a better strategy
And that "not always" is the crux of the matter, I think. You are arguing that we're not there yet, because there are lines of questioning you can apply that will trip up an LLM and demonstrate that it's not a human. And that's probably a more accurate definition of the test, because Turing predicted that by 2000 or so (he wrote "within 50 years" around 1950) chatbots would be good enough "that an average interrogator will not have more than 70% chance of making the right identification after five minutes of questioning". He was off by about two decades, but by now that's probably happened. The average interrogator probably wouldn't come up with your (good) strategy of using counterfactuals to trick the LLM, and I would argue two points: 1) that the average interrogator would indeed fail the Turing test (I've long argued that the Turing test isn't one that machines can pass, it's one that humans can fail) because they would likely stick to conventional topics on which the LLM has lots of data, and 2) that the situation where people are actually struggling to distinguish LLMs is one where they don't have an opportunity to interrogate the model: they're looking at one piece of multi-paragraph (usually multi-page) output presented to them, and having to guess whether it was produced by a human (who is therefore not cheating) or by an LLM (in which case the student is cheating because the school has a rule against it). That may not be Turing's actual test, but it's the practical "Turing test" that applies the most today.
If you understand TT to be about tricking the unwary, in what's supposed to be a trusting and non-adversarial context, and without any open-ended interaction, then it's correct to point out homework-cheating as an example. But in that case TT was solved shortly after the invention of spam. No LLMs needed, just markov models are fine.
Alan Turing was a mathematician not a psychologist, this was his attempt of doing philosophy. And while I applaud brilliant thinkers when they attempt to do philosophy (honestly we need more of that) it is better to leave it to actual philosophers to validate the quality of said philosophy. John Searle was a philosopher which specialized in questions of psychology. And in 1980 he pretty convincingly argued against the Turning test.
In the end though, it's probably about as good as any single kind of test could be, hence TFA looking to combine hundreds across several dozen categories. Language was a decent idea if you're looking for that exemplar of the "AGI-Complete" class for computational complexity, vision was at one point another guess. More than anything else I think we've figured out in recent years that it's going to be hard to find a problem-criteria that's clean and simple, much less a solution that is
Although I am scrutinizing Turin’s philosophy and, no doubt, I am personally much worse at doing philosophy then Turing, I firmly hold the belief that we will never be able to judge the intelligence (and much less consciousness) of a non-biological (and probably not even non-animal, nor even non-human) system. The reason, I think, is that these terms are inherently anthropocentric. And when we find a system that rivals human intelligence (or consciousness) we will simply redefine these terms such that the new system isn’t compatible any more. And I think that has already started, and we have done so multiple times in the past (heck we even redefined the term planet when we discovered the Kuiper belt) instead favoring terms like capability when describing non-biological behavior. And honestly I think that is for the better. Intelligence is a troubled term, it is much better to be accurate when we are describing these systems (including human individuals).
---
1: Though in honesty I will be impressed when machine learning algorithms can interoperate and generate appropriate human facial expressions. It won’t convince me of intelligence [and much less consciousness] though.
That’s not my experience at all. Unless you define “typical human” as “someone who is untrained in the task at hand and is satisfied with mediocre results.” What tasks are you thinking of?
(And, to be clear, being better than that straw man of “typical human” is such a low bar as to be useless.)
the AI bros like to talk about AGI as if it's just the next threshold for LLMs, which discounts the complexity of AGI, but also discounts their own products. we don't need an AGI to be our helpful chatbot assistant. it's fine for that to just be a helpful chatbot assistant.