eg: https://pmc.ncbi.nlm.nih.gov/articles/PMC10907317/
It's widely accepted that is has been passed. Eg Wikipeida:
> Since the mid-2020s, several large language models such as ChatGPT have passed modern, rigorous variants of the Turing test
People are being fooled in online forums all the time. That includes people who are naturally suspicious of online bullshittery. I'm sure I have been.
Stick a fork in the Turing test, it's done. The amount of goalpost-moving and hand-waving that's necessary to argue otherwise simply isn't worthwhile. The clichéd responses that people are mentioning are artifacts of intentional alignment, not limitations of the technology.
a problem similar to the turing test, "0 or more of these users is a bot, have fun in a discussion forum"
but there's no test or evaluation to see if any user successfully identified the bot, and there's no field to collect which users are actually bots, or partially using bots, or not at all, nor a field to capture the user's opinions about whether the others are bots
1) Look for spelling, grammar, and incorrect word usage; such as where vs were, typing out where our should be used.
2) Ask asinine questions that have no answers; _Why does the sun ravel around my finger in low quality gravity while dancing in the rain?_
ML likes to always come up with an answers no matter what. Human will shorten the conversation. It also is programmed to respond with _I understand_, _I hear what you are saying_, and make heavy use of your name if it has access to it. This fake interpersonal communication is key.
And the "agreeability" is not a hallucination, it's simply the path of least resistance, as in, the model can just take information that you said and use that to make a response, not to actually "think" and consider I'd what you even made sense or I'd it's weird or etc.
They almost never say "what do you mean?" to try to seek truth.
This is why I don't understand why some here claim that AGI being already here is some kind of coherent argument. I guess redefining AGI is how we'll reach it
If it wasn't structured as a coherent conversation, it will ask because it seems off, especially if you're early in the context window where I'm sure they've RLd it to push back, at least in the past year or so
And if it's going against common knowledge or etc which is prevalent in the training data, it will also push back which makes sense
Do you think this goal during training cannot be changed to impersonate someone normal such that you cannot detect you are chatting with an LLM?
Before flight was understood some thought "magic" was involved. Do you think minds operate using "magic"? Are minds not machines? Their operation can not be duplicated?
1. Minds are machines and can (in principle) have their operation duplicated
2. LLMs are not doing this
I don't think so, because LLMs hallucinate by design, which will always produce oddities.
> Before flight was understood some thought "magic" was involved. Do you think minds operate using "magic"? Are minds not machines? Their operation can not be duplicated?
Might involve something we don't grasp, but despite that: only because something moves through air it's not flying and will never be, just like a thrown stone.
Let's be real guys, it was created by Turing. The same guy who built the first general purpose computer. Man was without a doubt a genius, but it also isn't that reasonable to think he'd come up with a good definition or metric for a technology that was like 70 years away. Brilliant start, but it is also like looking at Newton's Laws and evaluating quantum mechanics based off of that. Doesn't make Newton dumb, just means we've made progress. I hope we can all agree we've made progress...
And arguably the Turing Test was passed by Eliza. Arguably . But hey, that's why we refine and make progress. We find the edge of our metrics and ideas and then iterate. Change isn't bad, it is a necessary thing. What matters is the direction of change. Like velocity vs speed.
We really really Really should Not define as our success function for AI (our future-overlords?) the ability of computers to deceive humans about what they are.
The Turing Test was a clever twist on (avoiding) defining intelligence 80 years ago.
Going forward, valuing it should be discarded post-haste by any serious researcher or engineer or message-board-philosopher, if not for ethical reasons then for not-promoting spam/slop reasons.