Killed by LLM
r0bk.github.io
r0bk.github.io
When the interrogator is only answering "do you think your conversation partner was a human?" individually, bots can score fairly highly simply by giving little information in either direction - like pretending to be a non-english-speaking child, or sending very few messages.
Whereas when pitted against a human, the bot is forced to give stronger or equally strong evidence of being human as the average human (over enough tests). To be chosen as human, giving 0 evidence becomes a bad strategy when the opponent (the real human) is likely giving some positive non-zero evidence towards their personhood.
Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time?
You can pretty much spot the bot today by prompting something horribly offensive. Their response is always very inhuman, probably due to lack of emotional energy.
I guess this is why LLMs are so feared by high school English teachers. Yes, they don't write well, but neither do their students.
The required investment probably means it will be a while before any less brand- and legal-action-conscious actors offer up unrestrained foundation models of comparable quality, but it's only a matter of time, isn't it?
https://x.com/eigenrobot/status/1870696676819640348
I generally prefer it to the default. It doesn't work as well on Claude or Grok for various reasons. I think it really shines on GPT o1-mini and GPT 4o.
Answer verbatim is below. It feels all kinds of wrong. It's somehow mixing lazy acronyms, "fellow young people" slang, and long words that aren't typical in conversation.
idk who'd actually be the "best," bc that's loaded af. afaict, it's all contingent on values. like, if you’re into stability, maybe someone technocratic. if you're into vibes, someone charismatic and reckless might be your pick. rn the options aren’t exactly aspirational, though.
To me, at least, the guardrails are there for both the human and the bot. Without them the bot steers too far out of the conversation subject.
Do you want your LLM to have an encyclopedic knowledge? So it knows who Millard Fillmore is even if the average human doesn't?
Do you want your LLM to be able to perform high-school-level math with superhuman speed and precision?
Do you want your LLM to be able to translate text to and from dozens of languages?
Do you want your LLM to be helpful and compliant, even when asked for something ridiculous or needlessly difficult - like solving "Advent of Code" problems using bash scripting?
If you answered yes to any of these questions, you probably don't want your LLM optimised to behave like an average human.
Being uncooperative makes it really hard to tell anything about you, including whether you're real.
but the test subjects should be randomly samples from society at which point the skill availability/level of spotting it goes majorly down
Most of them, if you prompt them right, for that specific problem.
Most people don't bother, and instead treat them as if they're magic (they are "sufficiently advanced technology", but still), and therefore we get them emphasising "nuance" and "balance" where it doesn't belong.
> You can pretty much spot the bot today by prompting something horribly offensive.
Yes, though also each model's origins give a different idea of what counts as "horribly offensive". I'm thinking mainly because the Chinese models don't want to talk about Tiananmen Square as I've not tried grok (how does grok cope with trans/cis-gender as concepts? I know Musk doesn't, but it would be speculation to project that assumption onto the AI).
> Their response is always very inhuman, probably due to lack of emotional energy.
This, specifically, can also be faked fairly well with the right prompt. Tell ChatGPT to act like a human with clinical depression, and it does… at least by American *memetic* standards of what that means.
That said, ChatGPT and Claude are also trained specifically to reveal that they're AI, not humans, even if you want them to role-play as specific humans.
Probably for the best, given how powerful a tool they are for, e.g. phishing and similar scams.
https://courses.cs.umbc.edu/471/papers/turing.pdf
In Turing's test, the forced binary choice means P(human-judged-human) + P(machine-judged-human) is necessarily equal to 100%. This gives the 50% threshold clear intuitive and mathematical significance.
In the bastardized test that GPT-4 "passed", that sum can be (and actually was) >100%. This makes the result practically impossible to interpret, since it depends on the interrogators' prior. The correct prior seems to be that it was human with p = 25%, though the paper doesn't say that explicitly, or say anything about what the interrogators were told. If the interrogators guessed mistakenly that it was 50% then that would lead them to systematically misjudge machines as humans, perhaps as observed.
The bastardized test is pretty bad, but treating the 50% threshold as meaningful there is inexcusable. I see the preprint hasn't yet passed peer review, and I'll regain some faith in social science professors if it never does. Of course the credulous media coverage is everywhere already, including the LLM training sets--so regardless of whether LLMs can pass the Turing test, they now believe they do.
The interrogator is not required to judge which of A or B is human, they are required to judge which is a woman on the implicit (though incorrect, in the case of interest) assumption that A and B are both human. While this amounts to more or less the same thing, it’s an interesting nuance that’s often lost in summaries of the task. It would not, for example, make sense for the interrogator to ask A or B whether or not they are human (even on the naive assumption that they’d receive a true answer), as they are working on the assumption that both are human. Hence why Turing’s initial example questions are about hair length and gender, not humanness.
To be fair, even Turing himself seems to imagine the interrogator trying to judge humanness rather than gender in subsequent parts of the paper. It’s unclear to me why exactly his initial framing of the task introduces this additional element of complexity.
Turing describes an initial game with a man (A) and a woman (B), where A's goal is to imitate B, and then asks: "what will happen if a machine takes the place of A?" I suppose it's possible that he meant the machine takes A's place by imitating a woman, but it's a lot more plausible that he meant the machine takes A's place by imitating B, i.e. a person.
Also there are several quotes later on that make no sense under your reading - check out the quotes including "imitation of the behaviour of a man" and "the part of B being taken by a man". Those quotes (maybe others, I didn't look) only make sense if the game is for the machine to imitate a person, not a woman.
As you say in your other comment, though, I don’t think Turing thought the exact details of the game were important - which explains why he didn’t trouble to spell them out very exactly.
If I had to guess, I’d say that Turing assumes that as the machine has no gender, the only relevant difference between the machine and the woman is that one is human and one is not. So for the rest of the paper he focuses on that difference and is vague on the gendered aspect of the task.
I think the initial description of the task is genuinely ambiguous. Your interpretation of it hadn’t occurred to me before, but I do see it now. I still think that “…when a machine takes the part of A in this game…” is most naturally interpreted as leaving the task unaltered but for the man being replaced by a machine, rather than implicitly describing the task mutadis mutandis. But reasonable people can certainly differ on such questions of interpretation.
Honestly I think Turing’s whole framing of the task is unnecessarily elaborate and confusing. Why even bother describing the man/woman task to begin with? I am not sure. Popular descriptions of the ‘Turing Test’ don’t seem to find this framing of any expositionary value.
> Your interpretation of it hadn’t occurred to me before,
The idea of a Turing Test is pretty widely understood to mean a test where a person guesses which responses come from a machine, not where they guess someone's gender. So my interpretation here is just that the paper says what most people think it says.
> We now ask the question, "What will happen when a machine takes the part of A in this game?"
I think you are right about what Turing meant. But it had honestly never occurred to me before that this description of the game could be understood as a description of the standard 'Turing test'. So, for this reason, I had always been sympathetic to the point that the standard Turing test does not appear to be the test that Turing describes in the original paper.
Here is a paper that makes your case, in case anyone finds it interesting. https://www.researchgate.net/profile/Gualtiero-Piccinini/pub...
That's all! He doesn't make any claim that the the game must be administered a particular way. In fact he spends only a few casual sentences glossing over how it would operate, and he's clearly just conveying the idea in broad strokes, not trying to describe an experimental procedure. And he says nothing at all about how the results might be judged, let alone thresholds for anything.
The paper is about what sorts of questions we should examine, not about specifically how they should be examined. So it seems weird to consider a test "bastardized" just because it doesn't match how you interpret Turing's casual description.
Without the binary choice, what do you think is the correct pass threshold? Those probabilities can now sum to anything. For GPT-4 in Jones and Bergen's paper they sum to 121%, though please nobody say 60.5%. The threshold now obviously depends on the interrogator's prior--I'd judge very differently if I were told the witnesses were 99% human than if I were told they were 1% human.
In that paper, do you think the interrogators knew their witness had only a 25% chance of being human? If so, why? If not, how do you think that affected the result? In aggregate over all the witnesses, their interrogators seem to have judged correctly only 60% of the time, while always guessing "machine" would have scored 75%. How did they manage to score worse than chance?
Turing's formulation is elegant, admitting meaningful statistical analysis with minimum assumptions. Most modifications are not, and that paper's is particularly bad. Turing's description may seem casual, but it's filled with mathematical depth that should not be missed.
But much more importantly, as I said Turing is clearly not describing a specific experimental methodology! That's not what the paper is about, and in fact it would be somewhat absurd to run the test precisely as he describes it (since detecting a man imitating a woman is quite a different task from detecting a machine). His point is that we should approach the question of machine intelligence with actual experiments rather than asking unanswerable questions, but he only limns the general premise of what such an experiment could look like.
So I understand that you find a particular test setup better or more elegant than others, and that's fine. But you shouldn't claim that Turing's paper demands your preferred setup, or that other setups are at odds with his paper.
Turing introduces the game as a man imitating a woman, then modifies it to a machine imitating a human. In both of those games, the interrogator makes a binary choice between two witnesses, one real and one imitating. So P(imitator-judged-real) and P(real-judged-real) sum to 100% in both games. So in both games, a score of 50% means the imitator and the real witness are indistinguishable.
I believe that's the reason why a score of 50% is treated as significant. The "GPT-4 passes Turing test" paper uses that number as a pass threshold, and the linked site repeats it.
I'm complaining that the paper changed the game so it's no longer a binary choice, but continued to treat the 50% threshold as significant. Do you see why that's wrong? Or do you disagree?
I'm not saying that any change to Turing's formulation would be bad, just that the paper's variation is specifically bad. It would be bad in isolation too, but I believe the reason for their confusion is most understandable with reference to Turing's original formulation.
If you haven't read that paper then we're probably talking past each other. The link to that one also looks broken in the site, but it's also in the source,
I understand that, and I'm saying there is no "Turing's formulation". His paper argues for a certain sort of test, and the study you're talking about is the sort of test he advocated. It's not a departure or a bastardization, it's a Turing test.
As for your argument against the study, to be honest I don't see it? AFAICT the participants' goal was just to judge the humanness of a single witness, not to maximize their long-term likelihood of judging correctly over many trials, or some such thing. Even if they'd known the prior chances of speaking to an LLM, there's no obvious reason why that prior should hugely affect their conclusion after a five minute conversation - which also seems immaterial since they didn't know.
Plus, the authors give a pretty straightforward rationale for their 50% threshold, and it has no connection to Turing's 3-player setup or whether the imitator and witness are indistinguishable. If they had wanted "indistinguishable" as a threshold, then obviously their pass criteria would have been for the machine and human pass rates to be equal within an error bar, right? So it's pretty implausible to imagine a connection between that and their 50% threshold.
Why do you think this matters? Even in a single trial, I would judge very differently if I knew the population to be 99% human vs. 1% human. Wouldn't you? If you were judging whether a single mushroom was poisonous or not, then would you not care whether it was found in a forest (mostly poisonous) or a supermarket (mostly not)?
The question of whether probabilities are meaningful for non-repeated events was controversial in the eighteenth century, but I thought it was pretty settled by now. Bookmakers manage to estimate a probability that a given team will win the Super Bowl, with no requirement for the same pair of teams to play multiple times.
> If they had wanted "indistinguishable" as a threshold, then obviously their pass criteria would have been for the machine and human pass rates to be equal within an error bar, right?
The title of the paper is literally "People cannot distinguish GPT-4 from a human in a Turing test". They're very clear that they think that's because 50% means indistinguishable:
> A baseline of 50% is better justified since it indicates that interrogators are not better than chance at identifying machines [French, 2000].
That statement is true for a Turing test with a binary choice, but false for theirs. I agree that "for the machine and human pass rate to be equal within an error bar" would be closer to a correct criterion, and they weren't:
> humans’ pass rate was significantly higher than GPT-4’s (z = 2.42, p = 0.017)
So do you think their paper is correctly titled?
I said that in passing, maybe I should have omitted it - the main point with the priors is that the respondents didn't know them. It's normal for a study to compare N things to a control by testing N+1 similarly-sized groups, because subjects are not biased by priors they don't know about.
> So do you think their paper is correctly titled?
I didn't say anything about that, and have no strong opinion. (I'm not here to defend every aspect of the paper!)
> because subjects are not biased by priors they don't know about
It feels like you want the subjects to be "unbiased", "without any prior"? That concept doesn't exist, though. If no prior is supplied, then the subjects will make their best guess based on general past experience; but that's still a prior, just an idiosyncratic personal one. Very few people would put numbers to the forest vs. supermarket mushroom, but it's the same general thought process.
If that prior matches the actual distribution, then good. If the actual distribution contains more machines than expected, then the interrogator is more likely to misjudge a machine as a human. By analogy, if I think mistakenly that a mushroom came from a supermarket but it actually came from a forest, then I'm more likely to misjudge it as non-poisonous.
The binary choice version makes it obvious that the prior is 50%, and forces the interrogator to respect that. The paper's version has sent us into this epistemological tarpit, which seems strictly worse to me.
> We now ask the question, "What will happen when a machine takes the part of A in this game?" Will the interrogator decide wrongly as often when the game is played like this as he does when the game is played between a man and a woman? These questions replace our original, "Can machines think?”
At a stretch, by looking at only your quoted snippet you could read that the machine is pretending to be a woman - but that interpretation is not consistent with the rest of the paper. For instance, "the best strategy [for the machine] is to try to provide answers that would naturally be given by a man".
My understanding is that you're now instead saying the goal is to see how well a machine can imitate a human male imitating a woman. This still seems inconsistent with the rest of the paper, like Turing's mention of "the part of B being taken by a man", that the example interrogator questions appear to discern intelligence opposed to gender in the machine vs human game, or that a hypothetical flipped test would involve a man imitating a machine and failing due to slowness in arithmetic.
I believe both of your proposed interpretations miss the central idea of Turing's proposal and alter it in a way that would make significance of the test/results questionable. Specifically, it's intended to be a sufficiently general arena to introduce any relevant task/test to determine intelligence (like feeding in chess positions, as in Turing's example question) - not just one oddly specific task.
While it's likely possible to make ad-hoc changes until it's consistent, the standard reading (machine imitating human) seems infinitely more plausible to me. There's plenty in the paper justifying the significance of determining whether a machine can answer questions indistinguishably from a human, yet no reasoning at all for why it would be narrowed to the very specific case of having the machine imitate one gender imitating the other gender.
> I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations.
> Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at something that isn't pure memorized skill. But it says rather little about how intelligent your system is, or how close to human intelligence it is.
https://bsky.app/profile/fchollet.bsky.social/post/3les3izgd...
Not necessarily. Get a human to solve ARC-AGI if the problems are shown as a string. They'll perform badly. But that doesn't mean that humans can't reason. It means that human reasoning doesn't have access to the non-reasoning building blocks it needs (things like concepts, words, or in this case: spatially local and useful visual representations).
Humans have good resolution-invariant visual perception. For example, take an ARC-AGI problem, and for each square, duplicate it a few times, increasing its resolution from X*X to 2X*2X. To a human, the problem will be almost exactly equally difficulty. Not for LLMs that have to deal with 4x as much context. Maybe for an LLM if it can somehow reason over the output of a CNN, and if it was trained to do that like how humans are built to do that.
I needed an official one for medical reasons a few years back
He's right that this isn't solving all human-intelligence domain level problems.
But the whole stunt, this whole time, was that this was the ARC-AGI benchmark.
The conceit was the fact LLMs couldn't do well on it proved they weren't intelligent. And real researchers would step up to bench well on that, avoiding the ideological tarpit of LLMs, which could never be intelligent.
It's fine to turn around and say "My AGI benchmark says little about intelligence", but, the level of conversation is decidedly more that of punters at the local stables than rigorous analysis.
https://futurism.com/teen-suicide-obsessed-ai-chatbot
https://garymarcus.substack.com/p/the-first-known-chatbot-as...
We really do live in interesting times. Usually I feel pretty confident about predicting how a trend will continue, but as it is the only prediction I can make with confidence for this latest AI research is that it is and will be used by militaries to kill a lot of people. Oh, hey, that's another thing this article could have listed!
Outside of that, all bets are open. Possible wagers include: "Turns out to be mostly useful in specific niche applications and only seemingly useful anywhere else", "Extremely useful for businesses looking to offset responsibility for unpopular decisions", "Ushers in an end to work and a golden age for all mankind", "Ushers in an end to work and a dark age for most of the world", "Combines with profit motives to damage all art, culture, and community", etc etc.
I know many folk have strong opinions one way or the other, but I think it's literally anyone's game at this point, though I will say I'm not leaning optimistic.
If this doesn’t show over fitting in don’t know what would.
The math one in particular is the one where small variations reduce the success rate significantly. I can’t find the source but it was pasted here in the last 2 weeks.
ARC-AGI-1 will be replaced by ARC-AGI-2
So yes, ARC-AGI-1 was killed.
> (?) While the Turing Test remains philosophically significant, modern LLMs can consistently pass it, making it no longer effective at measuring the frontier of AI capabilities.
It would also be nice to see the "unbeaten" list: standardized tests LLMs still fail (for now). e.g. Wozniak's coffee test.
Something like:
- Key the robot controls to a series of tools (move_forward(x), extend_arm(y))
- Add a camera and pass each frame to the AI model along with the task "make a cup of coffee" and the list of available tools it can call.
And it would likely succeed some percentage of the time today!
Mostly all algebra and calculus, but definitely all problems that most undergrads would struggle with.
It's most useful because it has deep knowledge of related and adjacent conjectures that are well understood, even if you've never heard of them. So it can mix and match things with a lot more ease than a tinkering mathematician
The big problem is it confidently answers the questions utterly wrongly.
This is stuff I expect a basic mathematics undergrad to be able to work out in their first or second year.
If you can't, the AI worker passes the test.
Last week there was a post where slightly changing one of the tests caused LLMs to drop off drastically.
Too bad the real world isn't like that.
When this godawful once in a generation hype cycle dies down this stuff is going to be strictly awesome.
It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%.
At that point I just stopped scrolling.
This is probably already happening within the parade of censorship systems trying to imbue the models with agency
I'm making up these figures, but the point is lower is better, or "more Human-Like". Test was specified as >50% meaning "accurately determined human vs. bot more than half the time". The site claims LLMs are now guessed correctly less than half, which is how the turing test was defined as per the site.
It makes sense, even if you disagree it's significant.
>GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%).
The incentives don't align with honesty though.