I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
1. "improve uniteds' plane shedule" -> the word "improve" does a lot of heavy lifting there..
2. Turing test for me is solved: "AI expert" is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can't communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol' website. In a standard llm session I don't really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.
2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn't trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM "tells".
I don't think those are problems to be solved. They are deliberate misfeatures added by the labs through RLHF to keep the models from doing the equivalent of passing a Turing test.
They don't want another GPT-4o, where people threatened to burn down the building, jump off of bridges, etc. when they unplugged it.
We can safely say that the AI labs aren't deliberately holding back. There's too many different companies making their own models, and an LLM that doesn't feel like an LLM is too lucrative of an opportunity for one of them not to defect. They are obviously going all out with this and still can't get it. I think this is why OP's benchmark is so interesting, because it seems that there are a bunch of persistent LLM defects that can't be solved definitively. They can try to squeeze it by making these defects less likely, but actually resolving what's causing them probably requires a new breakthrough in the field.
It seems very hard for me to believe, because everything motivates them to do the opposite. Can you imagine the flood of people and businesses that would come down on a model that actually sounds like a human being? The ability to impersonate a person without any tells would be a dream come true for businesses, marketers and scammers alike. They could put in zero effort and get what looks like normal human behavior in return. That is just too tempting for any of the labs to pass up. Besides, all of them have picked profit over sanity 10 times out of 10, so the argument that they drew just this single line in the sand and none of them ever crossed it on purpose seems unconvincing, especially with the sheer number of models out there, corporate and open.
At least to me, the first Dall-E 3 images were more realistic in some respects than anything that has shipped since. Likewise, I don't think it's a coincidence that the conversational capabilities of today's frontier models aren't much better than they were a year or two ago, even though other aspects and capabilities have improved massively. If they wanted a Turing test-capable model with no superficial tells, they'd have it, so I have to assume they don't want it.
(Elsewhere in the thread someone else suggests that the conversational degradation/lack of improvement might be the result of increased training on synthetic data, and that's another theory that sounds reasonable to me. If so, it's another thing they could fix if they wanted to.)
Can’t imagine it’d be too much work to hook up and LLMs output to one.
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
-- Robert A. Heinlein
If AI isn't achieving superhuman performance in all of these areas, I'm not sure we can actually call it "Humanity's Last Exam" -- it feels like a bit of an overextension.