Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).
Looking forward to the cope when the next big problem falls.
I hope I'm not being too blunt, but the other alternative is to "just trust me bro" the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don't think this is the way.
Simply put: For the same reason computers have not already solved all problems in mathematics.
More concretely:
Consider the Collatz conjecture. It's a very simple rule to write down:
Take some positive integer: If the number is even, divide it by two; otherwise triple it and add one. With enough repetition, do all positive integers converge to 1?
Trivial to write a program to test numbers starting at 1 and going up. We know it holds up to at least 2.36e21 (according to Wikipedia), but to prove it is true with such a program requires testing all of the infinite set of positive integers.But you may notice some things about the rules, that they suggest a subset of numbers will trivially always converge to 1, so that you don't need to even test them: any integer 2^n where n is also a positive integer.
You may find other easy wins, or ways to simplify the test, e.g. once you know the numbers up to m will converge to 1, you can terminate your loop early if you test m+1 and it ever has an intermediate value less than or equal to m. You can combine that with applying one of the rules in reverse, and know that all even numbers between m and 2m will on their first move be halved, making them smaller than m, which means you know they'll eventually converge.
But actually proving this is fully general? Nobody knows. You can't just throw arithmetic at the problem directly, you have to figure out patterns that would let you prove that it always holds, no matter what.
Or, you may find many such patterns and directly calculate some number not in any of them, to find one which doesn't converge to 1.
The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.
Talk about bad luck!
It only has 272 pages, but there is a big scandal right now about France's most prestigous literary award remobing a critically aclaimed(!) bestseller(!!) novel because it likely was "almost entirely written by AI"
https://www.theguardian.com/books/2026/sep/25/thelyson-oreli...
His publisher also said the original draft was submitted in 2019, though they’ve also become lukewarm in their backing of the author lately (he was also just accused of classic plagiarism for an unrelated short story).
So hard to say this one is settled.
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
1. "improve uniteds' plane shedule" -> the word "improve" does a lot of heavy lifting there..
2. Turing test for me is solved: "AI expert" is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can't communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol' website. In a standard llm session I don't really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.
Can’t imagine it’d be too much work to hook up and LLMs output to one.
2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn't trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM "tells".
I don't think those are problems to be solved. They are deliberate misfeatures added by the labs through RLHF to keep the models from doing the equivalent of passing a Turing test.
They don't want another GPT-4o, where people threatened to burn down the building, jump off of bridges, etc. when they unplugged it.
We can safely say that the AI labs aren't deliberately holding back. There's too many different companies making their own models, and an LLM that doesn't feel like an LLM is too lucrative of an opportunity for one of them not to defect. They are obviously going all out with this and still can't get it. I think this is why OP's benchmark is so interesting, because it seems that there are a bunch of persistent LLM defects that can't be solved definitively. They can try to squeeze it by making these defects less likely, but actually resolving what's causing them probably requires a new breakthrough in the field.
It seems very hard for me to believe, because everything motivates them to do the opposite. Can you imagine the flood of people and businesses that would come down on a model that actually sounds like a human being? The ability to impersonate a person without any tells would be a dream come true for businesses, marketers and scammers alike. They could put in zero effort and get what looks like normal human behavior in return. That is just too tempting for any of the labs to pass up. Besides, all of them have picked profit over sanity 10 times out of 10, so the argument that they drew just this single line in the sand and none of them ever crossed it on purpose seems unconvincing, especially with the sheer number of models out there, corporate and open.
At least to me, the first Dall-E 3 images were more realistic in some respects than anything that has shipped since. Likewise, I don't think it's a coincidence that the conversational capabilities of today's frontier models aren't much better than they were a year or two ago, even though other aspects and capabilities have improved massively. If they wanted a Turing test-capable model with no superficial tells, they'd have it, so I have to assume they don't want it.
(Elsewhere in the thread someone else suggests that the conversational degradation/lack of improvement might be the result of increased training on synthetic data, and that's another theory that sounds reasonable to me. If so, it's another thing they could fix if they wanted to.)
-- Robert A. Heinlein
If AI isn't achieving superhuman performance in all of these areas, I'm not sure we can actually call it "Humanity's Last Exam" -- it feels like a bit of an overextension.
In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!
If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.
If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.
We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.
Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.
Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.
At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.
Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).
In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.
What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”
To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.
>In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.
We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.
Now of course Alice has the advantage of massive parallelization, but that's not a reason to discredit her, that's a genuine advantage that she has, applicable to all problems.
To respond to a couple specific points:
> whether Bob’s contributions were required to Alice’s final step
Even if Bob's contributions were required, I don't think that's a reason to fully discredit Alice.
> Alice is capable of solving a second open math question
Given that AI has already solved dozens of open math questions, I don't think this is up for debate.