Vote on which of Hacker News' challenges for AI have been met
stoppels.ch
stoppels.ch
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!
If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.
If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.
We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.
Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.
Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.
At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.
Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).
In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.
What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”
To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.
>In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.
We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
1. "improve uniteds' plane shedule" -> the word "improve" does a lot of heavy lifting there..
2. Turing test for me is solved: "AI expert" is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can't communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol' website. In a standard llm session I don't really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.
Can’t imagine it’d be too much work to hook up and LLMs output to one.
2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn't trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM "tells".
I don't think those are problems to be solved. They are deliberate misfeatures added by the labs through RLHF to keep the models from doing the equivalent of passing a Turing test.
They don't want another GPT-4o, where people threatened to burn down the building, jump off of bridges, etc. when they unplugged it.
We can safely say that the AI labs aren't deliberately holding back. There's too many different companies making their own models, and an LLM that doesn't feel like an LLM is too lucrative of an opportunity for one of them not to defect. They are obviously going all out with this and still can't get it. I think this is why OP's benchmark is so interesting, because it seems that there are a bunch of persistent LLM defects that can't be solved definitively. They can try to squeeze it by making these defects less likely, but actually resolving what's causing them probably requires a new breakthrough in the field.
It seems very hard for me to believe, because everything motivates them to do the opposite. Can you imagine the flood of people and businesses that would come down on a model that actually sounds like a human being? The ability to impersonate a person without any tells would be a dream come true for businesses, marketers and scammers alike. They could put in zero effort and get what looks like normal human behavior in return. That is just too tempting for any of the labs to pass up. Besides, all of them have picked profit over sanity 10 times out of 10, so the argument that they drew just this single line in the sand and none of them ever crossed it on purpose seems unconvincing, especially with the sheer number of models out there, corporate and open.
At least to me, the first Dall-E 3 images were more realistic in some respects than anything that has shipped since. Likewise, I don't think it's a coincidence that the conversational capabilities of today's frontier models aren't much better than they were a year or two ago, even though other aspects and capabilities have improved massively. If they wanted a Turing test-capable model with no superficial tells, they'd have it, so I have to assume they don't want it.
(Elsewhere in the thread someone else suggests that the conversational degradation/lack of improvement might be the result of increased training on synthetic data, and that's another theory that sounds reasonable to me. If so, it's another thing they could fix if they wanted to.)
-- Robert A. Heinlein
If AI isn't achieving superhuman performance in all of these areas, I'm not sure we can actually call it "Humanity's Last Exam" -- it feels like a bit of an overextension.
I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).
Looking forward to the cope when the next big problem falls.
I hope I'm not being too blunt, but the other alternative is to "just trust me bro" the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don't think this is the way.
The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.
Talk about bad luck!
It only has 272 pages, but there is a big scandal right now about France's most prestigous literary award remobing a critically aclaimed(!) bestseller(!!) novel because it likely was "almost entirely written by AI"
https://www.theguardian.com/books/2026/sep/25/thelyson-oreli...
His publisher also said the original draft was submitted in 2019, though they’ve also become lukewarm in their backing of the author lately (he was also just accused of classic plagiarism for an unrelated short story).
So hard to say this one is settled.
I want AI to replace me in my chores, not in my enjoyable activities.
They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
They not only do that —saving you so many worries— but then you get to be medieval about it and say: no, I challenge the tax authority to a duel.
The Earned Income Tax Credit is one of the most important transfers embedded in the tax code. It provides about $70B to low-income workers every year.
Take a look at the eligibility criteria:
https://www.irs.gov/credits-deductions/individuals/earned-in...
You can claim the EITC if you are married, not filing a joint return, had a qualifying child who lived with you for more than half of the tax year and either of the following apply.
- You lived apart from your spouse for the last 6 months of tax year, or
- You were legally separated according to your state law under a written separation agreement, or a decree of separate maintenance and you didn't live in the same household as your spouse at the end of the tax year.
…
You can claim the head of household filing status if you're not married, had a qualifying child living with you more than half the year, and you paid more than half the costs of keeping up your home.
No, the government can’t derive this all “automatically”.
The government absolutely should know enough to apply this reasonably accurately, and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home') should simply not be part of tax law, or should be a checkbox when registering your address of "I am head of household" so the government knows.
The US government can't derive all this, but they should update it so they can.
Heck, with the NSA's hooks into banks and cell phone companies, there's a non-zero chance they actually could derive this all now if they wanted.
The American mind cannot handle such invasions of privacy.
In most situations where you paid significantly more than half deriving this information should be easy.
Obviously edge cases where someone paid 55% would not be easily derived, but probably most people wouldn't care too much about that.
“No, they can’t.”
“But they should be able to!”
Personally, I don’t want a battered wife fleeing her husband to have to file a form with the IRS for tax purposes, but mostly I’m just tired of this kind of hn thread.
If you're working a W2 job or contracting for a big company, the government is probably aware of your income and that's where we should be focused on automating tax returns but if your income comes from regular people, they likely have no idea.
Unless you turn the entire national security apparatus around to surveil citizens' income, I don't think it's ever possible for the government to know everything.
How it should work is that the government should show you everything they already know and prompt you to fill out anything they don't and then you submit it. The majority of people would immediately benefit from this. They don't do this for a few reasons: number one being the tax prep lobby that wants to make doing your taxes a burden such that even regular people feel the need to spend money or have their personal financial information gathered by private companies. The second reason is that if the government is already aware of how much you've made, chances are you've already paid your taxes and would either be looking at a refund or other tax credits, the government is effectively spending money to make it easier for you to take money away from it.
You can just ask the IRS for your "tax transcripts" and do the data entry. People don't do this because it leaves tons of money on the table.
Now you might say that a tax game that rewards skilled play is bad. But are you sure about that? Because everyone with influence over the system (who all happen to be skilled players) happens to be quite fond of the game, observably speaking.
In my country I literally got a letter from the government if I could please file my taxes, because they believe the automatic deductions are too high so I am likely owed a back payment. In a previous year (I am not very good with non-timing-critical paperwork) they even called me on my personal cell to inform me about something similar.
The slogan of our tax collection agency is "we can't make it more enjoyable, we can only make it easier". If you're a regular employee and your taxes - deductibles or not - are a hassle, then that's 100% a political choice.
It's worse than that. Nobody is actually sympathizing with Mark Zuckerberg, the reason US taxes are complicated is so Congress can confuse people about how they really work.
One of the examples in this thread is the complexity of the EITC. The EITC is one of the most "efficient" tax credits -- it's much better to give people money (or not take it from them to begin with) than having systems of complicated vouchers for some specific thing or another with a bunch of strings attached and bureaucratic paperwork. But that same efficiency means that if it was working as it was supposed do, most of the lower middle class would be getting a piece of it. So why is it so complicated?
First, to hide this part of it: The end of the phase out range for an individual with no dependents, i.e. the most money they can make and still receive it, is $19,540. Which is to say, is less than what you make by working a full time job at the minimum wage in 30 states. None of those people get any of it. Neither do the people who are unemployed, because you also don't get it if you don't have any earned income. The primary way to receive any of it is to have a child -- but not too many of them, because you get no additional credit for having more than three. Moreover, the caps for married couples are only a little higher than they are for single people, so a married couple in California with two incomes gets nothing even if they have three children because the credit is fully phased out by making their state's minimum wage.
And second, because if they make it complicated enough then even some of the people who are eligible for it won't notice.
The combination of these is the reason a credit that should be going to a significant percentage of the population is somehow only ~1% of the federal budget. Because then Congress gets to pretend to be helping people while minimizing the amount of helping people they actually do.
But realize that changing it is a policy debate not a “this should be going to a significant percentage of the population [but isn’t because of tax code complexity]” debate point.
(I happen to agree with you that this type of credit is highly beneficial.)
But that's a different matter than pretending to have it while disguising the fact that it's structured to make sure hardly anybody actually gets it beneath a layer of impenetrable complexity.
If EITC increases are capped at 3 children, it's because the policy that was passed made that choice. If we want to debate whether that cutoff should be 2 or 4, that's a valid policy debate.
If EITC credits are capped at just under $20K of income, that's perhaps because the policy targeted the very lowest income working poor and by policy did not allocate funds for the next tiers of income earners. Again, that's a policy debate to be had.
The issue is that Congress is making the details of the second type of policy complicated on purpose to conceal the fact that they're not really doing the first one, while making it seem like they are to mollify the people who want something like that.
That wasn't the framing of the EITC when introduced. The framing when introduced was in opposition to basic income proposed by Nixon and instead to offer a temporary stimulus via refundable credits to the working poor that would roughly offset their social security taxation, to ensure that "work pays more than welfare," and to try to stimulate the economy to exit the severe recession of 1974.
It had nothing to do with supporting working people in the 4th and 5th decile of income. Getting that policy agreed to is a policy debate that hasn't been had.
I have a relatively straightforward tax return and still found significant mistakes on 3 out of the last 4 years of returns filed by CPAs.
It beats the tax advisor we hired two years ago, who made a significant mistake.
The German tax system is one of the most complicated one in the world.
It also didn’t help that for a very long time, simple adaptive filters and basic neural networks were branded as “artificial intelligence” despite having very basic capabilities and little mystery on how/why they worked in their narrow use case.
I didn’t believe the recent hype for a very long time. But having tried out the latest models, I’m kinda shocked. In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less. Coverity would have cost dearly and generated far more false positives than substance.
They of course operate faster than humans, but this is hyperbole. Unless your senior engineer sucks.
If you take "arbitrary" seriously, we are still very far away from it.
You're maybe thinking that we can build and deploy arbitrary "kinds" of application. Being able to "build and deploy arbitrary applications" would mean I could ask for any scale or complexity in my application.
edit: I guess I misread this. I thought you were saying "well, AI isn't smart until it can solve any arbitrary problem in the whole world"
EDIT: to forestall further back and forth, I don't think there'd be any controversy if the problem statement said "common applications"
I would claim "common applications" is also controversial because what is a "common application" depends insanely on the area in which you work. Even if you exclude some highly advanced scientific applications (because you don't consider these to be common), in many industrial sectors there exist applications that have grown over multiple decades, and which encode an insane amount of knowledge about the respective sector and its workflows; this is a central reason why these applications are so hard to replace.
And just generally, anything that requires the LLM to understand something LLMs don't understand, it's going to fail. I'd hesitate to give an exact example without trying Astra/Fable but I'm sure they exist.
Right now, they can't build anything without help that isn't buggy in unintelligible ways. If you push the thing feature by feature, have a lot of tests and a lot of instrumenting, and you check that it isn't cheating or lying after every step, you can get a lot of work done.
- Even if we have P=NP, it is not even known whether there will ever exist a "practically useful/fast" algorithm for solving NP-complete decision problems.
- If P != NP, it is perfectly reasonable that there exists an algorithm that is for all practical purposes "fast" algorithm (say, some O(n^{log log log log log log log n}) algorithm with a very small hidden constant) for solving NP-complete decision problems.
- It is entirely possible that average case complexity is the much more important complexity measure than the worst-time measure that is used for defining the P and NP complexity classes. If you are into this kind of questions, you might enjoy the article
Fifty Years of P vs. NP and the Possibility of the Impossible
https://cacm.acm.org/research/fifty-years-of-p-vs-np-and-the...
and the paper
R. Impagliazzo
A personal view of average-case complexity
https://www.karlin.mff.cuni.cz/~krajicek/ri5svetu.pdf
--
Also, in the realm of complexity theory, the question of P vs NP is just a small puzzle piece.
Just to give one example: isn't the question of P vs PSPACE much more exciting. If you believe in P != NP, P != PSPACE is a trivial corollary. But we can't even exclude P = PSPACE.
Seriously: there exist so many complexity classes (some of high potential practical importance) for which we often basically know nothing except for the trivial inclusions.
I think there's a whole lot of survivorship bias going into our perception of how effective AI is at building arbitrary applications without significant hand-holding.
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
They are the kind of problems that, if your teacher was anything like mine, were usually skipped in order to keep the slower students from bogging down the class as a whole. 0.996 (MiMo-V2.5) is substantially better than what the vast majority of humans would do.
If you limited the question to adding arbitrary pairs of numbers of reasonable size, I imagine quite a few models could get to 1.000.
"there is a nand gate A connected to gate B through these wires, connected to another gate C, etc. what is the output value of gate C if i place a 1 at this gate"
Sort of like doing math "in their heads" (i.e. your "when adding numbers"), they would get it very right for simple/small cases (although the answer could have been in their training), then some LLMs would get it for harder cases (which were clearly not in their training), and all would fail at some point. This was all without any tool calling.
After a year of thinking about it, I made an eval [0] with ever-complexifying nand circuits - like, truly, bananas circuits [1] - and some models, do, in effect (through chain of thought? mostly?), get the right answer. ((what's nice is that you can always make a circuit at the very edge of what all models can correctly solve))
Tool-calling 100000% solves this problem for sure (evaluating a nand gate is trivial). But if you even prompt an llm to do math like a 5/6th grader (i.e. do it digit by digit, carry the 1, etc.) - I am quite certain most llms can, in fact, add numbers.
But yeah. These piles of weights are fascinating in how they seem flawed one day ("how many r's") and magical at once.
[0] https://lockstep.greg.technology
[1] https://lockstep.greg.technology/c/?id=rand_s4161_g160_d8
Decades before the modern AI push I was marveling at the distinction between the sheer overwhelming computational power of the human brain, if measured from the perspective of how much math it is doing under the hood, and its utter ineptness at basic arithmetic compared to the tools we can build. The earliest, klunkiest, most garbage mechanical adding machines we ever produced, long before we improved them by literally over a dozen orders of magnitude, were still already way better than we are at basic arithmetic.
There is something profound I still have not fully put my finger on in how basic arithmetic is so easy for a machine, yet the decisions we routinely make with our neural nets has been the laborious effort of decades with us still not arriving yet even with the trillions now poured into AI for machines. And vice versa. Even that practice I alluded to that allows you to train yourself to do this task on demand easily would incorporate mathematical advancements in representations that took our species thousands of years to come up with, rather than being something you get "for free" just for being smart. Our brains casually run an entire human body through an unbelievably complicated external universe, yet struggle with basic arithmetic.
There's so many places where we have one architecture that's a bit better than another at one thing, and a bit worse at another, but in the end they can both do the job. Like, if we had to do all our programming in immutable languages and run all our imperative code through an O(n log n) worst-case penalty for immutable languages emulating imperative RAM, we'd survive just fine over all. But between neural architectures and conventional arithmetic on dedicated silicon is this dozen+ order of magnitude difference on tasks. It's a pretty wild disparity. Some of the reason is somewhat obvious, I don't want to make it sound like I'm completely mystified... I just think there's probably, somewhere, an even more profound way to see it than the obvious differences that says something more powerful about the limits of computation than I've seen anyone say. Which is not to say somewhere out there someone already has had the idea I'm grasping for and written it in some brilliant paper or something. I'm just saying I haven't seen it.
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
"AI still hasn't passed a proper Turing test. A proper Turing test is 1 hour or more. Has any AI passed it?"
How do you answer that? "'Yes', AI has NOT passed a proper Turing test", or "'Yes', AI has passed a proper Turing test".
My comment doesn't actually make any prediction to judge, it's just an observation and an argument.
I thought the different variations on "could AI pass the Turing test?" were interesting in this regard. Surely any frontier LLM could pass a Turing test for some amount of time, and that's been the case for at least a year now. But I don't think we're anywhere close to a model that could pass an "adversarial" Turing test for an extended period of time.
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: https://chatgpt.com/share/e/6abeb955-7614-832e-a5e1-b1bd134f...
Like, is this an ice-cream? A tooth?
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
(_)(_)(_) represents the wheels
They do look rather wheel-like; I have to assume you see them as toes though?It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
I think I would have failed this test!
https://chatgpt.com/share/6abf02ae-9a40-83e9-a432-00bf064f60...
The images: https://imgur.com/a/ig6sn6I
... They're not what I would have described. For me, 99.something% flesh and blood with less than 1% metal, glass, and probably some microplastics...
The first one I see a person with a big tall hat and a big nose.
The second one... I do see the black statue with a figure in white in front of it.
The third one is immediately two fish looking at each other.
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
Maybe they mean a non-fiction "Learn Javascript in 21 Days" type of book? If not, I really want to read a novel-length (or even novelletee - 80k words or so) book written by an AI to see for myself if the results are comparable to non-self-published authors.
I've gotten AI to generate an entire book, with the plot points kept coherent (this character knows that when this chapter starts, that kind of thing). It didn't impress my wife, but it was a book.
Not sure how I'd share it. I suppose I could order a copy for you.
Admittedly, I don't think there's any chance it was one shot, rather a pastiche of many, many prompts guided by the quasi-author.
Challenge: “A stopped clock can’t be reliably used to tell the time.”
Group A: “A stopped clock is incapable of displaying the correct time and anybody that says they’ve seen one do so is a stupid lying liar! I saw a pre-release paper that proved it!”
Group B: “I looked at my stopped clock 5 times yesterday, 12 hours apart confirmed with the atomic clock, and it was within 2 seconds of the correct time every time except once. The people that can’t use stopped clocks to view the correct time are dumb dumbs that are looking at them wrong.”
But a standard coding harness and a good agent that isn't too narrowed in to coding (Gemini or Grok work) with a good skill file and a short prompt can get you a coherent book, with reasonably thought-out story arc, multiple characters, internally consistent world, etc. The pacing will suck, and there will be some misunderstandings about the physical world that read like plot holes.
I would say AI can currently write coherent, but not compelling books. The "good writing" claim isn't really true yet imho. But just like with code you can bring it from 80% there to 95% there with a little hand-holding and guidance along the way.
Not that I think good writing skills are even enough to make AI books compelling over human-written books
Would you consider it a one-shot event, or something you have got to prompt carefully until it gets the 80k wordcount?
But AI can run through those steps on its own if you don't want to give input in between. You have to give it a guideline of what steps it should take, but that's likely just because there is lots of reinforcement learning being done on the steps to produce good code, and basically none on the steps to write a complex story
This sounds like a "continuity supervisor" job in movie industry:
>A script supervisor (also called continuity supervisor or script) is a member of a film crew who oversees the continuity of the motion picture including dialogue and action during a scene. The script supervisor may also be called upon to ensure wardrobe, props, set dressing, hair, and makeup are consistent from scene to scene. The script supervisor keeps detailed notes on each take of the scene being filmed.
I'm quite certain it is an AI book. My aunt read it and enjoyed it.
I would say this rises to the level of a "coherent" book. Although I would not consider it a "good" book.
[0]: https://www.amazon.com/BOY-SIERRA-MORENA-Rodr%C3%ADguez-Wolv...
The test here was, specifically, coherency, which I feel is going to be difficult for even SOTA models to track over the standard novel length of 100k words (80k for YA novels).
Maybe if it's a novel about a single individual in a limited setting (so, pretty boring, but we are not measuring quality of plot), a SOTA model can keep it coherent?
I dunno how it would do with something like Stephen King's IT or Under the Dome (multiple parallel plots over 450k words, involving multiple main characters, with events on every single page needing to be coherent with the rest of the book), but before we even get there, standard novel lengths (around 100k words) would be the test for coherency.
Early LLMs were trivially non-coherent. The stories it wrote should shift constantly. The story could start on the moon, then the character could drive to Madison, Wisconsin.
A more obvious way to see this is in video generators. If the "camera" turns 180° twice, we often look at a completely different scene.
Can current models write an entire coherent book? I don't know. I have never wanted to generate such a book. But "do I like this book" is not a good test for coherence.
Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
Most business is correspondence with people who want money from you and people you want money from.
I'm not sure what risk a businessman is willing to accept, but I'd be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren't anywhere near that yet.
[0] https://www.bbc.com/travel/article/20240222-air-canada-chatb...
There’s nothing wrong with that, but starting a business means fading several once-a-year risks of failure and running even an established one means facing several once-a-century risks every year.
I've stopped reviewing the code in my mobile app project months ago. I now only look at files changed and lines count in MRs. Functionality is best verified via manual QA testing.
I know that people here are going to doubt the quality of my project and say that it is impossible, but they are clueless and have evidently not build a project in this way. Experiences from eg. corporate backend work are hardly relevant.
It is clear to me that the fewer consequential mistakes people find during code review, their attention to code reviews is going to go down, to the point of also skipping them.
I expect that for most development, not reviewing the code will be the standard by March of next year. Only critical code like authentication will be reviewed.
It's a paid app, and after 6 months of work I expect to publish it this month.
In other words, you will have to take my word for its quality.
One thing that I do is have each change reviewed by another model. If I implement with Codex, I would have Claude do the review, or the other way around.
I did not benchmark this against a review from a subagent with the same model, so I don't know if that in particular helps.
Take this calendar of events and rank your preferences for the event lottery. Events are spread across 5 weeks, some are 20 minutes away, some are 4.5 hours away, some are Friday/Saturday, some are Saturday/Sunday, some are historically extremely competitive, some fill up in round 1, others don’t even fill after round 2. For the last two years, I’d written some scripts to scrape the event sites, find the addresses, ask Google Maps to give me driving distances and times, scrape prior year registration information to find which teams went and the strength of those teams, etc. It was several hours of effort.
This year, ChatGPT was capable of doing almost all of that basic research and data conversion, filling out our internal spreadsheet. It probably still took 4 hours on the wall clock, but at 2% attention (5 minutes of human toil).
I doubt I go a single workday without some kind of “I have an idea and I know there are disparate data sources out there; go find those and cross-correlate or extract the relevant data points.” question that is now 10-50x more efficient than 2 years ago.
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
Even IF they need code, they need at best a CRUD app to track patients, that's it. There is no way Fable or Opus 5.5 can't one-shot a village vet clinic app in 30 minutes, and only with "I need a village vet clinic app" as a prompt, and whatever questions it decides to ask along the way with it's "ask user" tool.
Or a florist, to use the example from a sibling comment.
Code is tiny part of "business".
Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.
Dog-English machine translation.
I did support for around 15 independent vet clinics in the past.
And that one shot app is not going to be bug free.
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
People are trying, but I don't think they'd be happy with 91.5% success rate: https://www.emerald.com/ir/article-abstract/doi/10.1108/IR-0...
It's just amazing how quickly we accept that models are good at something.
My florist boss can't get Claude to automate rose pruning. But she sure as hell doesn't need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
Yes indeed, but I was responding to "So are we all going to be out of a job?", not "Will AI radically change the jobs market?"
We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
Apparently not.
Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
I didn't intend to say anything about AI or it's current capabilities in my post, honestly
Just that the turing test was always a flawed thought experiment and we've known that for a long time
Personally I don't really think any current AI is sentient but I do think it is more capable and sophisticated than any chinese room thought experiment was ever imagined to be.
I'm not sure if I care about its nature all that much, honestly. I'm much more interested in the social implications of its existence and how we continue to make humanity relevant in a world where AI and Robotics are doing more and more
I'm very afraid that if humanity loses relevance in the current social environment it will be very bad for most of us
Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.
Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.
Eventually I just had to search youtube videos until I found someone doing it in real life.
I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
When coding? I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid. As long as I feed it enough context its really good but it does overfit a lot still.
This means you need a skilled operator to get good results out of the AI, which casts some doubt about the way in which they are intelligent.
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
[0]: https://arxiv.org/pdf/2503.23674 (now published at https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.
[2]: "Dull Rigid Human meets Ace Mechanical Translator" (https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)
Turing test does not mean perfectly human it just means you can talk to one without knowing that has been passed for a long time now.
I have not been able to tell for a long time now especially if I directly give it human like writing instructions for outbound content.
Eliza can sometimes pass the Turing test!
Conversations with strangers can be hard to get going, but they aren't this bad.
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.
I didn't respond to it because it's a bad argument.
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it. > I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.
Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!
Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.
I don't think philosophy has any real value in assesment here. I'm bias but even before AI I thought it wasn't accurate model of how thought works and I think AI has reinforced that.
Reminds me of the 4 humors of medicine in medieval europe. It has some truth but it's not really accurate.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
[1] parrotchess.com, no longer available. Previous discussions: https://hn.algolia.com/?q=parrotchess.com
I think the theory is that the LLM having a high chess ELO was a pet project of a researcher that left.
My "goalpost", unmoved for decades and with nowhere to move it to, was always to independently and repeatedly make contributions to research maths. That one has now been met. That doesn't mean I suddenly think LLMs think like humans (or at all), that their way of doing maths is equivalent to or a replacement for the human one, that they are conscious, that their differences vs humans don't matter, that there will be a singularity or anything else, but it does mean I no longer have a specific, well defined "task" that I don't think they'll ever be able to do.
Like, who's still denying that AI can "write software" (https://stoppels.ch/goalposts/?c=13650937) or pass the turing test (https://stoppels.ch/goalposts/?c=11255120)?
However, I feel the the canonical Turing test bet is still not passed: https://longbets.org/1/
In 2002 Kurzweil bet Kapor that no computer will pass the Turing test. The test proposed is a 2 hour unrestricted conversation with three relative experts, Kurzweil, Kapor, and one third person they agree on. These are pretty extreme conditions, frontier LLMs can fool the most people in shot convos. But, I do not believe any LLM can make it two hours without slipping up against three people who are fairly familiar with LLMs.
I've personally had many 2 hour conversations with frontier models under various personas and I don't think any are close to passing for anyone who has read any quantity of LLM writing. Context rot is still very real, LLMs still have a ton of tells, and they tend not to be willing to push back enough. That being said these are things you notice talking to LLMs a lot, I do think that most of the time frontier LLMs could fool someone who does not use them much for two hours, but they'd probably notice weirdness.
Like many questions here it's somewhat ambiguous. Which Turing test was the original poster talking about? Which Turing test are we thinking about? What is the original spirit of the Turing test? I marked that one as unsure but I could see someone fairly marking it as yes or no.
Year Fraction answering yes
---- ----------------------
2016 56.89%
2017 50.78%
2018 51.21%
2019 39.59%
2020 37.32%
2021 48.48%
2022 42.58%
2023 43.93%
2024 42.11%
2025 38.57%
2026 28.84%To be fair, none of them have actually been met. Mostly what’s stopping them is the “reliably” part.
https://news.ycombinator.com/item?id=48517353
I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic
https://news.ycombinator.com/item?id=48500827
I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
How do you measure that?
Just scale it up a few more orders of magnitude, that should get us there! /s