We are currently at nonsensical pacing while writing novels.
We are currently at nonsensical pacing while writing novels.
Most human-written books don't do that, so that seems to be a ceiteria for a very different test that a Turing test.
I'm pretty sure if something like this happens some dude will show up from nowhere and claim that it's just parroting what other, real people have written, just blended it together and randomly spitted it out – "real AI would come up with original ideas like cure for cancer" he'll say.
After some form of that comes another dude will show up and say that this "alphafold while-loop" is not real AI because he just went for lunch and there was a guy flipping burgers – and that "AI" can't do it so it's shit.
https://areweagiyet.com should plot those future points as well with all those funky goals like "if Einstein had access to the Internet, Wolfram etc. he could came up with it anyway so not better than humans per se", or "had to be prompted and guided by human to find this answer so didn't do it by itself really" etc.
> With little or no human involvement, write Pulitzer-caliber books, fiction and non-fiction.
So, yeah. I know you made a joke, but you have the same issue as the Onion I guess.
What if we didn’t measure success by sales, but impact to the industry (or society), or value to peoples’ lives?
Zooming out to AI broadly: what if we didn’t measure intelligence by (game-able, arguably meaningless) benchmarks, but real world use cases, adaptability, etc?
2026 news feed: Anthropic cited as AI agents simultaneously block traffic across 42 major cities while trying to capture a not-even-that-rare pokemon
I currently assert that it's not, but I would also say that trying to follow your suggestion is better than our current approach of measuring everything by money.
No. Screw quantifiability. I don't want "we've improved the sota by 1.931%" on basically anything that matters. Show me improvements that are obvious, improvements that stand out.
Claude Plays Pokemon is one of the few really important "benchmarks". No numbers, just the progress and the mood.
Of course, this is just some pedantry.
I for one love that AI is progressing so quickly, that we _can_ move the goalposts like this.
There were popular writeups about this from the Deepseek-R1 era: https://www.tumblr.com/nostalgebraist/778041178124926976/hyd...
Not sure what is better for humanity in long term.
I could build a machine that phones my mother and tells her I love her, but it wouldn't obsolete me doing it.
I am amazed at the progress that we are _still_ making on an almost monthly basis. It is unbelievable. Mind-boggling, to be honest.
I am certain that the issue of pacing will be solved soon enough. I'd give 99% probability of it being solved in 3 years and 50% probability in 1.
Yeah, but 10% plus 20% plus 20%... next thing you know you're at +100% and your server is literally double the speed!
AI progress feels the same. Each little incremental improvement alone doesn't blow my skirt up, but we've had years of nearly monthly advances that have added up to something quite substantial.
(For those too young or unfamiliar: Mary Poppins famously had a bag that she could keep pulling things out of.)
Yes, Z is indeed a big advance over Y was a big advance over X. Also yes, Z is just as underwhelming.
Are customers hurting the AI companies' feelings?
No. It's the critics' feelings that are being hurt by continued advances, so they keep moving goalposts so they can keep believing they're right.
Lets not forget the OpenAI benchmarks saying 4.0 can do better at college exams and such than most students. Yet real world performance was laughable on real tasks.
That's a better criticism of college exams than the benchmarks and/or those exams likely have either the exact questions or very similar ones in the training data.
The list of things that LLMs do better than the average human tends to rest squarely in the "problems already solved by above average humans" realm.
The pace is moving so fast I simply cant keep up. Or a ELI5 page which gives a 5 min explanation of LLM from 2020 to this moment?
In that light, even a 20 year old almost broken down crappy dinger is amazing: it has a radio, heating, shock absorbers, it can go over 500km on a tank of fuel! But are we fawning over it? No, because the goalposts have moved. Now we are disappointed that it takes 5 seconds for the Bluetooth to connect and the seats to auto-adjust to our preferred seating and heating setting in our new car.