Humanity's Last Exam
lastexam.ai
lastexam.ai
These tests are getting ridiculous (especially when you take into account the name they choose for them...).
For me things are quite simple. I witness the limitations of LLMs every day. What this benchmark checks if what data it was trained on, that's pretty much it. No need to get "humanity" involved into this.
So I would say the goalposts are moving just in step with the market expectations/hype. Nobody's disputing that LLMs are amazing at generating useful text anymore.