There's about ~10% point improvement left (i.e, from 80% to 90%) before it starts to stagnate. We've seen the same with predictive models benchmarked on ImageNet et. al.
The tests aren't trying to measure intelligence, but rather whether you've learned the material.
Not saying you hold the same opinions -- but I wouldn't be surprised if people's take on these tests is more about what is convenient for their psyche than any actual principled position.
Not only that, but the innovation around this tech is also just getting started. It's immediately applicable for business use. The classical techniques still have their uses, of course.
You mean since the 2010's ?
Do you have any source for that number. What even quantifies a percentage in a generative model. Closeness to human ability?