Text2SQL was 75% on bird-bench 6 months ago. Now it's 80%. Humans are still at 90+%. We're not quite there yet.
I suspect text-to-sql needs a lot of intermediate state and composition of abstractions, which vanilla attention is not great at.
A user having to come up with novel queries all the time to warrant text 2 sql is a failure of product design.
But this is precisely why we're seeing startups build insane things fast while well established companies are still questioning if it's even worth it or not.
People got good results on the test datasets, but the test datasets had errors so the high performance was actually just the models being overfitted.
I don't remember where this was identified, but it's really recent, but before GPT-5.