If we look at the math problems they're solving their just now reaching the human frontier... they weren't doing that before.
And your comparison point is model released 2.5 months ago... saying for some use case you didn't see noticeable improvement in 2.5 months (even while other people and benchmarks disagree) isn't a great argument that they aren't improving.