the improvement of gpt4 over even 3.5 is significant and for practical applications these improvements exceed what is indicated by benchmarks
There are many things which earlier models maybe somewhat did, but only large models do reliably to the point of being usable for more than tech demos