The difference is a lot more than just throwing scale at it, pretty much everything useful comes from an evolving landscape of post-training techniques.
Of course, param count and context length are also important because they increase the model's overall fidelity, but a base model without SFT, RHLF etc is effectively useless.