I'm fairly lay but imo GPT-3 demonstrates pretty soundly that huge models are no magic bullet- it's got >2x as many parameters as the human brain has neurons and it can't do long division. Dogs and other animals get by just fine having less than 1% as many neurons as humans.
Even a billion parameters is a huge model, and a factor of 2.4x increase is not going to make a tremendous difference in your performance. In particular the data-heavy nature of vision stuff means that you'll be bottlenecked by training more than memory, AFAIK (again, lay).