at this point i wonder what’s different between all these model. all of them have quite similar model architecture. It is just how much money you can burn to train it?
I think while technically they differ mostly only in terms of the training materials thrown into it, the outcome is that each model is good at something and bad at others, just like human being. Soon you'll need standardized tests and HR department to evaluate individual LLM performance. :)