Hasn't it been repeatedly shown that many small models perform worse than a large model of the same total parameters?
What is this based on? Every researcher I've heard talk about this says it's exactly not true, as an uncontested rule, because the larger models will more effectively contain the smaller models, and use them together in ways that the connections between the smaller models can't. Remember, even MOE is to save compute/memory, not to help performance/parameter.
But, it depends on what you're measuring. By spatial/navigational memory, yes, elephants are far far better. Reasoning, no. It would be interesting to see what an elephant or whale eugenics program could result in, since humans have that pesky (or maybe instrumental?) birth canal problem.