And this implements a transformer. Actually it is a very cool didactic example.
What is this based on? Every researcher I've heard talk about this says it's exactly not true, as an uncontested rule, because the larger models will more effectively contain the smaller models, and use them together in ways that the connections between the smaller models can't. Remember, even MOE is to save compute/memory, not to help performance/parameter.
But, it depends on what you're measuring. By spatial/navigational memory, yes, elephants are far far better. Reasoning, no. It would be interesting to see what an elephant or whale eugenics program could result in, since humans have that pesky (or maybe instrumental?) birth canal problem.
This leads to some interesting optimizations. You can quantize all the parameters (or certain layers) of a model and halve or quarter the memory requirement but maintain most of the model's intelligence. This increases the token rate inversely with the size reduction.
Popular quantizations for local models are 8-bit and 4-bit parameter sizes. The Blackwell series of nVidia chips now even support native FP4 math making 4-bit quantizations even faster.