It's not like same parameter count models are identical, so that doesn't appear to be an indicator for quality, or even compute requirements?
There seems to be more to producing a better model than brute forcing parameter count after all.
There seems to be more to producing a better model than brute forcing parameter count after all.