> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.
This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.