You don't need to go to such complicated lengths. Just perform enough tests (as in, a statistically large enough amount) and a distribution will form. That also captures the variability of real world network effects.
Why is a different question and it's not that relevant, I often don't care why something I USE works better I just care that it does.
If it something I BUILD then I would care much more but again this is a whole different issue.
The case for more normalized tests is to find out which browser is factually better designed/written.