Thanks for the fun write-up though. Overboard rage aside, it's a good reminder that we all should double-check our QC processes.
Thanks for the fun write-up though. Overboard rage aside, it's a good reminder that we all should double-check our QC processes.
Back in the days when "system" was a reasonable font choice, fixing a bug after shipping would take weeks, months, or years, and could cost $10^4 to $10^8. You might have had to pull actual boxes off actual shelves somewhere!
In this story as it actually happened, the bug was fixed and the fix was deployed in a single day, at the cost of a few laughs.
I'm sure they tested on Windows when they added Segoe to the list, so the problem is that they didn't test again after fiddling with other fonts that they (incorrectly) thought would be irrelevant to Windows. And remember, nothing was broken in a way that a machine would notice, so automated acceptance tests wouldn't have caught this regression.
I think it's worth considering the possibility that this is a story of Medium acting more or less in their economic interest (maybe even in the economic interest of their users?) rather than a story of a bunch of hipsters with a hilariously broken QC policy.
I'd like to nit here that visual regression test harnesses for web application development do exist. [1] [2] [3] Using something like Sauce Labs you could conceptually have these tests run on Windows machines in the cloud.
However, having implemented visual regression testing in the past in various projects, I can personally vouch that it's a huge waste of time. Your tools will break all the time, your tests will break all the time, and you'll generally have a bad time.
[1] https://github.com/chenglou/node-huxley [2] https://github.com/Huddle/PhantomCSS [3] https://github.com/bradgignac/janus
In analogy with scientific modeling, there's a risk of over-fitting your tests against the exact circumstances of your product as it exists today. You want tests to capture the important aspects of your product, while having a little bit of flexibility for the "noise" related to differences between platforms, and changes to mostly but not entirely unrelated parts of the product.
The tool I used, PhantomCSS, did allow for a threshold underneath which pixel differences would not get picked up, but I opted instead to run my tests locally in vagrant so that the environment would be closer to identical to CI.
So yes, platform noise is just one of the myriad of challenges you have to solve when implementing such tooling.
Why would you be surprised by 80% of Medium readers using Windows? Even 20% non-Windows would represent a massive anti-Windows bias.
Not that this justifies skipping testing on Windows, but it's not impossible to see how it's no longer a top priority, since all of the growth is sitting on the mobile side.