From their methodology section:
> The data include responses only from the official Python Software Foundation channels. After filtering out duplicate and non-reliable responses, the data-set includes more than 18,000 responses collected in October and November of 2018 via promoting the survey on python.org, the PSF blog, the PSF’s Twitter and LinkedIn accounts, official Python mailing lists, and Python-related subreddits. No product-, service-, or vendor-related channels were used, in order to prevent the survey from being slanted in favor of any specific tool or technology.
That might have influenced significantly in the selection of the Python users taking the survey: why answer a survey setup and sponsored by a private company you are not really familiar with and that you don't necessarily trust as a result?
There is also a second selection bias: the people who follow the PSF and the official python.org communication channels are probably nerdier than the average Python programmer. This is reflected in the OS statistics. The survey results suggests than 53% of the Python users do not use Windows at all.
On the scikit-learn online documentation, after removing the mobile traffic we have: Windows at 61%, macOS at 23% and Linux at 15%.
On stackoverflow, the primary OS statistics are: Windows at 47%, macOS at 27% and Linux at 23%.
This is the js survey results all over again: no. Unless you can statistically prove the results are biased, you don’t get to ignore the results because you dont like them.
Finding data points with no methodology that contract the survey result does not invalidate the survey results.
Thats. not. how statistics work.
A great deal of effort was put into this survey, and the stats you’re looking at are more likely biased than the ones in this survey.
The stats and the methodology here are clearly documented; if you want to argue with them, be specific and provide concrete statistical proof for your assertions.
Specifically, why do the stats you have prove anything, and what confidence do you have that they are representative?
In this situation often the smaller dataset is wrong.
...not always. But often.
Human intuition based on limited data can seem compelling, but it’s always worth acknowledging you might be the outlier.
18000 respondents is a lot, especially when a specific effort has been made to sample from various sources.
The parent post didn’t even bother to check their own biases.
I don't know if that's the case here, but I do often think that people gravitate towards the first thing that they can understand that seems to, at a glance, check out, without investigating.
Similarly, stackoverflow has a lot more users than just Python users, so it says little about how many Python programmers use Windows.
The ads wouldn't have appeared on the Requests docs but it could have on many others.