In fact, the wikipedia article even says "A scatter plot is used when a variable exists that is under the control of the experimenter". https://en.wikipedia.org/wiki/Scatter_plot
If the author had read that part of the wikipedia article, I guess his claim would have been more specific :-)
It's simple to do and mimics reversing the effect of truncation of the data (at least for continuous quantities). Just use uniformly distributed values that are as wide as one bin width.
For most purposes, I prefer adding dither, and then using transparency, to moving to a density plot, for exactly the reason you mention -- the density plot introduces another parameter, the smoothing method, which puts another layer between you and the data.
Furthermore, if my data has two outliers that are near each other, they may well be indistinguishable in the hexplot from one (or five) clustered outliers.
Your post was very interesting and your examples are great. I'll definitely use hexplots in the future. But I will still default to scatterplots. It's just easier to see if there's something wrong with the data, and they require less interpretation.