> It's interesting to look at the history of measurements of the charge of an electron, after Millikan. If you plot them as a function of time, you find that one is a little bit bigger than Millikan's, and the next one's a little bit bigger than that, and the next one's a little bit bigger than that, until finally they settle down to a number which is higher.
> Why didn't they discover the new number was higher right away? It's a thing that scientists are ashamed of—this history—because it's apparent that people did things like this: When they got a number that was too high above Millikan's, they thought something must be wrong—and they would look for and find a reason why something might be wrong. When they got a number close to Millikan's value they didn't look so hard. And so they eliminated the numbers that were too far off, and did other things like that...
https://hsm.stackexchange.com/questions/264/timeline-of-meas...
A more theoretical argument is as follows. Suppose we systematically remove from analysis any data point which are more than 1.5 times the IQR below the first quartile or about the third quartile. (https://en.wikipedia.org/wiki/Outlier#Tukey's_fences) Points censured by such a rule are often called "outliers," but because we do not yet know the true underlying probability distribution this terminology is misleading. Then our estimates for skew, kurtosis, and all higher moments will be biased downward, and furthermore the censured sample is more likely to "pass" normality tests such as Shapiro-Wilks which can lead to us applying inappropriate models. However, what is the interpretation of a model when applied to a new observation? If the data point is well within the Tukey fences, then we can use the model to make a valid prediction. However, if it is outside the fences, we can say nothing. What then is the population mean of the predicted value? It is the weighted mean of the values for those data points for which our model is valid and those which lie outside of it. But those that lie outside of it can be enormously influential.
An example can make this concrete. Suppose an actuary estimates the amount of an insurance payouts for house insurance. 99% of the time, these effects are small, a few hundred or a few thousand dollars. But 1% of the time they are much larger, perhaps $500,000 or more. If the actuary applies Tukey's fences to his data and fits a model, he may infer that 1% of claims payout each year, and those payouts average $1,000 dollars. So he sets the price of the insurance at $20/year and makes a 100% profit margin. Until the first house burns down, and his company goes bankrupt.
The point of this is merely to illustrate the dangers of ignoring any effect which simply seems "too large." The beneficial effects of penicillin are "impossibly large" - would we reject the evidence on those grounds? The change in resistance of a semiconductor in the presence of an electric field seem "impossibly large" until the solid state physics is carefully analyzed - would we reject this clear and highly reproducible evidence of a qualitative change in behavior because it is too evident?
A much better approach is to scrutinize the methodology for, say, omitted variables. This is what the more detailed criticisms he cites in the third paragraph do. Or to attempt to reproduce the study, which is perhaps the strongest approach.