44 karma · joined February 26, 2015
Of course, this is not correct, and all we can really say is that the test did not reject the null, given the size (type I error rate) and power (type II error rate) of the test. It's entirely possible that the null should be rejected, but the test is just not very good (i.e. it might have the correct size, but very poor power).
So given some complex and eccentric real-world data, how can we figure out what the power of a given test might be in practice? If you have some idea of what the data generating process might look like then one option is to do some simulations. This enables you to see what the size and power properties of your test are by empirically measuring the type I and type II error rates.
I can certainly see why this is a good comparison, because it's true that both engineering methods and statistical methods rely on sets of given assumptions, but it's also really important not to take this analogy too far. Engineering is ultimately something that is done in a mechanistic world with primarily deterministic outcomes, whereas statistical modeling is conducted in a stochastic world with probabilistic outcomes, so it wouldn't be good to think about machine learning as predominantly mechanistic in nature (in spite of its name). Of course, a lot of the seven points that follow in the post actually emphasize the importance of stochastic factors (e.g. outliers, variance issues, collinearity, etc), so the author is clearly not making this mistake, but it might be good to clarify for anyone else who is reading.
"6. Use linear model without considering multi-collinear predictors"
This is a great point, and just to expand on it a bit, you can also have situations where you have simultaneity, i.e. two or more of your features or predictors are either functions of each other and/or functions of some third variable. This type of problem is more difficult to detect but can cause serious problems with interpreting the regression coefficients as it's ultimately a type of endogeneity, which means that common approaches like OLS will not be consistent.
In terms of explaining this to non-stats people, you might want to avoid explaining the p-values directly to them (as it's very easy for people to get confused about what p-values actually mean), so instead you might simply show them which features are "statistically significant". In other words, try to explain the results in a qualitative way rather than a strictly quantitative one.
I know that Augur is planning to allow people to report event outcomes as "invalid", and that might clear up some of these cases, but what about an event where the outcome initially appears to be objective, but after-the-fact it is unclear and open to interpretation, resulting in two distinct camps of reported outcomes (and hence multimodality)? Perhaps one solution would be to simply declare events with strongly multimodal properties in the distribution of reported outcomes as invalid, and thus avoid the somewhat arbitrary decision where one mode would be declared the "consensus" by the algorithm, costing reputation to those who reported the other mode as the outcome.
In the presence of a poor prior the Bayesian probability would be biased in some way, so frequentists would say that the p-value in the absence of priors is actually superior in this case. Bayesians would reply that if they thought the prior might be poor then they would simply consider multiple different priors, but it's not clear how this would improve things much over the frequentist approach that simply assumes that the prior is unknown.
> So when you don't know the prior and you observe a low p-value on something, isn't that just "preliminary research" that needs to be further confirmed with other methods or at least the same test but using other data?
Yes, when you observe a p-value with low significance it should definitely indicate to you that more testing is necessary, either by using different testing methods, gathering new samples, or even just increasing the original sample size if that's possible. What I was trying to suggest in my last paragraph was that this should be the case even when we have highly significant p-values, because even significant p-values are not decisive. So even when we have "confirmatory research" that is highly statistically significant, we should still do all of the things that we would do when we have a p-value with low significance. It is sometimes the case that this subsequent research will overturn even very highly statistically significant results (though often this is unfortunately because mistakes in the original statistical methodology are uncovered).
The big problem with machine teaching in many practical applications is what the paper refers to as the "glaring flaw", and that is that you often don't know what the learning algorithm might look like (e.g. in the provided nefarious example of trying to defeat a spam filter). In fact, the learning algorithm could be arbitrarily complex.
In the case where you do know the learning algorithm exactly (e.g. the learner is a robot where you have its precise specifications), the problem is the deterministic optimization problem described in this paper. But when the learning algorithm is unknown, the problem becomes stochastic, and then you're facing all of the traditional problems with optimization in a probabilistic space (e.g. overfitting, robustness problems, etc). That's not to say that it's strictly impossible to apply machine teaching approaches in such a case, it's just that it's a much more difficult problem to find a somewhat optimal training set.
However, I would say that Bayesian approaches do have a big advantage in terms of helping with the interpretation problems that plague frequentist significance testing. Namely, as the OP article points out, Bayesian approaches reformulate the testing question in a way that is more intuitive, i.e. "what is the probability of the hypothesis given both the prior probability and the new data?". So yes, Bayesian methods surely do not fix everything, but since interpretation of statistics is such a major concern, they can be quite beneficial.
As suggested, a good approach is to take p-values not as conclusive or decisive, but rather as a tool that must be supplemented by other statistics. In particular, the article emphasizes Bayesian methods, which can certainly provide additional information, but this approach can also be rather limited when priors are not well-defined or are entirely unknown, which is unfortunately often the case in many problem domains.
One potential question is how to determine the nature of the distinction mentioned in the conclusion between "preliminary research" and "confirmatory research", particularly in cases where statistics provide the primary evidence, as in, e.g. psychology. Further studies in the same vein as the preliminary research can certainly provide additional supporting statistical evidence, but this doesn't escape the problem that all of the evidence is probabilistic in nature. The key issue here is that since statistical approaches can only give probabilistic evidence that a hypothesis is correct, then they strictly cannot tell you what is certainly true, so even confirmatory research is quite open to falsification. So we wouldn't want the label of "confirmatory research" to somehow suggest to the public the idea that it is certainly correct.