HNHacker News
TopNewBestAskShowJobs

mdbco

44 karma · joined February 26, 2015

submissionscomments
mdbco··on Elsevier sold me a Creative Commons non-commercial licensed article
It looks like the corresponding author of this article (Didier Raoult) is also the Editor-in-Chief of the the journal (Clinical Microbiology and Infection), so it seems entirely possible that he might have relicensed the article to Elsevier when the journal moved over there from Wiley. This would be permitted, since Creative Commons does allow for dual licensing.
mdbco··on The harm done by tests of significance (2003) [pdf]
The single paragraph in the postscript of this paper (part 6) is actually really important. It's very common for people who are using statistical testing in applied settings to entirely forget about type II error (and correspondingly, the power of the test), and so when they see a p-value that isn't significant at a certain level (say 5%), then they just assume that the null hypothesis is true.

Of course, this is not correct, and all we can really say is that the test did not reject the null, given the size (type I error rate) and power (type II error rate) of the test. It's entirely possible that the null should be rejected, but the test is just not very good (i.e. it might have the correct size, but very poor power).

So given some complex and eccentric real-world data, how can we figure out what the power of a given test might be in practice? If you have some idea of what the data generating process might look like then one option is to do some simulations. This enables you to see what the size and power properties of your test are by empirically measuring the type I and type II error rates.

mdbco··on Machine Learning Done Wrong
"Statistical modeling is a lot like engineering."

I can certainly see why this is a good comparison, because it's true that both engineering methods and statistical methods rely on sets of given assumptions, but it's also really important not to take this analogy too far. Engineering is ultimately something that is done in a mechanistic world with primarily deterministic outcomes, whereas statistical modeling is conducted in a stochastic world with probabilistic outcomes, so it wouldn't be good to think about machine learning as predominantly mechanistic in nature (in spite of its name). Of course, a lot of the seven points that follow in the post actually emphasize the importance of stochastic factors (e.g. outliers, variance issues, collinearity, etc), so the author is clearly not making this mistake, but it might be good to clarify for anyone else who is reading.

"6. Use linear model without considering multi-collinear predictors"

This is a great point, and just to expand on it a bit, you can also have situations where you have simultaneity, i.e. two or more of your features or predictors are either functions of each other and/or functions of some third variable. This type of problem is more difficult to detect but can cause serious problems with interpreting the regression coefficients as it's ultimately a type of endogeneity, which means that common approaches like OLS will not be consistent.

mdbco··on Machine Learning Done Wrong
Point #7 is just referring to the magnitudes (or absolute values) of the coefficients. You can still determine which features are relatively important using the coefficient p-values if those are available. This of course is dependent on the necessary assumptions of the regression method that you are using being satisfied, as otherwise the p-values will be biased.

In terms of explaining this to non-stats people, you might want to avoid explaining the p-values directly to them (as it's very easy for people to get confused about what p-values actually mean), so instead you might simply show them which features are "statistically significant". In other words, try to explain the results in a qualitative way rather than a strictly quantitative one.

mdbco··on A Decentralized Lie Detector
This consensus algorithm looks pretty good, but it seems it could get into trouble in cases where the distribution of outcomes is multimodal. One thing that is mentioned is that users would be reporting on several events simultaneously (let's say k events), but it seems entirely possible that there could be strong consensus among k-1 events, but multimodality in the reported outcomes for the kth event, making it very difficult to see who is wrong and who is correct for that event.

I know that Augur is planning to allow people to report event outcomes as "invalid", and that might clear up some of these cases, but what about an event where the outcome initially appears to be objective, but after-the-fact it is unclear and open to interpretation, resulting in two distinct camps of reported outcomes (and hence multimodality)? Perhaps one solution would be to simply declare events with strongly multimodal properties in the distribution of reported outcomes as invalid, and thus avoid the somewhat arbitrary decision where one mode would be declared the "consensus" by the algorithm, costing reputation to those who reported the other mode as the outcome.

mdbco··on Psychology Journal Bans Significance Testing
> Is the cautious approach then to treat a p-value in the absence of priors on the same level as a p-value in presence of unfavorable priors?

In the presence of a poor prior the Bayesian probability would be biased in some way, so frequentists would say that the p-value in the absence of priors is actually superior in this case. Bayesians would reply that if they thought the prior might be poor then they would simply consider multiple different priors, but it's not clear how this would improve things much over the frequentist approach that simply assumes that the prior is unknown.

> So when you don't know the prior and you observe a low p-value on something, isn't that just "preliminary research" that needs to be further confirmed with other methods or at least the same test but using other data?

Yes, when you observe a p-value with low significance it should definitely indicate to you that more testing is necessary, either by using different testing methods, gathering new samples, or even just increasing the original sample size if that's possible. What I was trying to suggest in my last paragraph was that this should be the case even when we have highly significant p-values, because even significant p-values are not decisive. So even when we have "confirmatory research" that is highly statistically significant, we should still do all of the things that we would do when we have a p-value with low significance. It is sometimes the case that this subsequent research will overturn even very highly statistically significant results (though often this is unfortunately because mistakes in the original statistical methodology are uncovered).

mdbco··on Machine Teaching: An Inverse Problem to Machine Learning [pdf]
This is a nice little paper that provides a great introduction to machine teaching. I think the Socratic dialogue format was an excellent choice as it makes it very easy to follow.

The big problem with machine teaching in many practical applications is what the paper refers to as the "glaring flaw", and that is that you often don't know what the learning algorithm might look like (e.g. in the provided nefarious example of trying to defeat a spam filter). In fact, the learning algorithm could be arbitrarily complex.

In the case where you do know the learning algorithm exactly (e.g. the learner is a robot where you have its precise specifications), the problem is the deterministic optimization problem described in this paper. But when the learning algorithm is unknown, the problem becomes stochastic, and then you're facing all of the traditional problems with optimization in a probabilistic space (e.g. overfitting, robustness problems, etc). That's not to say that it's strictly impossible to apply machine teaching approaches in such a case, it's just that it's a much more difficult problem to find a somewhat optimal training set.

mdbco··on Psychology Journal Bans Significance Testing
You're absolutely correct that Bayesian approaches are not magical and do not suddenly supply you with vastly more information than frequentist approaches (particularly when you have a really poor prior, in which case the Bayesian approach will be similarly poor). Bayesian statistics is certainly very popular right now, but it should not be looked at as some sort of panacea for all statistical problems.

However, I would say that Bayesian approaches do have a big advantage in terms of helping with the interpretation problems that plague frequentist significance testing. Namely, as the OP article points out, Bayesian approaches reformulate the testing question in a way that is more intuitive, i.e. "what is the probability of the hypothesis given both the prior probability and the new data?". So yes, Bayesian methods surely do not fix everything, but since interpretation of statistics is such a major concern, they can be quite beneficial.

mdbco··on Psychology Journal Bans Significance Testing
It's actually entirely possible to do some "checking of your answer" for p-values, as well. As you mentioned, for practical math, you can often validate your answer against the initial assumptions. This is true for statistical testing too, as it typically relies on many theoretical assumptions. So what you can do in practice is propose a different set of assumptions, perform the hypothesis test in a manner that follows those new assumptions, and see if you obtain a similar result. Typically, for any given hypothesis that you want to test, there are several possible methods for performing that test, so you can redo your test many times. This is one type of robustness checking, which includes many other things as well (e.g. running your test over subsamples or resamples of the data, checking for sensitivity to outliers, etc). Good statisticians generally like to do lots and lots of robustness checking.
mdbco··on Psychology Journal Bans Significance Testing
The article is certainly correct that p-values and confidence intervals (or confidence sets, in multi-dimensional contexts) are widely misunderstood, not just in psychology or other social sciences, but in the hard sciences as well. The problem is even worse when you look outside of academia at common practices in more applied settings.

As suggested, a good approach is to take p-values not as conclusive or decisive, but rather as a tool that must be supplemented by other statistics. In particular, the article emphasizes Bayesian methods, which can certainly provide additional information, but this approach can also be rather limited when priors are not well-defined or are entirely unknown, which is unfortunately often the case in many problem domains.

One potential question is how to determine the nature of the distinction mentioned in the conclusion between "preliminary research" and "confirmatory research", particularly in cases where statistics provide the primary evidence, as in, e.g. psychology. Further studies in the same vein as the preliminary research can certainly provide additional supporting statistical evidence, but this doesn't escape the problem that all of the evidence is probabilistic in nature. The key issue here is that since statistical approaches can only give probabilistic evidence that a hypothesis is correct, then they strictly cannot tell you what is certainly true, so even confirmatory research is quite open to falsification. So we wouldn't want the label of "confirmatory research" to somehow suggest to the public the idea that it is certainly correct.