The thing that makes it hurt is that what seems like a minor rewording of that statement, "We're 95% confident that the range we calculated contains the true population mean," would be correct. (It's not the definition of a CI, but it is implied by the definition of a CI.) And that the reason why one is true and the other is false is down to the technical distinction that the true population mean is a discrete value and not a draw from a random variable, so, from a strictly mathematical sense, it is nonsensical to apply probabilities to it. By a the same token, an even more subtle rewording gets us back to falsehood: "There's a 95% chance that the range we calculated contains the true population mean."
Which, delving into to that level of hair-splitting makes for interesting math, but also leaves me with the opinion that, sure, the most common intuitive interpretation is wrong, but it's wrong in a way that isn't really of much practical importance.
By contrast, the most common intuitive interpretation of the p-value is disastrously wrong.
Another discrepancy between frequentists statistics and the article is that yes, the values at the boundary of your interval are as credible as in the center.
Another note: 'confidence interval' typically refers to the frequentist meaning, whereas 'credibility interval' is used in the Bayesian setting, when describing an interval of the posterior with 95% probability (which is arguably more interpretable). The usages of the two terms do not seem to generally be strict, however.
What would that mean in the frequentist framework?
To provide an extreme case, during the Iraq war, epidemiologists did a survey and came up with an estimated number of deaths. The point value was 100K, and that's what all the newspapers ran with. But the actual journal paper had a CI of (8K, 194K). There's no reason to believe the true value is closer to 100K than it is to 10K. Or to 190K.
But we can't neither say that the true value is equally likely to be closer to 21 than to 3.
The point is that, from the frequentist definition of a confidence interval, there is nothing at all that we can say about how likely the true value is to be here or there.
It could be 3, 21, or 666 and there is nothing that can be said about the likelihood of each value (unless we go beyond the frequentist framework and introduce prior probabilities).
Yes - sorry if I wasn't clear. I did not mean to imply that each value in the interval is equally likely (and looking over my comments, I do not think I did imply that).
The complaint is that the article is stating otherwise as fact.
>One practical way to do so is to rename confidence intervals as ‘compatibility intervals’
>The point estimate is the most compatible, and values near it are more compatible than those near the limits.
They simply are not in a frequentist model (which is the model most social scientists uses). I agree with the main thrust of the article in that there are many problems with P values. But I am surprised that a journal like Nature is allowing clearly problematic statements like these.
I don't know enough about the Bayesian world to be able to state if his statement is wrong there as well, but if it is correct there, it is problematic that the authors did not state clearly that they are referring to the Bayesian model and not the frequentist one.
(Not to get into a B vs F war here, but I remember a nice joke amongst statisticians. There are 2 types of statisticians: Those who practice Bayesian statistics, and those who practice both).
When you said that "the values at the boundary of your interval are as credible as in the center" you kind of implied that, which is why I asked.
I won't defend the article being discussed, but you opposed their statement that "the values in the center are more compatible than the values at the boundary" with an equally ill-defined "the values at the boundary are as credible as in the center".
What I meant was "there is no reason to prefer values at the center more than values at the boundary" based on the CI (there may be external reasons, though). To me, this is equivalent to your:
>there is nothing that can be said about the likelihood of each value
I find this very silly, since if we ditch the arbitrary 0.95 and go with 0.999.. confidence interval of [-998, 1040] for example. How can one say that one cannot tell if which value is more likely, 21 or 1040?
If this is an actual limitation of the frequentist model like you said, everybody should be a bayesian thinker then. And the "confidence interval" is just a quick way to communicate how wide and where the posterior bell curve is.
The difference is hugely important, the central limit theorem and the implication of a normal distribution commonly applies to sample means.
You can calculate confidence intervals for most (not all) other statistics, like point estimates, but the distributions might not be normal.