This is p-hacking after-the-fact, right? Seems like the classic example: The broader hypothesis doesn’t hold, but if you look for ways to slice and dice the data you’re likely to find a (spurious?) correlation eventually.
This is p-hacking after-the-fact, right? Seems like the classic example: The broader hypothesis doesn’t hold, but if you look for ways to slice and dice the data you’re likely to find a (spurious?) correlation eventually.
The blog post list this as the original Stata code:
> replace event`i’ = 1 if delta_mct`i’ != 0 | spouse_delta_mct`i’ != 0
and this as the correct one:
> replace event`i’ = 1 if (delta_mct`i’ != 0 | spouse_delta_mct`i’ != 0) & delta_mct`i’ != . & spouse_delta_mct`i’ != .
It looks like the authors didn't properly handle missing values in Stata, leading to marking people with missing health information being marked as being "severely ill" instead of being excluded from the analysis.
It's an unfortunate mistake, but it happens.
(Honest question)
You're saying the original analysis was wrong due to a coding error. I believe that's also true, but that's not what they were discussing. The variable names are inscrutable, but the article text also seems to imply that line (mis)codes divorce, not severe illness:
> People who left the study were actually miscoded as getting divorced.
So they actually found a correlation between severe illness and leaving the study. That's perhaps intuitive, if those people were too busy managing their illness to respond.
Although you're right that there's likely no relationship between wives developing heart conditions and subsequent divorce, there's not enough information from the article to know whether there's anything statistically meaningful about heart conditions specifically. It seems more likely that it's just statistical noise. I read that section and got the impression that the relationship is interesting but it doesn't necessarily mean anything, rather than implying that the original study still has some validity.
For another, the study doesn’t look at that many diseases, just four: cancer, heart disease, lung disease, and stroke. It’s possible the original study was doing some p-hacking in order to reach significance but these are pretty major categories of disease so they seem defensible. It would not surprise me if these are the only diseases in the original dataset.
Finally, the results are significant at the .01 level. Combined with the number of diseases, this is not anything as blatant as the classic XKCD green jellybean comic, although more subtle p-hacking could be at play.
Retracted: https://journals.sagepub.com/doi/pdf/10.1177/002214651456835...
Corrected: https://journals.sagepub.com/doi/abs/10.1177/002214651559635... (Not open access but it’s available from sci-hub.)
xkcd/882