Statisticians use a technique that leverages randomness to deal with the unknown
quantamagazine.org
quantamagazine.org
An ELI5 intro: https://abidlabs.github.io/EM-Algorithm/
Any sources for that? As far as I remember, EM is used to calculate actual cluster parameters (means, covariances etc), but I'm not aware of any usage to estimate what number of clusters works best.
Source: I've implemented EM for GMMs for a college assignment once, but I'm a bit hazy on the details.
For PCA or factor analysis, there's lots of ways but without some way of determining ground truth it's difficult to know if you've done a good job.
If so then estimating the marginal likelihood of each one and comparing them seems pretty reasonable?
(I mean in the sense of Jaynes chapter 20.)
EM imputation (or single imputation in general) fails to account for the uncertainty in imputed data. You end up with artificially inflated confidence in your results (p-values too small, confidence/credible intervals too narrow, etc.).
Multiple imputation is much better.
ps. Gelman was Rubin's doctoral student.
The table is particularly useful as it describes what the article is all about in a way that can stick to students' minds. I'm very grateful for QuantaMagazine for its popular science reporting.
Rubin Causal Model
Propensity Score Matching
Contributions to
Bayesian Inference
Missing data mechanisms
Survey sampling
Causal inference in observations
Multiple comparisons and hypothesis testing
Not the least of which is that it's far too easy to do the equivalent of p hacking and get your data to be significant by playing games with how you do the imputation. Garbage in, garbage out.
I think all of these methods should be abolished from the curriculum entirely. When I review papers in the ML/AI I automatically reject any paper or dataset that uses imputation.
This is all a consequence of the terrible statics used in most fields. Bayesian methods don't need to do this.
No. AI/ML folks don't do imputation on our datasets. I cannot think of a single major dataset in vision, nlp, or robotics that does so. Despite missing data being a huge issue in those fields. It's an antiqued method for an antiqued idea of how statistics should work that is doing far more damage than good.
In another related intuition for a probable foot gun relates to learning linearly inseparable functions like XOR which requires MLPs.
A single missing value in an XOR situation is far more challenging than participant dropouts causing missing data.
Specifically the problem is counterintuitively non-convex, with multiple possibilities for convergence without information in the corpus to know which may be true.
That is a useful lens in my mind, where I think of the manifold being pushed down in opposite sectors as the kernel trick.
Another potential lens to think about it is that in medical studies the assumption is that there is a smooth and continuous function, while in learning, we are trying to find a smooth continuous function with minimal loss.
We can't assume that the function we need to learn is smooth, but autograd specifically limits what is learnable and simplicity bias, especially with feed forward networks is an additional concern.
One thing that is common for people to conflate is the fact that a differentiable function is probably smooth and continuous.
But the set of continuous functions that is differentiable _anywhere_ is a meger set.
Like anything in math and logic, the assumptions you can make will influence what methods work.
As ML is existential quantification, and because it is insanely good at finding efficient glitches in the matrix, within the limits of my admittedly limited knowledge, MI would need to be a very targeted solution with a lot of care to avoid set shattering from causing uncongeniality, especially in the unsupervised context.
Hopefully someone else can provide a better productive insights.
Single imputation is garbage for accurate inference, as it reduces variance and thus confidence intervals as P(missing) increases.
MI is a useful method for alleviating this bias (though at the cost of a lot more compute).
That's why it gets used, and it's performed extremely well in real world analyses for basically my entire life (and I'm middle-aged now).
> especially in the unsupervised context.
I wouldn't use MI in an unsupervised context (but maybe some people do).
The incentives in medicine are more similar to those in academia, where your job is to cook up data that convinces someone else of your results, with highly imbalanced incentives that reward fraud.
The problem is that data is never actually missing at random and there’s always some sort of interesting variable that confounds which pieces are missing
More specifically, how do you determine if the pattern you seem to be identifying is actually related to the phenomenon being measured and not an error in the measurement tools themselves?
For example, a significant pattern of answers to "Yes / No: have you ever been assaulted?" are blank. This could be (A), respondents who were assaulted are more likely to leave it blank out of shame or (B) someone handling the spreadsheet accidentally dropped some rows in the data (because lets be serious here, its all spreadsheets and emails...).
While you could say that (B) should be theoretically "more truly random", we can't assume that there isn't a pattern to the way those rows were dropped (i.e. a pattern imposed on some algorithm that bugged out and dropped those rows).
If the “which data is missing” information can be used be to compress the data that isn’t missing further than it can be compressed be alone, then the missing data is missing at least in part due to the phenomenon being measured. Otherwise, it’s not.
We’re basically just asking if K(non-missing data | which data is missing) < K(non-missing data). This is uncomputable so it doesn’t actually answer your question regarding “how to determine”, but it does provide a necessary and sufficient theoretical criteria.
A decent practical approximation might be to see if you can develop a model that predicts the non-missing data better when augmented with the “which information is missing” information than via self-prediction. That could be an interesting research project...
What is the distinction?
The scenario described in the paper could be represented in a Bayesian method or not. “For a given missing value in one copy, randomly assign a guess from your distribution.” Here “my distribution” could be Bayesian or not but either way it’s still up to the statistician to make good choices about the model. The Bayesian can p hack here all the same.
A more interesting approach, let's call it OPTION2, would be to sample from the predictive distribution of a regression (regression mean + noise), which would result in more variability in the imputations, although random so might not what you want.
The multiple imputation approach seems to be a resampling methods of obtaining OPTION2, w/o need to assume linear regression model.
This must be about the confidence of the approach. Maybe interpolation would be overconfident too.