Assessing Cardiovascular Risk Factors with Computer Vision
research.googleblog.com
research.googleblog.com
"Funny story. We had a new team member joining, and @lhpeng suggested an orientation project for them of "Why don't you just try predicting age and gender from the images?" to get them familiar with our software setup, thinking age might be accurate within a couple of decades,and gender would be no better than chance. They went away and worked on this and came back with results that were much more accurate than expected, leading to more investigation about what else could be predicted."
- The standard CV risk calculators that we actually use in medicine don't do a tremendous job at discriminating risk. The AUC for the pooled risk equations tends to be around 0.74. As a refresher, AUC can be interpreted as the probability that someone with CVD has a risk score that is higher than someone without.
- This deep learning model had an AUC of about 0.7, similar to the pooled risk equation which was about 0.72
- Adding this deep learning model to the pooled risk equation did not change the AUC (0.72 -> 0.72). There may have been recategorization, but we can't tell from the paper as published.
- The pooled risk equations in this study did not have the second most important component: the lipid profile. (Age is the most important risk factor.) That absence hamstrings the baseline model compared to the deep learning model. We aren't seeing the standard pooled risk equations here.
- As my PI argues on Twitter, you could see this deep learning model as a very good age predictor.
- The fact that the model is paying "attention" to blood vessels and other structures that are relevant to humans is a promising sign.
Why is that necessarily a good thing? The point of machine learning is to unveil relationships that we may or may not have thought of. For all we know there's a structure we've never looked at that provides better correlates with disease, and if a model pays attention to that it would be an equally promising sign.
I understand that in medical research it would be best to have a physiologically relevant correlate that is being recognized by the model, but it doesn't mean that we should completely disregard the other pixels that a model could find important either.
Wonder if smoking would decrease if people could see the change over time
¹well, unswayed about actually quitting now
However 71% seems like a very low number... Does that mean that in 7 out of 10 cases the system guesses correctly if the patient is smoker or not? That's not really impressive IMHO. What am I missing?
To put that in perspective, in this study, a random choice would be right 50% of the time.
So this is interesting, but would it be useful diagnostic tool?
But unfortunately there isn't anything close to this yet.
High blood pressure, for example, probably more accurate than the retinal image technique here. https://www.ncbi.nlm.nih.gov/pubmed/25632496
The major goal here would be to screen for people to go for a more advanced test, like an MRI. But even that wouldn't hit the 90-95% level of predictor.