Regarding external validation set:
> This dataset differed from the UK Biobank development set with respect to both fundus camera used, and in sourcing from a pathology-rich population at a tertiary ophthalmic referral center.
Regarding UK Biobank set (training set)
> UK Biobank dataset, which is an observational study in the United Kingdom that began in 2006 and has recruited over 500,000 participants—85,262 of which received eye imaging38. Eye imaging was obtained at 6 centers in the UK and comprises over 10 terabytes of data39. Participants volunteered to provide data including other medical imaging, laboratory results, and detailed subjective questionnaires.
So the bottom line is that ironically due to this science gets stuck on false beliefs and they can be a bit hard to dislodge.
We have even this principle: https://en.wikipedia.org/wiki/Planck%27s_principle
I pulled a third sample from a completely different time window and it performed terribly.
It turned out that both datasets were dominated by class A being sorted into always selecting great fit or poor fit, so the ML model learned to memorize the class A instances.
This problem when away when I subselected down to only instances of class A that had examples of both good fit and poor fit.
My problem with modern scientific publications is that they focus more on the "discovery" instead of describing the logical rigor as to why their "discovery" could be true or false.
The paper in this post itself has its own Limitations section.
I haven't actually read the paper. I was reacting to the title.
> what kind of future research could help expound on the pitfalls of the current research.
I think a better way of documenting research to people is by describing what scientific boxes where checked.
In the given case for example one of the boxes maybe,
"Is there an anatomical distinction between retinas of sexes?" — If that is true then we can see if machine learning can detect such differences.
Take computer science for example. A publication with the title "Professors create a machine that can think" would maybe published as "Professors create machine that passes Turing test [0]"
Another example, can be that is medicine.
The research in question maybe if a microbe A causes a flue.
Instead of a publication being "Microbe A causes disease B".
A better Publication IMO should revolve around, "Microbe A had passed Koch's postulates [1] for disease B"
Once upon a time—I’ve seen this story in several versions and several places, sometimes cited as fact, but I’ve never tracked down an original source—once upon a time, I say, the US Army wanted to use neural networks to automatically detect camouflaged enemy tanks.
The researchers trained a neural net on 50 photos of camouflaged tanks amid trees, and 50 photos of trees without tanks. Using standard techniques for supervised learning, the researchers trained the neural network to a weighting that correctly loaded the training set—output “yes” for the 50 photos of camouflaged tanks, and output “no” for the 50 photos of forest.
Now this did not prove, or even imply, that new examples would be classified correctly. The neural network might have “learned” 100 special cases that wouldn’t generalize to new problems. Not, “camouflaged tanks versus forest”, but just, “photo-1 positive, photo-2 negative, photo-3 negative, photo-4 positive…” But wisely, the researchers had originally taken 200 photos, 100 photos of tanks and 100 photos of trees, and had used only half in the training set. The researchers ran the neural network on the remaining 100 photos, and without further training the neural network classified all remaining photos correctly. Success confirmed!
The researchers handed the finished work to the Pentagon, which soon handed it back, complaining that in their own tests the neural network did no better than chance at discriminating photos. It turned out that in the researchers’ data set, photos of camouflaged tanks had been taken on cloudy days, while photos of plain forest had been taken on sunny days. The neural network had learned to distinguish cloudy days from sunny days, instead of distinguishing camouflaged tanks from empty forest.
"""
When I was tasked to create stimuli or to take measurements on a series of items it was always important to try to eliminate systematic differences on non-interest, most typically thru randomization of the order of creation or measurements.
> For instance, in our work, we noted that the algorithm appeared more likely to interpret images with rulers as malignant. Why? In our dataset, images with rulers were more likely to be malignant; thus the algorithm inadvertently “learned” that rulers are malignant.
But still, there are ways of getting the resulting NN and applying some techniques to figure out which variable/patterns are responsible for most of the output.
Scientist: How do you know that?
NN: I just do, trust me.
Most things in science is almost impossible to reproduce because of cost or specialized equipment.
Is that what reproducibility means?
I think a good control would be to see how this model treats trans people at varying stages of transition
But indeed it would be interesting to see how trans people are treated at all.
Estrogen levels affect the eyes, so given enough samples and context it may be surprising accurate for people getting that type of therapy.
For a ML model to correctly guess gender identity, it'd need other cues that indicate a person would prefer to be referred to by a certain pronoun, such as clothing or facial hair.
Women's are statistically longer.
Their eyes are also more almond shaped.
Retinal photographs are hard to take. Often they contain significant amount of eyelashes and surrounding structure