AI can diagnose childhood autism from retinal photos
petapixel.com
petapixel.com
I would bet that even physicians aren't 100% consistent in their diagnosis of autism. If that's the case, then it should be more or less impossible for any other diagnostic approach to be 100% consistent with the physician diagnoses.
Edit: After reading the study closer, this criticism might be a bit harsh. In their autism subjects, they excluded those with mild/moderate autism. Limiting to severe cases should mean there's a higher degree of confidence/consistency in the diagnoses.
[0] https://jamanetwork.com/journals/jamanetworkopen/fullarticle...
That's because autism is diagnosed by using the DSM. You can take an x-ray of an arm and see the fracture, but in order to diagnose autism you have to determine 'persistent deficits in social communication and social interaction across multiple contexts'.
It is all dependent on how society defines things, and is fluid (and IMO, somewhat dubious).
There are subfields of psych that actually function as sciences, like cognitive psychology (not the same as cognitive therapy). There are also some that are far less scientific: looking at you, "evolutionary" psychology! It's a strange mess that propriety says no one should talk about openly.
“The data sets were randomly divided into training (85%) and test (15%) sets. We used 10-fold cross-validation to obtain generalized results of model performance. Data splitting was performed at the participant level and stratified based on the outcome variables. Because the data classes were imbalanced for symptom severity (ADOS-2 and SRS-2), we performed a random undersampling of the data at the participant level before conducting data splitting. Moreover, we examined different split ratios (80:20 and 90:10) to assess the robustness and consistency of the predictive performances across diverse splitting proportions.”
* undersampling is problematic here and probably introduced some bias. These imbalanced class problems are just plain hard. Claiming one hundred percent on an imbalanced class problem should probably cause some concern. * data split at the participant level has to be done really careful or you’ll over fit * multiple comparisons bias by testing multiple split ratios on the same test data. Same with the 10-fold cross Val. * not sure if they validated results on any external test data * outcome variable stratification also has to be done really carefully or it will introduce bias; seems particularly sensitive in this case * using severity of symptoms as class labels is problematic. These have to really have been diagnosed the same way / consistently to be meaningful.
I also note a long time history in collection of these images (15 years iirc). Hard to believe such a diverse set of images (collection, equipment etc) led to perfect results.
ML issues aside, super interested in the basic medical concept. I wasn’t aware retinal abnormalities could be indicative of issues like ASD.
> The photography sessions for patients with ASD took place in a space dedicated to their needs, distinct from a general ophthalmology examination room. This space was designed to be warm and welcoming, thus creating a familiar environment for patients. Retinal photographs of typically developing (TD) individuals were obtained in a general ophthalmology examination room. Each eye required an average of 10–30 s for photography, although some cases involved longer periods to help the patient calm down, sometimes exceeding 5–10 min. All images were captured in a dark room to optimize their quality. Retinal photographs of both patients with ASD and TD were obtained using non-mydriatic fundus cameras, including EIDON (iCare), Nonmyd 7 (Kowa), TRC-NW8 (Topcon), and Visucam NM/FA (Carl Zeiss Meditec).
So two questions:
1. Are we positive that the difference in rooms does not effect these images?
2. If we are in a dark room, and ASD patients are in it for 5-10 minutes longer, are we sure this doesn't effect the retina?
3. Were all cameras used for both ASD and TD images?
Want to make sure the AI is being trained to detect autism, and wasn't accidentally trained to identify camera models, length-in-dark-room or room-welcomingness.
Hopefully not, but I assume you have to be so careful with these sort of things when the model is entirely black-box and you can't actually validate what it's actually doing inside.
Just being in a dark room longer is sufficient to make changes that an AI could pick up on.
Ideally, they should capture the images from children before diagnosis, then see if they can predict the diagnosis.
Also, the study checked ASD participants were autistic by using structured interviews with psychologists against the DSM-5, but the TD participants were never assessed by psychologists, so if autism under-diagnosis is a thing, there could theoretically be false-negatives.
If a model was 100% accurate, considering the nature/accuracy of manually diagnosing autism you would probably expect the AI to either find new cases or identify a few incorrect diagnosises.
There was great excitement as it was near 100%.
It later transpired the pictures with submarines in had a white border.
Overfitting is, AIUI, a training method and data issue, not a model issue alone. I doubt any model is resistant to overfitting if you give it data where the answer is reliably encoded some aspect it can use but outside of what you want it to look at.
Now, you can notice suspicious results and investigate (or you can just publish a 100% success rate and call it a day.)
See the paper, "On estimating model accuracy with repeated cross-validation"
Even the sklearn docs for cross validation show this split: https://scikit-learn.org/stable/_images/grid_search_cross_va...
I remember stumbling upon multiple esoteric accounts on both Twitter and Tiktok with communities seemingly obsessed with characterising various psychological traits from purely looking at facial features, importantly without racial undertones.
While this on the surface sounds ridiculous and has various horrible historical echoes, i've always had a hunch there was actually something to this science from a purely intuitive perspective and knowing lots of people - again very importantly disregarding anything about race - instead focusing on the myriad of hormone linked features, neurotypicalism, alcohol, environmental factors, whatever traits that seemingly somehow go "across races".
Or maybe there's nothing there.
Autism is 4 times more common in boys than girls, and women are diagnosed with autism later in life and less frequently than menA few years later I met my wife's nephew, who was about 5 or so - and he was constantly spinning things... and I asked if he was autistic, and they said no.
A few month passed by, and he was officially diagnosed as autistic.
I wonder if something in the retina, and the way that it signals to the brain is soothing if there is a spinning image signal coming through.
I wonder if one were able to apply a HUD either in a contact or such, where there might be a slightly spinning halo/ring that one might look through, if it is that spinning things sooth an autistic signal processor. There are two different spin stimuli it appears that autistic people find soothing, visual or physical (spinning of themselves).
"" The data sets were randomly divided into training (85%) and test (15%) sets. We used 10-fold cross-validation to obtain generalized results of model performance. Data splitting was performed at the participant level and stratified based on the outcome variables. Because the data classes were imbalanced for symptom severity (ADOS-2 and SRS-2), we performed a random undersampling of the data at the participant level before conducting data splitting. """
100% indicates a major case of data leakage. Particularly, "When we generated the ASD screening models, we cropped 10% of the image top and bottom before resizing because most images from participants with TD had noninformative artifacts (eg, panels for age, sex, and examination date) in 10% of the top and bottom." tells me that there are known issues with the photos. Quite possible the photos of the groups were taken on distinct days and/or times and the lighting conditions is enough to distinguish the groups. Possibly, the background for the ASD candidates have a different background or a different camera sensor.
On one hand if it's real that's huge, on the other hand I wouldn't take this test seriously until the AI has been tested with pictures of children 1) taken in exactly the same conditions 2) the picture should be taken before the children has been diagnosed (this way you truly have a double blind study).
https://jamanetwork.com/journals/jamanetworkopen/fullarticle...
In addition to demonstrating 100% (n = 59/59) sensitivity for detecting melanoma, the AI software had a detection rate of 99.5% (n = 189/190) for all skin cancers and 92.5% (n = 541/585) for precancerous lesions.
I mean, you take a data set of how people with X health condition look like, train your model on that set, then the model can, with a certain probability threshold, tell you if other people have X condition as well. But to me, a non-AI average joe programmer, that's more pattern matching, and less AI.
That kind of application was something I remember some colleagues in university were working on over 10 years ago in OpenCV and Tensorflow, using labeled sets of X-rays and CT scans to pattern match conditions for aiding in diagnostics for radiology lab techs, well before the AI explosion.
What am I missing here?
About 15 years ago I learned about several of those topics in an AI course in college, and now they're all frequent occurrences in debates about whether they're really AI.
For years most people agreed that a computer beating the best human at chess would be "real" AI. Then they managed it, but people said "not that does not count, you just brute forced it".
What you describe, the automatic construction of classification models was textbook AI for over 20 years (using a whole range of methods), now it's often just seen as "data science".
Perhaps how the pattern matching is performed?
Whether this is "AI" or not, I don't know. But dismissing it as mere "pattern-matching" grossly simplifies the technology and it's potential usefulness.
>I remember colleagues in university were working on over 10 years ago
Yes, this is how we get better technology, by continuing to work on problems.
I would disagree with this statement. The model is _reasoning_ via the method of processing the data since the weights are informed from the relationships of prior similar information, but at its heart the model is still just doing pattern matching. The primary distinction and what I disagree with in your statement is that these AI models are closer to AGI because there aren't a bunch of if/else statements or basic pattern matching.
Can you at least interact with it and have it explain to you its reasoning like "so my diagnostic is based on seeing a lump in the lower right hand corner"? That would be an innovation indeed.