Building an AI to predict human age from a blood sample
colekillian.com
colekillian.com
The goodness of the result of a machine learning model like this one should be compared with the goodness of a simple "standard" model like linear regression.
Yeah it is kinda cool that we can use 10 lines of TF to spin up huge computation, but I guess that a simple linear regression would have provide results that are at least similar to the one of the neural network.
It’s sort of like production software in general. Sometimes the core product is pretty easy to prototype... but serving it to all your users with very high reliability and uptime is not so easy, and that’s what actually gets them to pull out the credit card.
Plus, LR is not a black box, so it both brings you a class and a reason why that class was chosen, which is a very desirable property in many problems.
It's more fun to use a neural net :) but after many similar comments I plan on implementing a simpler approach and seeing how it compares
Proceeds to use (1024^2 * 2 + 1024) parameters in the neural network.
/s
/s for this post, I mean. I have had this very suggestion made to me non-sarcastically under similar circumstances.
I just thought I'd highlight a bit of funniness.
As the author is using RELU, he will have a decent number of neurons 'die'. So some 'over-provisioning' is not a bad idea in theory. Also, if my keras isn't so rusty, I think the author is using less parameters than you are stating.
Still a more reasonable dropout rate and maybe some regularization/batch normalization might help, but I would say not over fitting on only 700 samples is a hard task, even with a network much smaller than that.
I would be pretty shocked if this neural net wasn't over fit
[1]-https://towardsdatascience.com/pruning-deep-neural-network-5...
after many similar comments I plan on implementing a simpler model and seeing how it compares
I've used it myself for fun after doing a blood test. It's a free alternative to InsideTracker's InnerAge product.
Over time, cell populations with BCR/TCR that recognize and bind such antigens will cease proliferation. Moreover, some cell populations will be localized to certain tissues and not in circulation.
Our startup (Chronomics) has built the most accurate epigenetic clock from Saliva (no needles..) which looks at 20 million positions (or features) https://www.chronomics.com/science
Really interesting area and we are starting to be able to define many more novel indicators of actionable health risks such as smoke exposure, alcohol consumption and metabolic status from DNA methylation.
You can now find the jupyter notebook code here: https://github.com/Ruborcalor/Age-Prediction-Via-Blood-Sampl...
I'll try and get back to you with the performance of k-folds validation and shuffling.
I don't think it can be learning the order or samples because the train and test data sets are separated very early on. If it were learning order or samples of the training set it would have to perform very poorly on the test set.
On a separate note, I think there may be a source file missing in your notebook. I kept getting an error when trying to load "GSE87571_series_matrix.csv". Might just be me.
[sklearn ref](https://scikit-learn.org/stable/modules/generated/sklearn.mo...)
[tf.keras](https://www.tensorflow.org/api_docs/python/tf/keras/Model#fi...)
Also, with so few samples, how do you do your hyperparameter tuning and validation?
I mean you could eliminate certain features in isolation but that doesn't capture dependent features. And how would you do dimensionality reduction?
Honestly I didn't prioritize hyperparameter tuning enough. I pretty much went with one of the first models I identified.
Could you elaborate on the idea of not capturing dependent features please?
If people want to learn more about how DNA methylation relates to aging, I recommend reading Lifespan by David Sinclair.
[1] https://www.semanticscholar.org/paper/DNA-methylation-aging-...
Thanks for sharing the book i'll have to sheck it out!
> The data was then split into training and testing sets at a ratio of 9:1, and fed into a sequential neural network.
What?? I thought it was split already.
(How training and test sets were obtained sounds fairly confusing. Did the author make sure there's no "data snooping" ?)
The data is only split once, before using a correlation test to select the features that the model would be trained on. As far as I can tell there is no data snooping occurring because the data is split into train and test sets before any decision are made.
https://www.dailymail.co.uk/sciencetech/article-3349739/Woul...
Interesting article thanks for sharing.
That is, of course, in addition to the all-encompassing family trees we're providing them with 23andme.
I'm not sure I understand; what was the claim from the startups, and what does the rumor mill say is happening at a low level?
I'd be really interested to see how well a baseline linear model using those features would perform - it seems like it could do pretty well.
There was a paper last year or so that compared correctly tuned linear models to various deep belief net papers and found that the performance "gains" suddenly evaporated or were not nearly as great as originally published.
If I can track down that paper, I'll post it.