Large-Scale Machine Learning for Drug Discovery
googleresearch.blogspot.com
googleresearch.blogspot.com
The big thing will be if it helps prevent compounds dying in the clinic, as such late in the day failures are the most expensive. ADME/Tox prediction would be a huge win, although QSAR/QSPR has been fighting this war basically forever. The companies have the data while the researchers have the methods and bodies - getting these two entities to work together have prevented success in the past. Although the data is not as good as one would hope, particularly after the merger mania in pharma have mixed up all the datasets.
Edit: They published a whole book about it in 1999: Neural Networks in Chemistry and Drug Design (Wiley, probably $$). Also thought to add that I have no relationship to his work.
It's great to see that competition's result both supported and extended!
So, what's actually novel in this paper?
The AUC improvements they show are fairly modest. Training on millions of data points is pretty old hat in the neural network world.
ICML is not a molecular biology journal and can't judge this paper on those merits (and probably neither can most people on news.yc). From a machine learning perspective, though, there isn't much here.
Check out my talks at gravesmedical.com and let me know if you are interested
I contributed my whole genome sequence/variants a while ago.
On a related note, Atomwise has been doing this for over a year. And they've been able to make predictions that have been tested in the real world in areas such as multiple sclerosis, C. Difficile, and Ebola.
Full disclosure: They are my YC batch mates.