Predicting Properties of Molecules with Machine Learning
research.googleblog.com
research.googleblog.com
medchem is notoriously NOT generalizable. A crude example is the reason why the developed heroin is because in the early days of medchem, the reasoning was acetyl-salicylic acid is awesomer than salicylic acid, so therefore acetyl-morphine must be awesomer than morphine. Actually, in many ways it is awesomer (and that's why it's a bad drug).
Consider Gleevec. Even if you knew the structure of gleevec's target (BCR/ABL) you would not be able to predict Gleevec, because it works by displacing an entire segment of the protein out of place which happens to be thermodynamically more stable (but kinetically disfavorable). Gleevec is a medchem drug (discovered through combinatorial synthesis) but sadly the insight into this mechanism is only generalizable in the conceptual sense, if you take that molecular fragment and graft it onto another molecule intended for a different target, it probably won't work.
Deep learning depends strongly on generalizable knowledge, and medchem is notoriously not easily generalizable for well-understood reasons.
Some aspects of medchem - like optimizing bulk synthesis reactions, picking synthetic routes, guessing at bioavailability, stability in formulations, might be amenable to ML, but I am not bullish on discovery. Let's hope I'm wrong.
It's not to say ML has no value, but predicting molecular behavior, even in the simplest system is really dam hard.
Wheb you only under 10% of the factors influencing behavior, ML doesn't get you far.
I remember working with some computational scientists:"just put a methyl group on this nitrogen and we should increase binding by 100x!".
So we make the molecule, give it to the biologists and find out binding is actually 1000x worse!
Though property predicting is a hard problem,I think there are low hanging fruits in other fields. For example, Anthropology where only partial skeletons are found but we know there is symmetry there. Software regeneration is slow and expensive and doesn't exploit the symmetry a lot.
A joint project between google and CERN also sounds really cool to me. Or maybe google can set up a system where researchers with large data can approach google and see if a symbiotic relationship can be formed.
You do know that CERN, regularly publishes papers in machine learning , right?
Founded by a former professor from CERN, and staffed about 90% from CERN postdocs. I was the only member of my team who was not a co-author on the Higgs boson discovery paper.
So yeah, people at CERN are pretty well aware of what can be done with ML.
Quantum is a special case really as DFT computation is already very time consuming.
The main problem is that SAT problems come in many different sizes, and while resampling a picture "makes sense" to make it smaller, resampling a SAT problem does not. Also the general lack of shape or structure makes things hard.
NN can help bad SAT solvers get better, but the heuristics in the best ones are (at present) better than anything a NN can produce.