Atomwise (YC W15) Discovers Drugs for Diseases That Don’t Even Exist Yet
techcrunch.com
techcrunch.com
Our goal is to bring to medicine discovery the same kind of incredible efficiency gains that computation gave us in aerospace and mechanical engineering design. Today, people have to physically synthesize and physically test molecules to figure out how they’re going to behave. That’s incredibly laborious, expensive, and time consuming. We’re able to get the same results, but in days instead of months or years.
Given all of the new and re-emerging diseases we’re encountering (such as Ebola, measles, malaria, and drug-resistant infections, to name a few that we’ve worked on), I think our species needs all of the help we can get in finding new medicines. I’m happy to answer questions about what we’re doing, or the challenges we encounter when we take deep learning algorithms out beyond image classification.
Is this newfound excitement and interest in AI fuelled by actual, objective advances?
EDIT: Keep up the good work, I sincerely hope that the work you do will make the world a better place. I wish I could help somehow.
I hope your team succeeds, keep up the hard work!
Over the past few years, there's been a huge increase in the amount of data available for this kind of machine learning. We curate our data from a number of private and public sources. For example, as part of my doctoral work (http://en.wikipedia.org/wiki/SCRIPDB), I learned how to parse chemical information out of U.S. Patent data, which is public domain. That said, if you're interested in working on something like this and need a quick million data points, I'd point you to PubChem as a first step: https://pubchem.ncbi.nlm.nih.gov/
My understanding of D.E. Shaw's approach is that they're doing molecular dynamics, i.e. simulation. You get to watch the motion of every atom in the system. That allows for a close investigation of a given protein's movements, which is great especially if you're trying to learn about its biology. Unfortunately, it is rather computationally expensive; while I don't know DESRES's latest stats, I've seen reports on large parallel MD simulations completing about once per day.
In contrast, we've posed the question of binding as a machine learning problem. Neural networks are computationally expensive to train, but make predictions quickly. Our system can assess millions of protein-drug pairs per day, since we're not simulating the motion of every atom. You don't get to watch what each atom is doing, but you get insight into the behavior of lots of potential medicines.
This is the kind of work that is essential for our future --so thank you. I'm quite (positively) surprised Sam decided to fund this; we definitely need to put more effort and resources as a civilization to work like this.
I'll hopefully have a lot more to say in the future, and will definitely be reaching out in a more substantive way... But in the meantime, a quick word of advice: This will sound strange, since Hinton's work is quite powerful as it is, but my guess is that good-old boosting / \ell_1-regularized ensemble learning methods would work much better for this particular problem domain --so please run some experiments and look into it, if you haven't already. It's hard to find good and up-to-date literature on this (nowadays) less fashionable work (a good rule of thumb: if it mentions 'random forests', it is not well-informed enough), but Freund and Schapire's recent book [1] is self-contained and a jewel to read back-to-back. Best of luck.
With respect to boosting, we have more investigation to do, of course; the tricky issue with the biological domain is that we know the underlying data is incredibly noisy. How to walk the line of extracting maximum predictive performance without overfitting is the challenge, since we know that a lot of the raw data points are unreliable. Any algorithm we use has to be able to handle this scenario deftly.
[1] http://research.microsoft.com/en-us/um/people/adum/publicati...
What kind of information actually goes into your neural net to predict binding affinity? How does your model actually work?
I've some experience with MD simulations; even these, which have had 1000s of man-hours of parameterization, are often not very accurate. I'm curious what you are using to evaluate your model's predictions.
How are you deciding what your targets actually are? How do you relate a particular drug target to a disease in the case of complex phenotypes? This seems to me to be by far the most challenging part of pharmacology.
I'm definitely enthusiastic about applying computation to biology, but as someone also on the frontlines, it's definitely not a clean analogy to computation in engineering. Biological systems are far more complex, interconnected, and nonlinear which is something that is sometimes underestimated.
I got to work in the Baker Lab that developed this: http://www.ncbi.nlm.nih.gov/pubmed/22267011
The typical input to the neural network is the 3D structure of the molecule and of the protein. The model works by detecting patterns in the pair of protein and the drug that correlate with binding, e.g. hydrogen bonding, halogen bonding, cation-pi, pi-pi interactions, etc. But these are complicated to encode manually, given all of the factors that affect binding strength: distance, angle, water mediated effects, resonance, (de)stabilizing environmental charges, etc. That's why we need the neural net: you can think of it as the network automatically deriving the best pharmacophoric features to maximally explain which training examples bind and which ones don't, and then the prediction step is looking for the presence or absence of those patterns in new protein-ligand pairs.
We evaluate our models both retrospectively and prospectively. For example, the DUD-E benchmark (http://dude.docking.org/) gives us an assessment of our performance over more than a million individual predictions, comprising many diseases and many biological classes (GPCR, nuclear receptor, enzyme, etc). It begins with 102 disease proteins and, for each one, has a set of molecules that bind to the protein and a set that don't. We shuffle those sets together and ask the neural net to "pick the aces out of the deck". Separately, we perform prospective evaluations, for settings where no one knows the right answer, and run the experiment to confirm the predictions.
I agree with you that the proper selection of targets is critical, as is the mapping between drug target and disease. For us, however, this is easy: we work with smart biologists! If you have any, please send them our way!
Finally, I agree with your point that biology is not designed to be understood by people. That said, molecular binding is fundamental enough that we could think of it as an example of physics rather than biology. And theory works so well for physics that, in many a physics lab, if an experiment disagrees with theory then the first step is to double-check the experiment for errors. The trick is to scale that up to larger systems. Semi-relevant: http://www.smbc-comics.com/?id=2272
Also, as I described in my above answer to et2o, we do large retrospective tests to evaluate our predictive accuracy.