This is causing problems for the first day of classes, see for example https://www.reddit.com/r/uofm/
754 karma · joined January 21, 2012
This is causing problems for the first day of classes, see for example https://www.reddit.com/r/uofm/
The good news is that once the library is prepared, it is quick to screen at more targets--and we make the pre-computed library available at zinc15.docking.org.
Interestingly, as the library grows a limiting factor is storing the library on disk. It is now ~20T. We've set up several mirrors around the world for groups that are actively using it. An interesting problem will be to see if preparing compounds for screening on the fly (e.g. with machine learning models) can overcome this limitation to keep up with library growth.
A big question for us is what will the return on investment in screening larger and larger libraries be? One of the take aways from this work is if docking has moderate enrichment, than screening larger libraries not only gives more hits but actually can increase the hit-rate for the top scoring compounds.
The key idea is to add a glucosyltransferase from Lactobacillus, which consumes simple sugars and produces complex sugars.
It is not clear to me that this will be beneficial for human health.
[1] https://patentscope.wipo.int/search/en/detail.jsf;jsessionid...
You bring up very good issues and perhaps I'm being too optimistic. I definitely agree that there isn't going to be one single mapping of sequence --> energy landscape any time soon or even ever.
But I think there are subproblems that are easier because the search space is more limited[1] or the chemistry is easier (e.g. avoiding chemical reactions or interactions with high energy fields). I think often the major modeling challenge is identifying when it is feasible to take advantage of problem constraints or when lower levels of theory can be used. For example there are a range of "enhanced sampling methods" for molecular dynamics that e.g. constrain the the simulation to a reaction coordinate or assume Markov transitions between states so they can be computed on a distributed cluster.
Taking advantage of these opportunities often requires a fair amount of engineering to build appropriate representations. I wonder to what extent these representations can be learned?
An alternative approach is to directly map conformational states to their free energy. This leads to a problem of searching for candidate conformational states (e.g. the folded state, transition states etc.) and scoring them. Usually for a given computational budget there is a trade off between better conformational sampling or higher accuracy energy scoring.
Historically, searching and scoring methods have been designed separately. For example [1] improves sampling while [2] improves energetics. This is done because they historically involved different aspects of the simulation and each is lot of work. But searching and sampling are not really separable, in that the deeper one samples the more challenging the task of the scoring function becomes--discriminating stable from unstable conformations.
Another application that can be thought of as searching and scoring is the game of GO. My impression is that one of the major breakthroughs with AlphaGo is that they were able to integrate models for searching and scoring together and learn the models simultaneously. It would be awesome if similar architectures could be applied to molecular modeling.
A remaining challenge in applying GO models to molecular biology is that while the representation and scoring rules for GO are fixed and quite easy, the ground truth for molecular simulations comes from heterogenous experimental data (X-ray crystal structures, small molecule activities, directed evolution antibody screens etc.) and higher levels of theory QM simulations, which have their own challenges. However, I think the principles carry over--complicated scoring functions (e.g. free energy) over large state spaces (e.g. protein conformation space or chemical space) can be learned by combining models for searching and scoring. I think deep learning is poised to tackle these problems.
[1] (Conway, et al., 2013, DOI: 10.1002/pro.2389) Relaxation of backbone bond geometry improves protein energy landscape modeling
[2] (Park, 2016, PMID: 27766851) Simultaneous optimization of biomolecular energy function on features from small molecules and macromolecules.
(Wallach, 2015, http://arxiv.org/pdf/1510.02855.pdf) AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery
(Duvenaud, 2015, http://papers.nips.cc/paper/5954-convolutional-networks-on-g...) Convolutional Networks on Graphs for Learning Molecular Fingerprints
(Kearnes, 2016, http://arxiv.org/abs/1606.08793v1) Modeling Industrial ADMET Data with Multitask Networks
(Gómez-Bombarelli, 2016, doi:10.1038/nmat4717) Design of efficient molecular organic light-emitting diodes by a high-throughput virtual screening and experimental approach
and of course
(Gómez-Bombarelli, 2016, https://arxiv.org/abs/1610.02415) Automatic chemical design using a data-driven continuous representation of molecules
http://deepchem.io/ is trying to set up standard data sets for chemoinformatics/machine learning.
ChEMBL and PubChem are the big public repositories though some care must be taken in curating data from these for machine learning.
For virtual screening it is possible to speed things up by say not taking into account receptor flexibility or ignoring explicit interactions with water.
As for lower resolution representations of small molecules, there is ROCS[1] and friends which represents small molecules with a set of gaussians.
One of challenges with low-resolution representations is that the aims of virtual screening is often to find novel backbones that may interact with the protein. So any low-resolution representation should mix different backbones into the same cluster, but finding such a representation is difficult, given the diversity of small molecules.
As for U47700, finding the mechanism of action for drugs that treat complex processes like pain is quite difficult. Also small molecules often interact with numerous targets so deconstructing how it works is non trivial. Part of the motivation for PZM21 is to try to separate out the downstream effects of hitting the mu-opioid receptor as a "biased" ligand. I think PZM21 with its new scaffold will help disentangle the effects of classical opioids.
Here is an example from our lab using virtual screening to develop PZM21 to treat pain [1]. where we screened 3M compounds. We would have liked to have screened 10^6 fold more compounds to cover easily synthesizable chemical space in this as well as other campaigns, but that is currently computationally infeasible. If molecular autoencoders could help us more efficiently screen this space, it would be huge.
I'm co-organizing a free, 1-day workshop for deep learning for chemoinformatics at Stanford Nov 11th. We've got ~75 mostly computational chemistry researchers coming. I would love to have more machine learning researchers come as well. The website is deepchemworkshop.docking.org, or PM if any of you think you may be interested.
[1] Manglik, et al. Structure-based discovery of opioid analgesics with reduced side effects (doi:10.1038/nature19112)