AI is changing chemical discovery
thegradient.pub
thegradient.pub
A key challenge: very few labs have enough data.
Something I view as a key insight: a lot of labs are doing absurdly labor intensive exploratory synthesis without clear hypotheses guiding their work. One of our more useful tasks turned out to be interactively helping scientists refine their experiments before running them.
Another was helping scientists develop hypotheses for _why_ reactions were occuring, because they hadn't been able to build principled models that predicted which properties were predictive of reaction formation.
Going all the way to synthesis is nice, but there's a lot of lower hanging fruit involved in making scientists more effective.
Would love to chat more with you about this.
Chemists can just try out things.
I don't think you can compare the two.
Not to generate ideas, there's always more ideas than resources in chemistry.
Mainly to do more automated routines than ever.
9 out of 10 chemists aren't that great at the bench anyway.
Everyone would probably benefit from getting them in front of a computer full-time to leverage their training in a way, and freeing up the bench space to those who can really make the most of it.
You might just decide that a compound "needs" isopropanol/acetone, plus a bit of water, cause something vaguely similar you encountered years ago crystallized well. You often start with some educated guesses and refine based on what you see.
But there's often no clear hypothesis, no single physical law the system obeys.
Also shameless plug: I started a company to do just that, anchored to generating custom million-to-billion point datasets and using ML to interpret and design new experiments at scale.
It is also getting harder, not easier, to get.
I am working right now on a retro synthesis project. Our external data provider is raising prices while removing functionality, and no one bats an eye. At the same time our own data is considered a business secret and therefore impossible to share.
As someone who does NLP research where the code, data and papers are typically free, this drives me insane.
This lets you stumble over unknown unknowns. Taylor et al discovered high-speed steel by ignoring the common wisdom and doing a huge number of trials, arriving at a material and treatment protocol that improved over the then-state-of-the-art tool steels by an order of magnitude or more. The treatment mechanism was only understood 50-60 years later.
Finally: every kid can draw up novel structures. Then: how do you actually fabricate these (in the case of real novel chemistry and not some building-block stuff). Noone has a clue!
I for myself have decided that for now (with the data at hand and non-Alphafold-budgets) the 2 keys areas, where you can actually help computational chemistry are:
- creating really robust and generally applicable ML-MD-potentials, potentially using graphs https://arxiv.org/abs/2106.08903 (or a traditional approach: https://www.nature.com/articles/s41467-020-20427-2). Facebook is also working in this area: https://pubs.acs.org/doi/10.1021/acscatal.0c04525
- and approximating exchange correlation functionals (... Google and some guys at Oxford, which got stomped over by the deepmind-PR machine https://arxiv.org/pdf/2102.04229.pdf): https://www.science.org/doi/10.1126/science.abj6511
If anyone can tell me how those generative models spit out graphs which look like reality (actually this is imho part of AlphaFold), wake me up.
Yep. I worked at a biotech startup in the early/mid 2000s.
We had a 2-pronged approach to finding small molecule drugs: 1) traditional medicinal chemistry based on simple SAR (structure-activity relationships) and 2) predictive modeling (before ML was hot).
The traditional med chemists were, in my opinion, rightfully skeptical of the suggestions coming out of the predictive modeling group ("That's a great suggestion, but can you tell me how to synthesize it?").
As one of my co-workers said to me: "The predictions made by the modeling group range from pretty bad to ... completely worthless."
It's possible that things have gotten better, though, as I haven't done that type of work since about 2008.
Start by plugging it into askcos.mit.edu/retro/ then, do your job?
> As one of my co-workers said to me: "The predictions made by the modeling group range from pretty bad to ... completely worthless."
Workers feeling threatened by technology think the technology is bad or worthless, news at 11.
Of course, but we are talking about chemical discovery - after which you want to test if the compound theoretical capabilities work on cells. Yields are not yet a concern!
> expect any reasonable results
No it won't do all the work, but it will direction, and suggestion for which pathways could be used.
I appreciate the sentiment and I think it's understandable to think that. In this particular case, however, my co-worker was one of the smartest / most talented people I've worked with. I can assure you that he did not feel threatened in any way. His comment was sardonic, but not borne of insecurity.
To be fair, the members of the modeling group were also quite talented. They were largely derived from one of the more famous physical/chemical modeling groups at one of the HYPS schools. But even they acknowledged that on a good day, the best they could do was offer suggestions / ideas to the medicinal chemists.
In fact, one of the members of the modeling group said this to me once (paraphrasing): The medicinal chemists are the high-priests of drug discovery. We can help, but they run the show.
As mentioned by someone who responded to my original comment, the usefulness of ML/modeling has likely gotten much better over the past 10 - 15 years.
With high-throughput screening and automation, even small/medium-sized players can start building internal databanks for multi-objective models.
I personally have a clue, and the entire field of organic chemistry has a clue, given enough time and money most reasonable structures can be synthesized (and QED+SAScore+etc and then human filter is often enough to weed out the problem compounds that will be unstable or hard to make). Actually even some of the state of the art synthesis prediction models are able to predict decent routes if the compounds are relatively simple [0]. The issue is that in silico activity/property prediction is often not reliable enough for the effort to design and execute a synthesis to be worth it, especially because as typically the molecules will get more dissimilar to known compounds with the given activity, the predictions will also often become less reliable. In the end, what would happen is that you just spend 3 months of your master student's time on a pharmacological dead end. Conversely, some of the "novel predictions" of ML pipelines includign de novo structure generation can be very close to known molecules, which makes the measured activity to be somewhat of a triviality.[1] For this reason, it makes sense to spend the budget on building block-based "make on demand" structures that will have 90% fulfillment, that will take 1-2 months from placed order to compound in hand and that will be significantly cheaper per compound, because you can iterate faster. Recent work around large scale docking has shown that this approach seems to work decently for well behaved systems.[2] On the other hand, some truly novel frameworks are not available via the building block approach, which can also be important for IP.
More fundamentally, of course you are correct, and I agree with you: having a lot of structures is in itself not that useful. Getting closer to physically more meaningful and fundamental processes and speeding them up to the extent possible can generate way more transparent reliable activity and novelty.
[0] https://www.sciencedirect.com/science/article/pii/S245192941... [1] http://www.drugdiscovery.net/2019/09/03/so-did-ai-just-disco... [2] https://www.nature.com/articles/s41586-021-04175-x.pdf
We've done work in this area and will be publishing some results later in the year.
Have a look at how driving a taxi has changed. (And I include the likes of Uber here.)
The hard part for humans used to be knowing all the roads in a city and selecting the best route quickly and reliably. Almost any adult can do the actual second-to-second driving reasonably competently in almost any city on the planet.
Now, we have outsourced the 'hard' part to Google Maps. But we are still far from a machine that drives in arbitrary locations on the globe. As far as I know, Waymo has the most mature system currently in development, but requires absurd amounts of precise mapping data for any location they want to drive in.
And let's not even talk about the even more 'trivial' task of chatting to the passengers.
> [...], but it doesn't seem to be very good at science/engineering (drug discovery/self driving cars/radiology).
Technology has already automated huge chunks for science and engineering. We just don't call any of the already solved chunks by the name of AI anymore.
OP describes what could potentially be AlphaFold with small molecules instead of with just proteins, and specifically calls it “chemical discovery” and not “drug discovery” more broadly. I think it’s fair to cheer for such a big advance while recognizing it wouldn’t solve everything.
[1]: https://www.science.org/content/blog-post/more-protein-foldi... "More Protein Folding Progress - What's It Mean?"
[2]: https://www.science.org/content/blog-post/alphafold-exciteme... "AlphaFold Excitement"
Support vector networks were searching for good shapes by 1997, maybe before then, but the vocabulary we used to describe the search technique was different.
A huge computational hurdle was protein folding. We had brute-force searches for plausible shapes, lots of supervision by the chemists, weeks per iteration on the best workstations we could get. $250,000+ SGI Onyx, then DEC sent over an Alpha workstation.
It's come a long ways.
Why do you need a 3070/3080 specifically? If it's to run something like Tensorflow or CUDA code more generally, could you do it with an older card, or the more available 3060s?
This is false.
You can easily find thousands of brand new Nvidia 3070/3080 GPUs online.
The problem is, you wish to pay MSRP. Supply and demand doesn’t work in your favor here.
In general: price competition might be effectively outlawed in some circumstances.
Eg you can't really buy a replacement kidney for any amount of money (outside of Iran).
[1] - https://buildredux.com/
About CPU vs GPU: things seldom work out on first try. A GPU gives you much lower latency for trying out a series of ideas.
[1] https://scholar.google.com/scholar?q=materials%20project%20h...
I know this doesn't necessarily apply, but if the solution space for certain niche problems is so small that we can just drown the problem in compute, I couldn't care less that the algo was N^2 or whatever, or that the UI was less than ideal. Maybe I'm not thinking deeply enough about what you mean by "better software solutions".
If I run a batch job and it spits out 10,000 compounds that I can try to a certain affect, then it then becomes a filtering problem where I can apply humans and do more traditional science, and if it was feasible to just try everything in parallel that option is nice as well, feels like how you got to that 10,000 compounds doesn't matter much.
Looking forward to hearing just how wrong my simplistic view is.
Not sure if they mean 1060 factorial or 10 to the 60.
I am not sure that is very true at all.