Using Deep Learning to Predict the Olfactory Properties of Molecules
ai.googleblog.com
ai.googleblog.com
- They created an entirely new dataset of ~5000 molecules, all hand-labeled by perfume experts.
- They held a competition (presumably Kaggle or a similar platform) to classify this dataset, and used the results as a strong baseline.
- Their GNNs get comparable (slightly better, but not statistically significant) results than the winning random forest model of the competition.
The embeddings show promise, but I'm curious why they omitted a simple "fully connected layer" on the Morgan bit descriptors as a baseline classifier. Seems like that would outperform the random forest.
[0] https://pytorch-geometric.readthedocs.io/en/latest/notes/cre...
[1] https://pytorch-geometric.readthedocs.io/en/latest/modules/n...
Maybe we're missing the most interesting aspect. Olfactory Receptor Genes in humans comprise ~1% of the total genome. The benefit here is in understanding how environmental changes trigger beneficial mutations and enhance sensory features.
It is two different things – you either recognize a scent OR you identify a scent. Why not speed things up by running it as a two step process? Recognize = Rapid results, Identify = Best Guess at what it might smell like. I’m not sure how a human determines that there is a smell of gasoline, freshly cut grass and a hint of something else/unknown in the air rather than thinking hmmm, I do not recognize a scent that has aspects of cut grass AND gasoline AND an unknown therefore the whole scent is ‘unknown’.
*anon-experts that go ‘gosh, why didn’t those idjuts with multiple PhDs think of that’
The real issue here is that olfaction involves a small molecule interacting with multiple different proteins for your body to "read" it. We're not so great at the whole predicting drug-protein interaction thing just yet - some closely related fields include virtual screening, molecular dynamics, protein crystallography, X-ray diffraction, and cryoEM (among others).
Yes, it's good for companies to do some R&D, and sometimes the R part of the R&D gets pretty theoretical, which is usually a good sign that a company is trying to really innovate.
But usually also there's some indirect way to tie it back to some kind of application that could possibly somehow make the company money in the long term. Otherwise, you're just a for-profit company spending money on something because it's interesting.
So what's the application here? Are there novel ideas or techniques here that can be applied to other AI problems? Is there some kind of application for smell in a Google product?
So while the scent identifying doesn’t seem directly relevant, nailing the underlying techniques is valuable.
(complete speculation, but seems plausible)
For example, I ran an idle-cycle-harvesting service at Google called Exacycle that ran problems like protein folding, protein design, drug discovery, telescope discovery, and more. The only pushback I got was to run problems where the results could reasonably be considered "useful" (scientifically).
One way to think about it is that many of the people with power at Google really like science and have the resources to support it. Once you've built things like TPUs, it would be a waste not to dedicate some about of their resources to problems that people wouldn't be able to address.
Another way to think about it is these things have indirect effects- even if Google didn't want to make some sort of product with smell (like a phone with a builtin GC-MS?), publishing this work gets the attention of scientists, who will then read the paper, and consider Google Cloud as a place they'd like to do their work.
Google wants to help the world by reducing our reliance on farming animals for meat, and thought that research like this would help humanity's nascent artificial meat endeavours.
If you wanna get Black Mirror-esque, perhaps a Soma-like medication from Brave New World (essentially pacifies/zombifies you by creating endless bliss) could be made. Or the "bliss" drug episode of Doctor Who.
Shulgin gives ratings to compounds in PiHKAL. Those could be used as well.