One question is how robust the predictions are that DeepMind produces. I would also assume that right now it can't e.g. determine protein structures in the present of other small molecules, or protein complexes. A lot of the interesting stuff lies in the interactions between molecules.
And in general in life sciences any new development will take at least a decade until it hits day to day life, likely even more. We're living with a exception to this rule right now due to the pandemic, but in general things take quite a bit of time in that space.
What an accurate model of protein folding allows us to do, is to take our big database of DNA, predict protein foldings for all of it, and then stand up a search index for this database, keying each amino-acid "row" by the "words" of its predicted protein's structural features.
We could then, with a simple search query that executes in O(log n) time, find DNA targets that produce molecules with interesting structures that might be worthy of study.
This would, for example, be a game-changer in how biopharmaceutical macromolecule-therapy R&D is conducted. Right now we have to notice that some bacterium or another produces some interesting protein, and then engineer a bioreactor to get more of that protein. With this tech, we can work backward from an entirely hyothetical, under-specified "interesting protein", to figure out what catalogued-but-unstudied DNA sequences produce never-before-catalogued proteins that fit that particular functional "shape", and therefore might do the interesting thing. Then we can either directly synthesize that same DNA, or find the organism we originally sampled it from and study it more.
To determine potential targets for drugs we have to understand what the proteins do. Having the structure is not really enough for that, it doesn't tell you the purpose of the protein (though it certainly can give you some hints).
In most cases the proteins were determined to be interesting by other experiments, and then people decided to try and solve their structure. So the structures we already solved are also biased towards the more biologically relevant proteins.
Anyone uninitiated with think the same, and thise already informed. Well, they are already informed.
> In most cases the proteins were determined to be interesting by other experiments, and then people decided to try and solve their structure.
Yes, that's what we're doing right now, because structure is not a useful predictor, because we don't have structure available in advance of studies on the protein itself. There was no point to a "functional taxonomy" of proteins, because we were never trying to predict with protein-structure as the only data available.
In a world where protein structure is "on tap" in a data warehouse, part of the game of bioinformatics will become "structural analysis" of classes of known-function proteins, to find functional sub-units that do similar things among all studied proteins, allowing searches to be conducted for other proteins that express similar functional sub-units.
If someone produces an AI that you give a sequence and it tells you what the protein does exactly, I'd be extremely impressed. I don't see that happening soon.
The specifics matter a lot here. We can often determine rough functions for subdomains by homology alone. But that really doesn't tell you the full story, it only gives you some hints on what that protein actually does.
"If someone produces an AI that you give a sequence and it tells you the protein conformation, I'd be extremely impressed".
Sure there are many more things to solve in this space; but that doesn't take away that this is an impressive achievement and does unlock quite a few things (including making more tractable the problem you just brought up). I'm excited to see what DeepMind works on now and what the new state of the world will be just five years from now.
I'm maybe overcompensating for the tech-centric population here, with some comments speculating for very near and drastic impacts from discoveries like this. Biology and life sciences are much slower, and there's always more complexity below every breakthrough. That does tend to push me towards commenting with the more skeptical and sober view here.
It'll take a while before those results can be trusted, though, right? There's probably a selection bias in the training data for proteins which are easy to crystallize, so many proteins probably aren't well represented by the training examples.
Background, Neural networks have two modes 1) training - where you learn all the model weights and 2) inference - where you run the model once on new data. Training takes takes a long time, because you're computing derivatives to implement updates rules on millions or billions of parameters based on iteratively examining massive datasets. Inference is extremely fast because you're just running matrix multiplies of those parameters on new data. And TPUs/GPUs are specially designed to compute matrix multiplies.
The article said: "We trained this system [...] over a few weeks." I searched for, but did not see them identify the inference time. I do expect inference time to be well under one second, though I'm not personally experienced with running inference on this type of network architecture.
For comparison, GPT-3 and AlphaStar have month long training times and real-time (sub-second) inference times.
https://science.sciencemag.org/content/369/6502/440.abstract
Where "a few" is around 0.1% of the known 180 million proteins. So a relative few and a whole lot.
But the catch is which proteins could we figure out by experiment, and which not. In particular membrane proteins are hard to experimentally determine. But knowing how they fold is very important for figuring out how to get things to react with or get through membranes such as cell walls. Which is an important problem for everything from understanding how viruses work to targeted delivery of drugs. We now have a way to find those structures.
Some possibilities: artificial muscles for robots, man-made blood substitutes, designer enzymes to break down plastic and other compounds. Software defined biology, where the pipeline from DNA code to actual protein can now be modeled in silico ahead of time. The biology classes of the future may be less observation of animals and more training in usage of whatever the equivalent of Autodesk for biology will be. Healthcare economics in developing nations will be changed as biochemistry itself may finally become deterministic (to some extent). Orphan drug development price would drop (and if you take into account right to try laws and ignore ethics in favor of progress, then people with rare disease may be cured en masse without bankrupting the health insurance company).
So imagine you gave an iPhone to someone in the 1800's, they wouldn't understand how most of it works, but this may be analogous to them finally figuring out some key aspects of the transistor. So it's another tool in the toolbelt and like all good tools will be used in all sorts of unpredictable ways.
Someone else I'm sure could do a lot better at explaining how important shape is to understanding the function and behavior of proteins.
Of course, there is a caveat. The static, crystallized structure is only one aspect of a protein. The dynamic behavior dissolved in H2O, at different pH, different ionic strength, with different ligands/cofactors are all also important, and not (afaik) directly addressed by this research.
So your second guess is correct - one of the steps is much cheaper now, which marginally improves the entire pipeline. As a result, drugs should now arrive to the market faster.
As a side note, I am curious what happens to the field of structural biology in 10 to 15 years from now. Every research university has a large structural biology department with super expensive Xray/NRM/Cryo-EM machines, and armies of students who routinely spend 4-6 years of their PhD trying to solve a structure of a single protein. If AlphaFold works as advertised, NIH will gradually shift funding to other problems.
(It was predicted that it'd be taxi drivers, not professors, that AI got first. Ironic.)
Back in the 1990s, when I worked on structure data, I remember that at least some crystallizations were easy enough they could be done as a rotation project.
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6287266/ suggests that life is now a lot easier than the 1990s. Quoting the abstract:
> Macromolecular crystallography evolved enormously from the pioneering days, when structures were solved by “wizards” performing all complicated procedures almost by hand. In the current situation crystal structures of large systems can be often solved very effectively by various powerful automatic programs in days or hours, or even minutes. Such progress is to a large extent coupled to the advances in many other fields, such as genetic engineering, computer technology, availability of synchrotron beam lines and many other techniques, creating the highly interdisciplinary science of macromolecular crystallography. Due to this unprecedented success crystallography is often treated as one of the analytical methods and practiced by researchers interested in structures of macromolecules, but not highly competent in the procedures involved in the process of structure determination.
Certainly some proteins are extremely hard to crystallize, and the new single-atom EM work will help a lot. But are there really "armies of students who routinely spend 4-6 years of their PhD trying to solve a structure of a single protein" these days?
I honestly don't know. I'm sure some do. But if so, that army is pretty small compared to the vast numbers who more routinely use crystallography.
Anyway that anecdote is pretty much the entire sum of my protein crystallography knowledge, but perhaps it explains how your experience and GP's statement can both be true?
Solving protein folding is huge, Nobel in chemistry scale achievement. It would be massive leap for biochemistry.
It seems that Deep Mind solved competition benchmark and made huge leap, but it's just partial solution that works on limited set.
After you have solved protein folding, there is still problem of solving chemical interactions between molecules accurately. Quantum chemistry is extremely compute intensive.
They always had single digit, bizarre artifacts, where the program can't sometimes recognise the very data it was trained on with most minute differences.
Other artifact was that the most "stereotypical cases" were least reliably recognised, and they hot a lot of flak for screwed up live demos, where a radiologist put a very, very obvious tumor shot onto the scanner, and it didn't work without a half an hour of wiggling the film, and a camera.
The "bruteforce" solution may well be always, 80-85% off, but off consistently, and always. NN algo so far beat them, but fail with double digit frequencies on "artifacts" which they themselves can't do anything about.
How well it deals with the later, is what I believe will measure its real world usefullness.
This is more the case the more close to bruteforce you come, like encryption cracking. Imagine, spending years of HPC cluster time, trying to break a password, while knowing you have a single digit chance to miss the right key, in a way which would be completely impossible with with a conventional solution.
One could say the same thing about programmers automating a task, or a number of other trivial examples. I would lean towards assuming deep mind has competent model validation teams vs. not, even if data science is hard.
So in 5 years you'll see exactly zero new medicines pop up.