Scores of 30-45 for 15 years. Now scores of 87-92.
This isn't a minor improvement, it's a leap forward.
Scores of 30-45 for 15 years. Now scores of 87-92.
This isn't a minor improvement, it's a leap forward.
>a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods
So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.
The threshold for "real" in particle physics is +5 sigma. Which takes a lot of data.
These proteins are so massive that we often use Daltons [1] as an averaged measure of molecular weight.
Conceptually one of the most promising applications of quantum computing is theoretical chemistry, and we are only now starting to make progress in this avenue [2]. I anticipate it would require quantum computing to explicitly optimise large folded proteins.
1. https://en.m.wikipedia.org/wiki/Dalton_(unit) 2. https://arxiv.org/abs/2004.04174
Think of it like satellite imagery of a tree: A score of zero would be a single green-ish pixel, while a score of 100 would show each leaf within the range it naturally moves in due to wind etc. (proteins tend to wiggle quite a bit under natural conditions, as well)
In this case, the model is predicting values of multiple structures, but patterns could still theoretically be found which allow for predictions beyond the accuracy of a single measurement.
Let’s say you have 100000 proteins in the training set. Now remove #1 and train on 99999, and then check that it still predicts the same protein result for #1 as the experimental result.
Or remove from training whole sets of proteins by particular teams to find systematic errors made by teams?
While this is an accomplishment, nobody is going to be confusing these models for structures produced experimentally. The CASP metric is for backbone atoms. To have a useful model of protein structure, you really need to have the positions of the protein side-chain atoms modeled correctly. Experimental methods will do that, but this method, as I understand it, does not.
Which gets into the concept of whether the ML model has actually learned some deeper conceptual ideas than we have, some deeper truth about how this works. If so, can we somehow extract that truth, or is it truly a black box that does the thing we want?
I'm reminded of a sci-fi book I read long ago in which humans are discussing the fact that the science they are utilizing is beyond the scope of a human mind to comprehend- only the AIs can intuitively deal with 12-dimensional manifolds (or something to that extent). Maybe we've reached the doorstep of that future.
So i do think the results could be more accurate than measurement.
Well I think that the results speak for themselves; ultimately the question you raise is one of semantics. ML models don't think in terms of "conceptual ideas" like humans do, these models simply perform at such a massive statistical scale that they can identify patterns far beyond any human conception. Clearly, the model embodies some verifiably reliable information about the way the world works, but this is "just" a trick of statistics not anything resembling actual "understanding" in the way the word is typically used when referring to human understanding.
What's an experimental method for protein folding and why is it so good? Are they talking about creating an actual, physical protein in a lab and observing how it folds?
Exactly. Researches purify the folded protein and then use methods such as X-ray crystallography, nuclear magnetic resonance, and cryo-electron microscopy to determine its three-dimensional atomic structure.
We understand the physics of e.g. X-ray diffraction pretty well, so we can fit pretty decent forward models for the x-ray data given a proposed structure. The hardest task here is getting a good enough guess at the structure to optimize the physical model, and it’s my impression that people use an iterative model refinement workflow. At least that’s how it’s done in condensed matter materials.
There are many sources of experimental uncertainty, like the non-ideal nature of the x-ray source and optics, and the fact that the atoms in the protein are not static but have some thermal fluctuations. so at the end of the refinement you still have some uncertainty on your model parameters (the interatomic distances for proteins I guess), but if you are careful you can calibrate these uncertainties pretty well.
This paper looks like a really good detailed discussion: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4080831/
So you "make guesses at what the phases are", the best choice is to bootstrapping these phases measured with another technique (you can introduce crystal defects that do allow you to guess at what the phases are).
Less scrupulous is to use a computer generated model, like fitting another protein "that you guess is related", then you model the electron density, take the phases of that.
In any case you take these "phase" guesses, and then apply it to your intensities, re-run the fourier transform, refine your electron densities, twiddle the location where you think the atoms, are, then repeat with your new model. This process repeats until you converge on a structure that you're happy with.
Now alarm bells should be screaming in your head right now: Yes, it's entirely possible to converge on a wrong structure, especially if you're a young up-and-comer professor seeking tenure that has no ethical problems with "suggesting" their grad students to sleep in the lab and work 100 hour weeks and willing to do slipshod work to get you tenure: https://www.sciencedirect.com/science/article/pii/S002228360...
I guess my question is, how do you know if you’ve converged on the right structure or not? Is there a different experiment you could do?
This bodes extremely well for the future of computational biology, I'm very excited thinking about the prospects. If we know how a protein folds, we know its shape, meaning we know which shaped/charged molecules are needed to act as suppressors/enhancers of those proteins.
Could this lead to a virtuous cycle where AlphaFold is used generate a ton of random sequences where it has low confidence, those are then screened for ease of synthesis, measured and the results used to improve the model?
Edit: nevermind, according to another comment[0] there are still plenty of real proteins without experimental data left to explore.
It can verify how much it minimizes the potential energy, which may not always line up with how it would fold in the real world but is a strong indicator.
There's no reason to believe the list will contain all solutions, however.
Does seem like the contest structure could include quite a bit of risk for hiding the effect of overfitting ... I wonder if there is anything inherent about the problem that reduces that risk ...?
The reason why the top score in one year, can be lower than in the previous year, is that the test (the 100 structures to guess) is always new and different, so it can end up being 'harder' than the year before. Luck will also play a small role.
Another explanation for a reduction in the top score would be, that previous winners are not re-submitted unchanged. For instance AlphaFold v1 seems to not have been submitted to the latest competition.
Is it really possible to select 100 new structures which together are likely to represent a meaningful increase in the sample generalization versus the prior years test set ...?
Using 1% of those (presumably from the more-often-reproduced subset) for this challenge seems reasonable? Note that the structures have to remain secret up until the challenge, and presumably all those teams uncovering the structures don't want to have to wait up to 2 years every time to actually make their results public.
I suppose it will take a few more years of repetition for the challenge to confirm that the problem has been been solved -- but I wonder if a new version of the contest is going to be needed as well? Maybe the model accuracy is now high enough to invert the contest to a form where models generate predictions for randomly selected unknown samples -- and experimental teams are then expected to make observations for those particular sequences over the next two years as part of their otherwise research agenda selected experimental workload?
But yeah, compared to other fields, the size of training/test sets is sometimes pretty small in ML for life sciences.