A watershed moment for protein structure prediction
nature.com
nature.com
In fact, the distance constraints to 3D structure part is in fact very old - I was calculating structures from experimentally determined distances 30 years ago. You need a surprisingly low number of fairly weak ( these atoms are between 3 and 5 Angstroms ) distances to determine a 3D structure if you have a decent number of long range ones.
What they have done is execute better.
However the problem they are working on
sequence -> structure
though it's been a long term 'holy grail', practically it's not that useful!
The models typically aren't quite good enough, it's not predicting interactions, and experimental methods to determine structures have also moved on in leaps and bounds.
As the article briefly mentions what you really want to do is go the other way.
Designed novel structure -> protein sequence to make it.
One way to do that if you have a function going the other way ( like alphafold ( let's ignore limitations for now - ie does a knowledge based approach work well for completely novel folds ? ) ), is some sort of heuristic search - however the search space is huge and a step size of hours isn't going to cut it.
> sequence -> structure
I recall reading about a result a couple of years ago supposedly demonstrating that "synonymous" DNA codons were in fact not synonymous, because the ribosome took systematically different amounts of time to process them, and the difference in construction time resulted in different folding for the protein.
This would imply that the problem "sequence -> structure" is not well defined, at least if the sequence in question is the sequence of peptides making up the protein and not the sequence of codons making up the gene that codes for the protein.
Do you know anything about this? Am I just making it up?
It's not clear to me that you could truly demonstrate substantially different folding due to codon usage like that, on an experimental level, to make a general statement about all proteins.
This would also seem to imply that "sequence -> structure" isn't quite the right problem.
That's a reason that evolution-based methods, which use statistics about families of related proteins to estimate distances between pairs of amino acids (in 3D space), are more effective- many times in biology we can use evolutionary relationships between proteins to infer things that would be hard to determine through experiments or rigorous, thorough simulations.
But it's important to appreciate there are a large number of proteins that don't fold to a single unique structure rapidly-- and there are many ways this can be the case and many different biologically relevant behaviors depend on these properties. The tools from CASP are much less useful for proteins that violate the assumptions of Anfinsen's dogma, although the evolutionary data is helpful there too, it can often be a lot more challenging to deconvolute the signal.
Ultimately, "what is the right problem"? THe one that makes the most money? Produces the most "useful" scientific result? Is accessible with today's technology? For now, there's plenty of value in these sorts of competitions.
Personally I think the "right problem" is: "given a collection of diseases, use experimentally derived data and clever math, to discover biological treatments that reduce the total suffering from those diseases, subject to monetary and ethical constraints". That's what pharma attempts to do, although not particularly well. Others might say simply solving interesting problems like protein folding is inherently valuable as the right problem.
When you're doing protein expression, it's a standard procedure to do a codon optimization in order to adapt the codons to your host expression system. When you're expressing a human protein in yeast or ecoli the codons will be quite different but the folding is the same. If translation wouldn't show this stability it would be difficult to use foreign expression systems at all.
Apart from protein folding there are other interfering folding mechanisms like chaperone proteins that also make it hard to phrase it as a straight "sequence -> structure" prediction problem, though they often only exist for a small number of the total proteins (in humans).
Imagine software that evolved by trial and error rather than being designed - it's not 'clean'... there are no absolutes.
It's well known that codons effect translation rate and translation rate could/can effect the kinetic pathway of folding.
When a protein folds ( even away from the ribosome ) it isn't in isolation - it's in a solution - being bombarded by Brownian motion of other molecules.
It's amazing you are alive at all really - it also suggests most proteins would need the functional state to be strongly energetically preferred.
So there are potentially lots of things beyond just the protein sequence that could be important variables.
However some proteins will reversibly fold/unfold in simple solution all day long - so it makes sense to start there.
What's the rush?
1. We had a model that worked in principle, but the search space was practically infeasible.
2. We made an observation that a different model might exist that makes the search space irrelevant.
3. We threw ML at it.
4. Now we might have a model that fulfills (2) but we cannot be sure because we used a black-box approach.
5. Somehow the results are exciting. Better results would be really exciting.
6. We hope that more data yields these better results.
Is that correct? Am I the only one to lament these black-box approaches? Should there not be a bunch of people now studying the learned models to figure out if much better results can actually exist?
It means a model which has a somewhat opaque internal working. A lot of modern ML approaches treat the model as a black box in that it is not particularly clear how the features actually interact to reach the prediction.
> “not particularly clear how the features actually interact to reach the prediction.“
This is quite true of many models, such as linear regression in the wild. You may also have a clear but wrong picture of how features interact, eg looking at coefficients of interaction terms in a misspecified linear model.
See for example, “The Mythos of Model Interpretability”
- https://arxiv.org/abs/1606.03490
Without a coherent definition of what “not black box” or “explainable” means scientifically, the buzzword of “black box” is also meaningless and is more of a political game to be a gatekeeper over what models are allowed to be used than any kind of honest intellectual inquiry.
I’ve worked professionally on systems where misspecified linear models and text models using simple ngram boosting were vastly more inscrutable than comparable neural networks or dimensionality reduction models for the same applications.
Nobody has any scientifically cogent idea of what makes a model “explainable” other than arguing semantics.
Bullshit, any client you show a regression to can easily fully understand how the inputs change the output.
Section 4.1 of the paper he linked discusses extensively the limitations of linear model interpretability. Here's an example:
"With respect to algorithmic transparency, [the claim that linear models are more interpretable] seems uncontroversial, but given high dimensional or heavily engineered features, linear models lose simulatability or decomposability, respectively."
Given how many interesting problems are high-dimensional or solved through heavy feature engineering, you must be lucky to live in the happy space where problems are either low-dimensional enough to not need heavy feature engineering or where the client is already familiar with the high-dimensional features you're going to limit the model to using.
- http://www.saramitchell.org/achen04.pdf
The functional mechanism of a regression has nearly nothing to do with whether the regression explains the phenomenon it’s being used to model or not.
Quote:
"1. There’s a tiny functional language based on a small number of side-effect free combinators
2. For a given task, a program template (which the authors call a sketch), further constrains the set of programs that can be learned for the problem in hand. This also very handily constrains the search space of course, helping to make learning a suitable policy program tractable.
3. To help guide the search within the set of programs conforming to the sketch, a standard reinforcement learning algorithm is used to learn a (black box) policy.
4. The black box policy is used as an oracle (the Neural Policy Oracle), and a neurally directed program search (NDPS) tries to find the sketch-conforming program that behaves as closely to the oracle as possible."
The paper described:
http://proceedings.mlr.press/v80/verma18a/verma18a.pdf
...this seems kind of similar to the recent success in solving mathematical equations using deep learning. The language translation paradigm seems to have a lot more potential than some realized.
Merely decomposing a prediction into say a linear combination of predictors does not, itself, provide any type of explanation unless a variety of more complicated assumptions about statistical model checking turn out to be true.
Even worse, because the naive approach is to just treat those regression coefficients as if they do automatically give explanatory power or feature-wise attribution, people misrepresent things unwittingly and don’t carry out enough robust model checking, leading to wrong explanations that appear to have deceptive degrees of confidence associated with them.
I'm not sure we are talking about the same thing, or which of the three sources in my comment you are reacting to.
My understanding of "explainability" in this context is that it means "humans can understand why the model gives a certain prediction".
It seems like you are using it to mean "humans can understand why the thing modeled does something (presumably by referring to the model)". That isn't what I thought people meant by the term, nor does it seem like a reasonable goal to me.
Do we agree these are distinct ideas? Feel free to elaborate on where I've gone wrong.
I got bored with physics really quickly when I realized this is what everyone does.
If the result is immediately applicable to real world problems, like in this case, that's a perfectly valid approach.
In this case we just need the model to produce answers we can use -- how we made the model, what's in it, etc, we don't need to care.
Far from it. Prediction looks like a tool in the arsenal for better understanding. One still has to correlate the structure with the complex interactions in vivo. Even using AI in classification mode, where we can segment a large atlas of tumor cells and identify a dozen or so classes of cell anomalies may lead to faster breakthroughs in immunotherapy.
What I am trying to wrap my head around is the synthesis problem. Say AlphaFold generates a promising candidate. One that does not exist naturally. You still need the DNA or mRNA transcription sequence to synthesize the protein, right? Won't some candidates simply be too complex and unstable to reliably produce using existing mammalian or baculovirus platforms?
You can add that to the objecting function that your model training function is optimizing, ensuring model output is not "too complex or unstable to reliably produce".
Methods from the domain where already not truly model-based (compared to things like physical equations). It is mostly observations on existing proteins coupled with a form of gradient descent so, for this particular application, things have not degraded. (I am aware that there is more to it, this is just a fast summary)
But to be honest, it is far from a solved problem and I expect more breakthroughs from deep understanding and modeling than ML.
Huh, I would've said exactly the opposite. The feature space of the variations of amino acid sequences is so big that I wouldn't bet on breakthroughs derived from understanding. Most recent advances seem to focus on how specific sequence motifs interact with each other, which are generally only applicable to certain kinds of proteins, but not protein folding as a whole.
My prediction would be that next breakthroughs that push the field over the usability limit are derived from more general machine learning advances.
What it can be is a force multiplier for every progress in our understanding of why those sequences are transcribed into those structures.
But I have to admit that, with years of head-on, the research has not gotten there so I might be wrong.
ML is that with significantly more horsepower.
But fluid mechanics has a very deep theoretical underpinning, and generally has interpretable models. I'd suggest skimming through a copy of either Lamb or Batchelor's book, to see how far you can get just with pen and pencil and no statistical input at all.
looking back through the google, I may be dimly remembering the calculations dealing with turbulent flow. It was a whole different experience to a College Junior that was used to the relatively simple equations from Physics, Statics and Dynamics.
The books I mentioned are classics in the field and were first published in 1895 and 1967, respectively. Both are still in print. No computers, just advanced math (vector calculus etc).
It's not like the alternative is any different. Being sure you have the right answer is not viable, the other methods of getting there are just more transparent heuristics.
In much of drug discovery/development (and disease research), being able to predict protein structure would be very valuable. Being able to quickly find candidate structures (that can then be searched for in the lab) speeds things up immensely. Reducing false positives (or just coming up with possibilities at all) is a huge win.
But you’re right that this probably doesn’t help the protein theoretician much if at all. We already “know” how it works (it’s just thermodynamics and quantum mechanics) and of course have no idea how it works (“well they wiggle around until they find a low energy state” doesn’t really tell you anything). But that doesn’t keep this from being exciting.
It’s the sort of thing researchers say to get grants, but as a distant goal, not a practical reality.
For structure-based drug discovery (which isn’t even the majority of drug discovery), the details are what matter (e.g. “does this water molecule mediate a binding interaction, or do the sidechains shuffle a bit, and kick the water out?”), and these methods don’t even come close to predicting detailed interactions.
Metrics in this space are focused on “general correctness” of protein backbone conformation. Success is to achieve a kind of blurry view of the overall shape of the molecule, and drug design is trying to predict specific atomic interactions. They’re two wildly different problems.
About the best you can say is that if we had a generalizable model of physics that could predict protein structure, it might also be able to do a good job of evaluating how a small molecule binds to a protein target. But even that is a huge leap, and when you start using black-box methods like AlphaFold to specifically solve the problem of structure prediction, it’s not really clear that generalization is even possible.
There are potential practical uses in drug discovery for a method that can design a protein which takes a particular shape, but even that is pretty different from what AlphaFold actually does.
There are indeed many different applications of ML to drug discovery, but AlphaFold is a niche technique, and protein structure prediction is not practically useful, in general.
There are many areas where ML is being used to advance drug discovery and other kinds of practical, difficult biophysics problems. Casual readers would do well not to get too fixated on this particular area.
https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp...
I especially enjoy the segments where he upfrontly addresses "an indictment of academic science" and "an indictment of pharma". He pulls no punches in saying how embarrassing it is for pharma and academia to be literally outclassed by DeepMind.
A great quote:
"If you think I’m being overly dramatic, consider this counterfactual scenario. Take a problem proximal to tech companies’ bottom line, e.g. image recognition or speech, and imagine that no tech company was investing research money into the problem. (IBM alone has been working on speech for decades.) Then imagine that a pharmaceutical company suddenly enters ImageNet and blows the competition out of the water, leaving the academics scratching their heads at what just happened and the tech companies almost unaware it even happened."
And I read that the size of the team was 10 people - that's not a big number.
The compute power applied was not why they had this outcome.
"What is worse than academic groups getting scooped by DeepMind? The fact that the collective powers of Novartis, Pfizer, etc, with their hundreds of thousands (~million?) of employees, let an industrial lab that is a complete outsider to the field, with virtually no prior molecular sciences experience, come in and thoroughly beat them on a problem that is, quite frankly, of far greater importance to pharmaceuticals than it is to Alphabet. It is an indictment of the laughable “basic research” groups of these companies, which pay lip service to fundamental science but focus myopically on target-driven research that they managed to so badly embarrass themselves in this episode."
From: https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp...
I think a lot of the commentary is missing two essential points:
1. Protein structure prediction is to a large extent a solved problem for small-ish, soluble targets. AlphaFold is a significant improvement on the current state of the art, but the state of the art was already far enough along that the best computational models in 2007 were good enough to bootstrap experimental structure determination (https://www.ncbi.nlm.nih.gov/pubmed/17934447). In other words, it's not like the entire academic community was stumbling around helplessly in the dark.
2. The value of these predictions to pharmaceutical companies is extremely marginal. Having a high-accuracy model is very helpful but it's rare that the researchers have so little information available that a completely de-novo prediction is necessary. And when they really don't have much information at all, it's usually because the target is sufficiently messy to defy traditional structure determination methods - which means it's almost certainly more than AlphaFold can handle too.
Without measurable benchmarks we have no idea if we're making real progress towards human level AI.
After all these years, we are still at it. New methods are regularly evaluated and the simulation software is being refined all the time.