Alpha Fold promises to revolutionize biochemistry
medium.com
medium.com
I’m equally skeptical of the claim that this follows the 80/20 (or 90/10) rule. If you use straightforward measurements (rmsd etc), reality may well be 10 % off considering these aren’t static structures, but molecules in perpetual chaotic motion on several scales.
Homology is never considered "solving the protein folding problem". It's 100% based off of past data which incorporates all sorts of biases, including as others have pointed out, experimental error. For instance, in my field, where there is terrible amino acid sequence conservation, AlphaFold is just as poor at prediction as existing solutions like Phyre.
Many aren't switching, but not because they disbelieve it. For example, in precision medicine, a Stanford professor was mostly skeptical of practical relevance because of the general gap between real-world data findings and the multi-year clinical trial process... but that's not about GNNs, but the well-known issue of institutional resistance to real-world data.
(We're biased: We're working with enterprise to bring in GNNs more on the supply chain, finance, fraud, security, etc. sides.)
It's still lousy when there's inadequate folding, it's lousy at oligomeric state (although that can probably be improved), and it's got tons of baked in, innate bias. It's also lousy at conformational states (but so is most empirical data). It's good enough for something that doesn't need to be high resolution, but so were preexisting methods. It's not good enough to predict a fold de novo.
Ex: https://www.icr.ac.uk/blogs/the-drug-discoverer/page-details...
""" Not surprisingly Lupas was impressed that AlphaFold enabled him to determine in half an hour the structure of a protein that he’d failed to solve for 10 years! """
From a computational / biology perspective, this is also important as part of the broader push to get end-to-end simulation & neural networks for areas beyond individual protein structures. Imagenet is a really good analogy: what started with a jump in mnist tasks is now at the level of deepfakes and multi-modal audiovisual inference. The leaders of the field defined the bar for what it means to do well in this area, and with alphafold doing well, better and composite solutions can now happen. (And more subtlely, they showed ~GNNs work in this area in general, which was not viable till recently.)
Observationally... when NN's enter a field like this, they seem to grow in use, not shrink. There's an uncomfortable epistemological, and arguably Darwinian, counterpoint to rejecting blackboxes or mathematically unprincipled methods, particularly when their objective results are superior. So for many people who claim NN's can't be used, that's often from a subjective methodology preference, and modern history seems to be on the side of NNs. I dislike that's true of a leading method, but that's where a lot of modern modeling is.
Hard disagree. It's common knowledge, and was even mentioned in this fluff article, 90% accuracy is within the experimental error of the "truth set". It's entirely possible alphafold is more accurate then the X-ray crystallography source for a large portion of the data set.
I would like to know what are you basing your very bold claim that 90 to 100% is bigger than 60 to 90%?
Talking to random people in the pharmaceutical industry also seems like a very poor way to find truth. It's a very big field/industry with entrenched interests and very narrow fields of view. Also, roughly 100% of the companies who are asked, "are you about to be disrupted by new technology that you didn't create and is totally different then the way you do things?" answer, "Nope!".
Even if it was, we don't have a way to distinguish and so I'm not certain what extra good that does.
>I would like to know what are you basing your very bold claim that 90 to 100% is bigger than 60 to 90%?
Because 5-10 angstrom resolution isn't terribly hard to do either with traditional biochemistry, or with homology modeling, with very few exceptions here. The problem I have is that the confidence level there isn't good enough for the actual work that matters. Exact ligand interactions, etc.
It really doesn't seem like you are well versed on this topic. I don't know any other way to put it.
Before alphafold, the best modeling methods would hit ~75% accuracy, and x-ray crystallography is both prohibitively expensive and also specimen specific in regards to it's efficacy or even possibility. So what alphafold can do, "traditional biochemistry, or with homology modeling" can not touch and I don't know why you are claiming they can nor what you are basing that claim off of. I feel like I'm repeating myself and I wish I didn't have to, but, alphafold's 90% is 90% because that is the highest possible score with the level of error within the source material, not because we know there is 10% error left to "fix".
> The problem I have is that the confidence level there isn't good enough for the actual work that matters.
Yeah, that is actually how the world works. Confidence starts low then builds upon successful results. It's really not worth bringing up unless we start to see a string of failures from alphafold's prediction ability, which is the opposite of what we've seen from alphafold. 90% accuracy has to be good enough to be useful as it already has been pointed out, that's within the error range of x-ray crystallography , the current "Gold standard", so let's be clear here, any "crisis of confidence" here isn't "this isn't precise enough to work" but "I don't trust dem neural networks, they are going to take our jobs".
Ok. I am, but I don't expect you to take my word for it.
>alphafold's 90% is 90% because that is the highest possible score with the level of error within the source material, not because we know there is 10% error left to "fix".
This is just making my point. Homology modeling is not knowledge of protein folding. It's knowledge of structures of experimentally observed folded proteins
>which is the opposite of what we've seen from alphafold. 90% accuracy has to be good enough to be useful as it already has been pointed out, that's within the error range of x-ray crystallography
I dont think you really understand how protein model building works. You are conflating prediction error from a model with observed error from empirical data. An xray structure with an Rfree of 10% is not the same as a predictive model that is 90% accurate, because in one case you have electron density to guide you, and in the other you do not.
>so let's be clear here, any "crisis of confidence" here isn't "this isn't precise enough to work" but "I don't trust dem neural networks, they are going to take our jobs".
Unnecessarily derogatory
I don't know anything about protein folding or drug development. But you brought up a very good point. Sometimes insiders and experts are the wrong people to ask, because they are not impartial. In the worst case we've seen so many virologists are on a campaign to paint lab leak as conspiracy theory just because their careers at risk. In similar vein it's also not the best idea to ask Wall St. who is responsible for 2008 financial crisis.
I wouldn’t discount it as progress toward something grander though.
They will get crushed by their pessimism. It is definitely a big deal and you throwing "muh 90% is not enough" just points that ignorance is still widespread.
We have this information. And we've had it for decades. There's hundreds of duplicate structures on the PDB. There are new cryo-em structures that are nearing sub angstrom. Also, the xray data we do have is never considered absolute, but the boundaries and margin of error have been quite well understood for some time.
Or the CASP competition?
Or posted further up: https://www.science.org/content/blog-post/more-protein-foldi...
Or take the RMSD of a superposition of any two identical structures on the PDB?
https://www.science.org/content/blog-post/more-protein-foldi...
The real ‘breakthrough’ IMHO is that deep learning is so generally useful, rather than any specific application per se.
Completely agree with the stance. Protein modelling is one part of the enormous jigsaw that is drug development. Structural biology is already a computational savvy discipline that is accustomed to using available tooling/predictions to enable better decisions with spare/minimal data. Sometimes the predictions are good, sometimes bad, but they are always being reviewed by a human expert. Maybe alpha fold is significantly better than the competition -then it is likely to make solving structures faster.
Hypothetical: the protein structures of all human proteins were solved tomorrow -is human disease cured? No, far from it, there is so much development work that goes beyond knowing the accessible pockets on a protein.
Genes are not the end all be all to medicine let alone biology. There are places where medicine won't go, such as the delicate structures of the brain. Much too unethical. Neurotransmitters fire across synapses. You will always be uncertain of what temperature or charge you need. No way of telling what's really on the inside. And that's with an incomplete model.
You also can't just zap or plunge in medicine. If you are going gene by gene or protein by protein, yet alone cell by cell, how long is that going to take?
Sure.
>and any disease
Impossible. Modeling a drug's effect on mental illness would require... fully simulating a human brain. Modeling the effect of thalidomide on fetuses would require fully simulating fetal development. That's not happening any time soon.
If that drug then also kills the patient, it's not much of a medicine.
Sure you can direct an algorithm to optimize for a metric (like curing a disease), but it will be completely ignorant to all other aspects of being (like not having harmful side-effects).
But a lot of what they do is outsource the initial R&D.
First, however you have to have a promising pre-clinical studies, then likely pass a Phase-1 trial. Lots of bio startups take a pre-clinical product through a Phase-1 trial. Traditionally, they would IPO to fund this. That game has changed a bit, where now some VCs will sponsor the Phase 1. There's also the whole SPAC thing as well. If Phase-1 is a success they usually get bought by a big biotech or pharma.
In terms of your particular example, how good is your simulation? If your simulation can accurately predict the effects of perturbations on metabolic pathways, then PM me and we can start a biopharma company.
Denaturation is not all that significant. Physiological conditions (i.e. pH, temperature, etc) tend to be maintained by homeostasis. Protein structure prediction is already incredibly useful.
> what is the status of our current metabolic pathway simulations for example ?
IMO we do not have the data to build useful metabolic pathway simulations. To do a good simulation you need to know (1) the proteins and metabolites involved, (2) their functions and key biophysical properties (ex. Km), (3) feedback / control elements. For most pathways we do not have (1). The data for (2) are mostly nonexistent, buried in literature, and / or unreliable. Same with (3).