A protein sequence is analogous to a computer program, but the "machine" is a mostly-water solution, and the instructions are interpreted by summing up all the intermolecular forces at play as the sequence is squirted out of a little extruder as a string (as in the stuff your clothes are made of, not text). Bits of the string repel and attract each other and it globs up in some biologically useful structure. The folding problem is the problem of predicting that structure from the string.
Unlike the halting problem, there is no way to generate an execution trace saying a particular glob would be formed in reality. In fact, there is no way to perform a polynomial time check of the result, so we've already escaped the land of P and NP.
Also, things like temperature and proximity to other proteins mean that there might not be a unique fold for a given sequence. Therefore, like the halting problem, we have unknown inputs, and we need to figure out which states an arbitrary program can reach.
When someone claims to have "solved" folding, you should be as skeptical as you would be if someone claimed to have solved the halting problem for arbitrary machine code, and that they don't need any extra information about the machines that run that code. Although their program runs on conventional computers, it also works on programs written for quantum computers.
(Edit: That's not to say this work isn't useful, or that this press release overclaims. I've been hearing some pretty wild claims about this work elsewhere...)
FWIU, Folding@home has additional problems for AlphaFold, if not the AlphaFold team;
> Install our software to become a citizen scientist and contribute your compute power to help fight global health threats like COVID19, Alzheimer’s Disease, and cancer. Our software is completely free, easy to install, and safe to use. Available for: Linux, Windows, Mac
Is there any NN architectural reason that AlphaFold could not learn and predict the Folding@home protein folding interactions as well? Is there yet an open implementation?
Yes, there are open implementations of nearly-AlphaFold at this point.
E.g. re-derivations of Lean Mathlib would be the strings to evolve.
While Wash U was a contributor, I am confused about why you call it the home of the Human Genome Project. The Project seems a lot more strongly linked to the Whitehead/MIT in terms of press and the site of key figures.
"The Cost of Sequencing a Human Genome" https://www.genome.gov/about-genomics/fact-sheets/Sequencing...
Let's say you have a car (or lego set etc). The number of possible ways the parts could go together are astronomical! Does that mean it's not possible to figure out how it fits together, or how you might build one?
But there is no reason to expect this process to produce the right end-result for a Lego set that has never been seen before.
There is also a lot of sub-structure that helps - similar parts of proteins tend to fold in similar ways, so even if you don't have real predictive power on unknown sequences, you may do quite well for a protein that is 90% the same as one in the training set - you will be quite correct on ~90% of the folds, even if your pretty way off on the remaining 10%.
Note that all of this is not to minimize the success of what AlphaFold achieved. I am just trying to explain how you can do well at this problem without having discovered some deeper deterministic structure in protein folds.
So, I can pull a protein out of thin air and there's a good chance it'll have an overall fold similar to another protein that's got a structure. Unfortunately, the devil is almost always in the details. An amino acid here or there, a short extension here or there, a missing charged residue or an extra glycine and now you have a different target and entirely different behavior in a biological system.
One cool thing I found actually, was a protein in an Archaeal virus had no known homology a few years ago, but when I checked the other day, it now matches most closely to an (otherwise thought to be) entirely synthetic protein out of David Baker's lab at UW. Which means this Archeal virus and David Baker converged on the same fold somehow (likely because it was "stable").
I’ve found it extremely hard to have a casual understanding of biology, unlike math where I feel like I have a solid high level sampling of the field. I’ve done a few bio and chemistry courses and books but it’s so deep and ill suited for a programmer who is used to asking how things work underneath at every level (you have to constantly stop yourself from asking why something does what it does and just go with it until it starts to connect later, which is more of a commitment than I could give).
Anyway thanks for your comment
Biology and biochemistry is unbelievably complicated and difficult to grasp without truly going deep into the fundamentals
I am looking at Molecular Biology of the Cell (Alberts) and Cell Biology (Pollard). Both were recommended to me, but wondering what the pros and cons of each are (if you are familiar with both of them).
- There is a (short) book called "A Computer Scientist's Guide to Cell Biology" by William Cohen which is a little pricey but very dense and helpful with a lot of concepts.
- Combine that with David Goodsell's "The Machinery of Life" which has a lot of great illustrations and practical examples.
https://www.edx.org/course/introduction-to-biology-the-secre...
Alternatively, Harvard Extension School has some great biology courses you can sign up and get credit for. Though those are mostly for pre-med career changers
#1: The quantum interactions of electrons that are the basis for chemical bonds behave in ways our computers and intuition are incapable of simulating
#2 It's a matter of degree, not kind, and nature is more sophisticated than our computers, reasoning and thought processes.
#3 Nature is magic, whatever you define that to be
#4 When stipulating the degrees of freedom involved (ie from dihedral angles), the possibility of additional information we haven't discovered is being overlooked. Is there a recipe or algorithm that could help?
But nature hasn't really "solved" the problem, it is just doing its thing, but the way it does things is completely different from what our computers do.
It is like trying to reproduce a guitar sound using a synthesizer. A guitar solves to problem of sounding like a guitar, but it doesn't mean it is more sophisticated than a synthesizer, in fact, a synthesizer can do much more, it is just that the process by which the guitar makes sounds are hard to simulate.
For example, the problem "is there a route in this graph that visits all nodes and has length <= L" can be quickly verified with a classical computer, as long as you're given a "yes" answer accompanied by such a route. Finding the answer from scratch might be much slower, but checking it is quick.
That's absurd. The halting problem is provably impossible with either conventional computers or Quantum computers.
This is clearly not true for protein folding, although it is possible that it is computationally intractable with a conventional computer.
Take a look at the configs for Amber (molecular dynamics simulation -- https://ambermd.org). QC might help map the space of inputs that would converge, but it probably couldn't identify a hypothetical 'done folding' state for any given protein.
I didn't see any claim by DeepMind that protein structure prediction is a solved problem. I think these guys are pretty diligent when it comes to communicating their science. What you may have seen, is a non-scientist reporter making inaccurate claims.
The problem with the structure prediction problem is not a loss/energy function problem, even if we had an accurate model of all the forces involved we'd still not have an accurate protein structure prediction algorithm.
Protein folding is a chaotic process (similar to the 3 body problem). There's an enormous number of interactions involved - between different amino acids, solvent and more. Numerical computation can't solve chaotic systems because floating point numbers have a finite representation, which leads to rounding errors and loss of accuracy.
Besides, Short range electro static and van der waals interactions are pretty well understood and before alphafold many algorithms (like Rosetta) were pretty successful in a lot of protein modeling tasks.
Therefore, we need a *practical* way to look at protein structure determination that is akin to AlphaFold2.