Are AlphaFold's new results a miracle?
olegtrott.substack.com
olegtrott.substack.com
From what I can tell it still depends heavily on having a good sequence and structure template (or templates). It tells us little to nothing about the specific details of the folding process. To me the only part that seems miraculous is that it seems like we can predict novel structures (previously unknown conformations) using small fragments of templates rather than entire protein domains.
Just to add a bit of context: Rosetta’s de novo methods, which had the highest success rates of template-free structure prediction before the arrival of ML-based protocols, use a similar approach. Picking small fragments from protein structure databases reduces the conformational sampling space a lot.
I agree, and that’s probably why they switched to de novo prediction/design at some point.
It’s a small nitpick, but I think that the author actually meant “sequence identity” here, because his statement would make much more sense then. Sequence similarity is physicochemical in nature, and tends to be concentrated, in addition to functionally relevant sites (usually ligand-binding residues, as he mentioned), at key structural regions of the protein such as the hydrophobic core (where a high frequency of similarly hydrophobic residues is expected).
This is one of the reasons why proteins from the same family can share highly similar structures while having very low sequence identity, with highly conserved motifs (where the sequence identity is concentrated) taking care of the functionality.
I think his central point is fair and interesting. The test train split is apparently legit, as they used structures released before 2021 for training and the rest for testing. However, there was no real check for duplicates, and the success rate might be inflated by a bunch of "me too", low hanging fruit structures that are very slight variations from what we know.
However, I'm not sure I agree with his skepticism. LLMs suffer from the exact same problems - getting it to write a Snake game in any language is trivial, but it is almost certainly regurgitating - , but can be useful as well. I mean, if for various reasons people are publishing very similar structures out there, there's certainly value in speeding up or reducing that work considerably.
AF3 stands as one of the greatest achievments in machine learning/structural biology we've yet seen.
They do remove duplicates by sequence similarity (filtered PDB).
Please assume the DM folks really do know what they are doing.