Not to give DeepMind too much credit, but it's also entirely possible that "both are correct". By the nature of how targets are selected for CASP, year over year there is higher likelihood that "pathologically difficult" proteins are presented for the competition, for example - a protein that exhibits polymorphic structure where the technique (say Xtal vs NMR) biases the "solved" fold in a huge way. I believe CASP tries to weed out proteins that we know are polymorphic, but you can never be sure, and again, as time marches on, those are the proteins that fundamentally are harder to solve, so it's likely they will be enriched in the test pool.
An extreme example is insulin. The structure of insulin is has been solved for 60? I think years, but when it's in the environment of its receptor, it looks TOTALLY different (solved 5? years ago). Having said that, doubtful that DeepMind could ascertain that structure, since the environment is super-super different.
I think that the modestly high error rate is an indication that deepmind is mostly interpolating and has solved the broad protein folding heuristic. It probably will get better.
Right but we also can't generalize from 2 out 3, and without knowing this figure it's really hard to say how useful this is, no?
> Knowing how ML works, I'm surprised it didn't give any indication of low confidence.
This is actually a hard problem in ML, for example in NLP, many people assume a high log prob score means a high confidence but it is not true at all.
Thank you for the clarification. I wasn't aware of this, as I'm most familiar with super-basic/old NLP techniques like BOW/RNN/LSTM/GRU techniques, where log prob scores seemed to me to be roughly correlated with result quality. I'm aware the landscape has changed recently with insights about high dimensional searches...
But after training you can recalibrate the temperature of the softmax on the test set and still get meaningful confidence scores (temperature calibration). Or you can use a variation of cross entropy called Focal Loss that will leave your logits un-squashed.
If it got answer right 1 time in 100, that would be amazing and you'd be foolish not to use it!
Except you have no way of knowing if the answer it gives you is one of the 66 right predictions vs one of the 33 wrong predictions. You could say it's likely correct, but not to a high enough degree of confidence that you could really trust it without verifying using the old established techniques.
Also important - verifying that a given model has a signature that matches the established techniques is far easier than using those techniques to generate the complete model from scratch.
I'm not really sure what your point is.
This isn't really how it works.
To quote the CASP competition organisers:
The organizers even worried DeepMind may have been cheating somehow. So Lupas set a special challenge: a membrane protein from a species of archaea, an ancient group of microbes. For 10 years, his research team tried every trick in the book to get an x-ray crystal structure of the protein. “We couldn’t solve it.”
But AlphaFold had no trouble. It returned a detailed image of a three-part protein with two long helical arms in the middle. The model enabled Lupas and his colleagues to make sense of their x-ray data; within half an hour, they had fit their experimental results to AlphaFold’s predicted structure. “It’s almost perfect,” Lupas says. “They could not possibly have cheated on this. I don’t know how they do it.”[1]
So you have experimental results, but still don't know how it folds. You aren't trying to avoid the all the experiments, just understand them.
[1] https://www.sciencemag.org/news/2020/11/game-has-changed-ai-...
When you set a threshold you improve precision at the detriment of recall. It's a tradeoff you can play with, but the score depends on it.