That is not an outstanding question. The test on which DeepMind scored high marks is a test of how well the algorithm folds novel proteins -- proteins whose ground-truth structure has not yet been published.
> the outstanding question ... is going to be how good is the confidence metric at telling the user to trust or not trust the results.
There is generally a threshold, less than X, not the class, equal or more, is the class. Then you run the network with the same threshold on a known data set and compute a confusion matrix, which tells you about the error, I don't even want to know what a confusion matrix analogue for 3D geometry would look like but I'm sure they have something.
This is literally the process that one does in taking part of the this. And the error rate (specifically the lack of errors) is what is everybody is talking about. 90 is just as accurate as we can get with experimental measurement. It's likely at this point the source of error is in the data set (we can only train on data we experimentally measure and these are not perfect measurements). It's also possible, at this point, the model generalized so well that when it deviates from experimental measurements it's actually correct and the experimental value was the one that was wrong.
So no, the outstanding question is not "is going to be how good is the confidence metric at telling the user to trust or not trust the results.". Nobody is going to be looking confidence values when it model is giving an output, they are going to be looking at the overall error rate across a broad spectrum of proteins to get a sense of it's accuracy.
in reality, that's not how it works at all. The energy functions we have are crappy and require too much sampling before we can find the lowest energy configuration. And more importantly, it doesn't look like proteins typically fold to their lowest energy configuration (with the exception of some small fast two state folders), but rather explore a kinetically accessible region around there (or even somewhere else entirely, if the energy cost to transition is too high).
Methods like AF depend heavily on large amount of information correlation from evolutionary data, which has historically been of the highest value for making decisions about protein structure.
The article did claim:
> According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods.
So perhaps their score of 87 GDT is pretty significant. But “competitive with” is not the same as “always in agreement with”, as you point out. Could be the failure modes are problematic.