But can we trust the hallucinations of the “AI”?
Anyway, from their tests using standard CNN practices: "Our best model achieved on the test set an F1 score of 0.97 (accuracy = 97.5%) for the classification task and a R2 score of 0.84 (RMSE = 21.9 m, or about 1 image pixel) for the length-estimation task."
secondly, there is no perfect way to measure accuracy of remote sensing for many reasons