It is 2026. How are we still publishing articles on medical diagnostics data science and using area under the ROC curve as the primary metric of success. ROC-AUC of 0.9 under severe class imbalance (almost always the case in diagnostics) could still mean something like 4/5 predicted diagnoses are wrong (false positives). Precision-Recall curve + mAP or GTFO.
Science article in question: https://www.science.org/doi/abs/10.1126/science.aec6129
Also, the most interesting result here is that the CNN-based feature encoder significantly outperformed a vision transformer encoder backbone…