I would suggest these two things might be linked - the only way you get publishable results is by failing to do rigorous studies.
I would suggest these two things might be linked - the only way you get publishable results is by failing to do rigorous studies.
"Please use the original title, unless it is misleading or linkbait; don't editorialize." - https://news.ycombinator.com/newsguidelines.html
I hope I don't offend any other people in the field, but I think historically fields like histology/radiology/MRI screening via AI have been subject to less scrutiny by virtue of their multi-disciplinary nature... particularly in the past, it was hard to find reviewers that both (A) understood the biological/clinical validation (B) the technical validation.
I think things have drastically changed over even the last 5 years, which is why you're more likely to see headlines like these. We have a lot of discussions like these during our lab's journal clubs. The optimists among us (e.g. myself) argue that this opens up a lot of opportunities whereby more fair validations/benchmarks allows us to compete with "SotA" methods that are actually fragile and easy to surpass on even ground. More pessimistic members are quick to note that it's far harder to introduce new, fairer validations and that reporting lower metrics is far less buzz worthy (all true points).
This reminds me of a fundamental error that was recently found in methods that tried to predict protein-protein interactions. It was found that the train/test/validation method used by ostensibly every paper was leaking huge amounts of data [0]. When we plugged that leak, we saw much more modest metrics, but focusing on regularisation allowed us to beat the competition [1].
The kicker? The information leak were identified in 2012, yet you'll see papers written every year that have the same leak and report >90% accuracies, and >0.9 ROCs.
[0] https://www.nature.com/articles/nmeth.2259
[1] https://www.biorxiv.org/content/10.1101/2021.08.13.456309v1