I also wish people would stop using AUC, and start using a measure reflecting realistically useful specificities. I don't care if you have 99% sensitivity at 90% specificity.
It is already happening. A lot of people who can contribute to actual Science are moving into AI/ML field for the money and the industry/media hype are reinforcing this. Everything is "Deep${NONSENSE}" nowadays whether it is relevant or not. As a beginner, when i started to learn NNs, i couldn't get past my initial hurdle on how to validate the results on actual real-world data. What Statistical metrics do i use to "know" that the blackbox is working correctly? What are the assumptions and limitations that i need to be aware of to understand and have faith in the output? Most people don't seem to know or care; it is "magic" to them. In a world awash with data, reckless application of NN models to any and every problem is only going to drown us in spurious results and muddying all Scientific endeavours.
There are a lot of situations where before people would have assumed their best option is to carefully tweak a custom statistical model, whereas now they're just happy to throw a black box at it and see what happens. This is as much a cultural change as a technological change, and it's good that it is finally happening. That's what a "paradigm shift" is after all.
The goal of research is usually to rip those boxes open to figure out what's inside and how it works. Moving away from that towards opaque predictions doesn't make a lot of sense to me, especially when the predictions aren't even that much better. Plus, a lot of this work seems weirdly disconnected from what the rest of the field knows to be (im)plausible.
Obviously, black boxes can be useful tools. DeepLabCut is incredibly helpful and will save a lot of grad students a lot of tedium, and that probably wouldn't happen if it involved a lot of tuning. Predictions can also be very useful--frankly, we'll take anything we can get for most neuropsych conditions--but mechanisms and targets for intervention are so much more useful. I know there is some work on this, but it's drown out by the 0.99AUC!!1!! (in a small, cherrypicked group) stuff.
Building a custom model will help with feature selection. It will provide a baseline to compare the ml model to which can help debug problem points of the ml model. And finally it serves as a sanity check that you aren't leaving a lot of performance on the table.
If AI is subject to crazes, with investors as a whole drastically overestimating its potential, then it's certainly possible to over-allocate capital (human and otherwise) to it in the hopes of a payoff.
Consider the Dutch tulip mania of the 1630s. Imagine if it had lasted a bit longer, long enough for promising scientists and scholars of every type to be trained solely to optimize the growth of tulips.
This allocation of capital would provide a benefit to tulip investors for as long as the craze lasted, but would prove to be a detriment to society once the craze ended.
'AI' is just math + programming. Don't overthink it.