Three problems with this:
* salespeople constantly try to sell the automation as more complete than it is
* product owners try to push us developers into making it more fully automated
* users get lulled into thinking it's more complete than it is (and accepting suggestions instead of deeply thinking through the issues like they would if they had to think things from scratch)
Maybe fixing management is the more pressing issue then working on the task of selfreplacement in the name of profit for others. Thinking about it, the implications are interesting. What is the energyconsumption of a human thinking in comparison with the energy requirement of a possible machinic replacement?
I.e. for classification you can judge "certainty" by the soft-max outputs of the classifier, then in the less certain cases can refuse to classify and send it to humans.
And also do random sampling of outputs by humans to verify accuracy over time.
It's just that humans are really expensive and slow though, so it can be hard to maintain.
But if humans have to review everything anyway (like with the EU's AI act for many applications) then you don't really gain much - even though the humans would likely just do a cursory rubber-stamp review anyway, as anyone who has seen Pull Request reviews can attest to.
LLMs are able to counterfeit a truly impressive number of indirect signals which humans currently use to make snap-judgements and mental-shortcuts, and somehow reviewers need to be shielded from that.