I wonder how “real” ML people deal with the stochastic/gradient results and people’s expectations.
If I do ordinary software work the thing either works or it doesn’t, and if it doesn’t I can explain why and hopefully fix it.
Now with ML I get asked “why did this text classifier not classify this text correctly?” and all I can say is “it was 0.004 points away to meet the threshold”, and “it didn’t meet it because of the particular choice of words or even their order” which seems to leave everyone dissatisfied.