I'm not claiming that's what happened here, nor am I interested in nitpicking "what counts as 'science'". I'm just saying this is a reasonable thing to do.
I'm not claiming that's what happened here, nor am I interested in nitpicking "what counts as 'science'". I'm just saying this is a reasonable thing to do.
But this is, of course, 1000 times more expensive to do. And if you only train 100, or 10, or 1 model, then the deduction becomes increasingly unstable.
So from a practical point of view, it's probably not feasible, because you would put those resources into something else instead that has more ROI.
Your suggestion of running 1000 training runs with different subsets of data sounds excessive and unnecessary to me.
It really depends upon the data. A smaller set of data that mostly consists of "truth" might be better than a larger dataset that also has many "lies".
Perhaps what you mean is that the model might be more representative, rather than _better_.