the beginning of the arguments was good but the "because of randomness" is wild
Have you ever tried reproducing even a small neural network exactly if you train on GPUs on more than one machine? I have and it is pretty close to impossible, and I'd argue actually impossible at scale.