One idea is that the cause is batched inference in sparse MoE (mixture of experts) models.
https://152334h.github.io/blog/non-determinism-in-gpt-4/
HN discussion: https://news.ycombinator.com/item?id=37006224
Parallelism doesn't magically add non-determinism of this kind unless you intentionally build it to be non deterministic. Nothing prevents you from processing an array in order in parallel.
You would have to explicitly order the terms prior to reduction but you don't always have that level of control.
100% correct if you remove processing time from the equation.
In reality, Nvidia Cuda calculations run much faster if you let it schedule the order of floating points operations itself. This makes the ordering different from run to run.
This in turn causes the results to be non-deterministic.