LLMs are, in theory, deterministic. Sampling is not intrinsic to LLMs.
Greedy decoding a single batch in most libraries will give you mostly deterministic outputs. Higher batch sizes can increase variance.
But all of this is down to CUDA and/or kernel implementation issues.