I've only dabbled, but I thought for a given input, with no previous context and a hotness of zero, LLMs would be deterministic - am I wrong?
Where that gets murky is… perturbation inside the models themselves, such as Mixture of Experts (MoE) models who’s more internal parameter activation is not so deterministic (and therefore the tokens they generate are not deterministic across runs)