> LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unless you run the LLM with a temperature of 0 which would give out terrible results.
I'm not understanding how the word with the highest probability isn't the "best choice"?