> You will understand if you run some LLM models with a greedy sampler that does that. The text quality begins to deteriorate rapidly.
Right, I've done this, and this makes sense to me, but I'm not following how that falsifies the top probability word being the best choice in any particular instance.
"Picking only the best word at each decision point results in a worse final result" seems like an imminently reasonable hypothesis.