So what if the output is stochastic? LLM's have self consistency, so you can repeat the inference several times and pick the most frequent output.
They can't even perform basic arithmetic (which is not surprising since they operate at the syntactic level, oblivious to any semantic rules), yet people seem to think offloading more complex tasks with strict correctness requirements is a good idea. Boggles the mind tbh.