You need to change the temperature to 0 and tune your prompts for automated workflows.
What would really be useful is a very similar prompt should always give a very very similar result.
Your brain doesn't have this problem because the noise is already present. You, as an actual thinking being, are able to override the noise and say "no, this is false." An LLM doesn't have that capability.
It’s the same reason why great ideas almost appear to come randomly - something is happening in the background. Underneath the skin.
maybe it can work if you are running your own inference.