How does everyone think about putting GPT agent in production
I find it very difficult to have agent run reliably in my test env, let alone prod. The un-determinism of LLM got multiplied when running in multistep agents or multi agents. Anyone has experience where it works well consistently?