For these tests, why not tune the temperature and such to reduce the randomness and convert them to almost-always-succeeds vs almost-always-fails? Is it not the iteration count that drives up the cost?
1. If you turn the temperature down too far, the output is just bad and no amount of running prompts optimization will let you hill climb your way to good performance.
2. It’s not about determinism vs non-determinism. It’s about chaos. A perfectly deterministic model is still chaotic. Meaning that very small changes to the input result in very large changes to the output.
Turning temperature down doesn’t actually get you predictable or reproducible behavior across different inputs.