> Run the same agent n times to increase success rate.
Are there benchmarks out there that back this claim?
Are there benchmarks out there that back this claim?
Unfortunately, undesirable in practice due to people being token-constrained even before. One case is retrying only on failure, but even that is a bit tricky...