>the best reliability we managed to achieve was 90% with tests that run for 40 minutes, which is obviously not acceptable.
What was actually going wrong during that 10%?
I get something closer to 100% reliability, so I'm feeling a little perplexed by all of this.
Do you make heavy use of sleeps?