Why would you benchmark the LLMs for 50% success? I expect 100% success, or nearly so, to make an LLM a practical replacement for s human. 50% success is far too unreliable.
Edit: notice that I said "100%, or nearly so". I realize that 100% is an unrealistic metric for an LLM, but come on, the robots should be at least as competent as the humans they replace, and ideally much more so.