It's not a strawman. There are many fundamentally unpredictable things where we can't make the benchmark be 100% accuracy.
To make it more concrete on work I am very familiar with: breast cancer screening. If you had a model that outperformed human radiologists at predicting whether there is pathology confirmed cancer within 1 year, but the accuracy was not 100%, would you want to use that model or not?