Isn't that why we look at best/average _and_ worst case? Not only the worst case alone?
Also, Big-O notation can primarily be used to see how algorithms scale with larger datasets. Not to see which algorithm is faster. (Although in a lot of cases, the better O-notation algorithm is also the faster one.)