With that being said, even despite the equivalent mutants, you can get a pretty good insight. Here is an example of finding from a proprietory real-time OS used in the space industry: https://gist.github.com/AlexDenisov/b5d2e23457b88813b5ab9d5d... (this is a part of an email which I could safely post online).
You’d also want to watch out for the particular implementation of sorting algos... if you swap out a stable quicksort for a nonstable bubblesort, your benchmark tests might still pass, but you could wind up with different orderings of “equivalent” items in your list before/after that could cause subtle issues.
If you have a test that truly still passes regardless of any mutation, I’d have to wonder if it should exist at all.