We've found most public benchmarks, especially Libri, are not that representative of real world data we see in production. Most real world data we see is a lot noisier, and has worse recording quality like low bitrates and compression from mp3 encoding.
We do worse than state of the art benchmarks on Libri Clean today, for example (I think we are around 7% WER last time I checked), but are much more accurate on real world data than models reporting 3-5% WER on Libri. This is why we want to make sure we are thorough when we report our benchmarks on popular datasets like Libri.