Dunno why (probably dataset) but open source Speech Recognition models are performing very poorly on real world data compared to google speech to text or azure cognitive.
Most open source models are trained on Libri, SWB, etc. which are not really big or diverse enough for real-world scenarios.
But to max-out results the devil is in the details IMO (network architecture, optimizer, weight initialization, regularization, data augmentation, hyperparam tuning, etc) which requires a lot of experiments.
Google is training on datasets that are as big as 30kh and MS seems to work on a 10k h dataset.
At the moment, I am working on a similar e2e system but 80h big dataset makes it a really challenging task to generalize well.