Yeah, I would agree that the academic datasets are not representative, especially for STT. Our models do around 8.7% on LibrisSpeech Clean. But on our internal dataset PaddlePaddle's WER is 29% whereas we do around 15%. And we regularly see higher WER's in production, especially for accented and noisy files. Hopefully continuous re-training will help improve the generalization. Here's the output from our model the 6063 file.
0:00:00.7 S1: I'd say there's a such thing as eating too much, but I just have a massively fast the table of them and so, I constantly eating so.
Do you do diarisation and punctuations as well?