>
"The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer."Wait what? The headline is about text-to-speech aka speech synthesis, not speech recognition (speech-to-text.) Are they trying to do both? It seems to me that you'd train both using different sorts of datasets. If you wanted TTS to be intelligible to the most number of people, training to to speak like a 'perfectly miked BBC announcer' is probably exactly what you'd want to do.
Train it to recognize many regional accents, but train it to speak with the most prevalent and universally understood accent you can find. So either BBC English or Californian/Hollywood English.
Although traditionally TTS engines have shipped with numerous voices, such that you can select either a British or an America accent for the English voice. It may be worthwhile to have other English accents too, maybe one for India (125 million speakers.) But if you trained a TTS engine to have a computer amalgamation of all possible English accents I really doubt the result will be considered high quality by anybody.