> An ideal dataset would consist of thousands of hours of speech where source accent utterance is mapped to each target accent utterance and aligned with it accurately.
To put it in terms of text translation, roughly how many sentences or words is this?