This effort would be more interesting if it could collect speech data in a more specific domain, like web search queries.
This effort would be more interesting if it could collect speech data in a more specific domain, like web search queries.
Common Voice has optional fields for age and accent in the dataset, but I don't use them. I just toss everything into a single bin.
I've had users with many accents (including e.g. thick UK, India, Germany), report that my 3k hour model including Common Voice performed significantly better for their accent than a librispeech-only model, and also much better than the older macOS speech recognition English models. (And how a model feels for a user in the real world is something you can't really get by just testing against switchboard)
I can't afford switchboard, so I'm very glad that datasets like Common Voice, librispeech, and TED-LIUM exist. I hope at some point I can get the model running in a website, it really does feel pretty good to use.
For Talon we developed a custom speech collection site, that has a lot of prompts we control: https://speech.talonvoice.com - I tried it out with the TIMIT prompt list initially, but right now I'm recording dense command-like speech, which I've found some of my more experienced users are able to say naturally/quickly, not like they're reading from a prompt.
I don't have a separate verification process, because I've had a lot of success with a process that automatically prunes inputs that "obviously make the model much worse". The site is basically designed to record as fast as possible, with keyboard shortcuts to go to the next item and start/stop recording so you can almost record nonstop, which ends up being slightly less forced than just reading one sentence.
The site is pretty reusable, I previously used it at noise.talonvoice.com to record "noise recognition" samples. Here's the source of the current speech site, if someone wants to spin up a common-voice-lite for a specific domain it's pretty easy (just need to run a python app somewhere): https://github.com/talonvoice/noise/tree/speech-dataset
Fair but statistical systems are also extremely susceptible to placebo. I would be curious what a test set shows.
The point isn't that LibriSpeech isn't clean enough. Rather, it's that conversational speech is very different from read speech (which is based on written text). Everything has an effect, even planning of utterances (think: hesitations, "uhm", "uh"), turn taking behaviors (think: how speakers negotiate taking turns), how speakers self-correct, phonetic convergence (think: speakers adapting their speech to be more similar to that of their interlocutor), and so on.
The Common Voice data won't help with that, as it's read speech. It's far more expensive to collect conversational speech datasets, as transcription (or correction of automatic transcripts) involves a lot of manual labor.
That's literally the worst case scenario. Victims escalating the violence.
How so?