My favourite example of misrecognition is one of Travis Goodspeed's talks, where YouTube's VR output "Geek women are expensive, but not prohibitively so."
Voice recognition is OK but it falls a long way short of a usable level of accuracy, and even the accuracy it can muster goes right down the toilet if there is any background noise or music, or if the speaker is in any way unclear (accents, rushed or slurred speech, etc). There's a long way to go before you can just get a usable transcript of speech automatically.
Quite a lot of voice recognition engines seem to have been trained on thousands of hours of C-SPAN or Meet The Press or something, because when recognition conditions get challenging, some engines start to degenerate into outputting nonsense like "congress Muslims Kenya capitol great today Cheney".
There is no substitute for a human pair of ears and a lightning-fast means of text entry like a steno keyboard - nice to see Plover getting a mention in the source article too.
https://github.com/hausdorff/bangbangcon.github.io/blob/gh-p...
For instance, she'll separate out conversations:
>> Can you move the mic closer to your mouth? >> Yes. Is this better? Is this better? Okay. I will talk like this, then. >> You can move the mic. >> Like... >> Take it off the stand and hold it up to your face.
She can also figure out when something is an acronym (like LARP), make sure everything is capitalized correctly (Python, Ruby), separate out what's being said into paragraphs when the speaker starts talking about something new, and a ton of other things.
"Automatic speech recognition is not currently a substitute for human transcription, because computers are unable to use context or meaning to distinguish between similar words and phrases and are not able to recognize or correct errors, leading to faulty output. The best automatic speech recognition boasts that it's 80% to 90% accurate, but that means that, at best, one out of ten words will be wrong or missing, which results in a semantic accuracy rate that's often far lower than 90%, depending on which word it is."
(this is a subset of the answer to "Will speech recognition make CART obsolete?")
What software/sdk's have you used?
There are no automated solutions that can universally do a good job of transcribing natural speech from people who aren't specifically "speaking to be recognised", if that makes sense.
Maybe one day, but not yet. It's a problem waiting to be solved, so the reward for the first who can really crack it will be substantial.
It works to some degree, and has the advantage for the subtitling companies that respeakers are easier to train and don't need to be paid as much as a proper stenographer. The disadvantage is that the output is much slower and subject to a rather greater delay. There is still nothing that beats steno - but respeaking is cheaper and people don't complain enough about the inaccuracies.