Reaching new records in speech recognition
ibm.com
ibm.com
Code is at https://github.com/watson-developer-cloud/speech-javascript-...
Simple demos at http://watson-speech.mybluemix.net/
More complex demo at https://speech-to-text-demo.mybluemix.net/
I'm goin to be out some this evening but feel free to ask me questions and I'll answer them as available.
http://stream.watsonplatform.net/ (the domain that the speech APIs use) redirects to https. Ditto for http://gateway.watsonplatform.net/ which is what most of the other APIs use.
Both of the linked demos also redirect to https if you try a http URL.
1: https://www.ibm.com/watson/developercloud/speech-to-text.htm...
Is it also built for those situations, or mainly focused on accuracy in low ambient noise contexts?
I've been told that cell phones make a good mic in noisy environments FWIW.
I've never tried LUIM, so I can't comment there.
Is there a way to feed some kind of (text) dictionary to aid recognition? Or does it also need audio samples to learn from?
More details: https://www.ibm.com/watson/developercloud/doc/speech-to-text...
edit: They talk about this in the arxiv paper:
>The transcription protocol that was agreed upon was to have three independent transcribers provide transcripts which were quality checked by a fourth senior transcriber. All four transcribers are native US English speakers and were selected based on the quality of their work on past transcription projects.
>...The transcription time was estimated at 12-14 times realtime (xRT) for the first pass for Transcribers 1-3 and an additional 1.7-2xRT for the second quality checking pass (by Transcriber 4). Both passes involved listening to the audio multiple times: around 3-4 times for the first pass and 1-2 times for the second.
Upload a video, it strips out the audio, pushes to Watson for transcription, converts the result to a caption/subtitle track, and then allows people to comment on the discussion (like Google Docs).
I plan to polish it up a little and open source it soon.
Additionally, if you're using WebVVT format subtitles, I'd be interested in merging that code into the appropriate SDK. They're all on GitHub if you'd like to send us a PR: https://github.com/watson-developer-cloud/ (I'm on the team that is responsible for these.)
That said, I've tried the computer speech to text systems for transcribing interviews. And even with just one person talking they're nowhere good enough for me to use. Even budget human transcription (e.g. CastingWords) is just so much better that it's nowhere worth my time to use a machine-based system.
Also, with wavenet in the mix, no way this is used in production.
* $29.95 to register
* $9.95 per month subscription which entitles you to $14.00 worth of free transcription per month
* $3.50 per page (double-spaced, 225 words) for any pages in excess of your $14 allocation
$14 is four 225 word pages, so, 1-900 words for $0.11 per word, then $1.55 per word. Ish.
I can't see anything that says how IBMs minutes are done though, whether every message is rounded to 1 minute or not. Edit - based on a comment from someone at IBM in this thread, they're not rounded up to a minute, rather only at the end of the month is the total time rounded. I can't see how they'd be more expensive than google.