Amazon Transcribe Streaming Now Supports WebSockets
aws.amazon.com
aws.amazon.com
Performance-wise they seemed mostly similar - both can pick up text fairly accurately in somewhat noisy environments. We did notice that Google's offering seemed to be slightly more accurate. But the difference was marginal and only noticeable when we compared side-by-side on intentionally poorly pronounced words. There were two big differences that made it impossible for us to use Google's product:
1. Google has a 1 minute limit on streaming speech-to-text. It just closes the connection at 1 minute. It doesn't even send you a "final result", so if speech was being recorded at the time the connection dropped, that transcription is lost. Speaking of which...
2. Google doesn't provide incremental updates. So if someone speaks for a while, you only get an update at the end of it.
Note that this is their API - my impression is the product they use in their apps is superior in functionality to the product available on Google Cloud.
Amazon Transcribe, on the other hand, has a 4 hour limit and sends incremental updates, so in a longer sentence, like this one I'm currently writing, you would get a message every couple words, which is essential when the goal is to show a live transcription.
Also, they definitely have incremental updates.
Can you explain this comment? From my recollection, this statement is not true. Also, their site says:
"Returns text transcription in real time for short-form or long-form audio Cloud Speech-to-Text can stream text results, immediately returning text as it’s recognized from streaming audio or as the user is speaking. Alternatively, Cloud Speech-to-Text can return recognized text from audio stored in a file. It’s capable of analyzing short-form and long-form audio."
Probably best explained as an example. Consider someone saying this:
"How are you doing? I'm doing well. Do you plan to go to the park with Alice and Bob tomorrow to see the fireworks show?"
When we were using it, Google would give:
1. "How are you doing?"
2. "I'm doing well."
3. "Do you plan to go to the park with Alice and Bob tomorrow to see the fireworks show?"
as three separate messages.
So it's certainly not waiting until the end of the entire streaming to give you a result - it sends you those three messages as you speak. In that sense it is "returning text as it's recognized", because once it recognizes you've finished a sentence it computes the words and gives them back to you. But the issue for us was that we could only get results after full sentences or long pauses.
Amazon, on the other hand, would give for the last sentence something like:
1. "Do you plan"
2. "Do you plan to go"
3. "Do you plan to go to the park"
So for our purposes (showing a chat bubble above a person's ahead), the latter version was much more useful. The reason is that in spoken language sentences often ramble, so we wanted to be able to show some incremental updates to the user as a long sentence was spoken so they wouldn't have to read 30 words at once.
Disclaimer: My comment reflects my own views, and not those of my employer, etc.
Also one thing I liked about it IIRC is that Amazon was the only service at the time that offered punctuations (to mark end of sentence) in English which is very useful in some cases.
Can test Speech reco API in the page here: https://azure.microsoft.com/en-us/services/cognitive-service...
Google has a cloud text to speech API that supports streaming audio. Google blows Amazon out of the water here with accuracy, speed, and features at the same price or cheaper. They also have translation APIs.
I'm more than happy to help out if needed!
The browser highlight->speech plugins have always been a bit iffy.
I personally hate it when I have to use a service for something that could be done locally on my computer or smartphone. And I don't get that fuzzy magical feeling, but instead I think of a (very nearby) dystopian future where a single company knows what all citizens say or do in real time.
Needless to say, I didn't read the rest of the article.
Then a few paragraphs later:
> For real-time transcription, Amazon Transcribe currently supports British English (en-GB), US English (en-US), French (fr-FR), Canadian French (fr-CA), and US Spanish (es-US).
So basically it boils down to two English variants, two French variants and US Spanish variant.
And then one wonders why such projects never pick up steam around the world.
Actually I must admit that at least they did look beyond North America for the 2nd English and French variants, and not the usual American English and the rest comes when it comes, if ever.