Updates to Cloud Speech-to-Text and general availability of Cloud Text-to-Speech
cloud.google.com
cloud.google.com
Maybe I'm just paranoid, but I just can't imagine using a speech-to-text system for anything serious that I can't self-host. It feels like we've just seen example and example over and over again why this is a bad idea -- to the point that when I hear a company like Google talk about a locked-down cloud platform as "making AI accessible to everyone" it feels almost dishonest.
Especially once we start talking about text-to-speech. We can already do a lot of that locally - we should be pretty hesitant about coupling new text-to-speech techniques to strategies that require us to move logic away from local devices onto the cloud.
The quality is far below Google’s speech API as the model is somewhat out of date and more importantly the training data set is much smaller and less general.
The best pretrained speech to text model I’ve seen is from Baidu’s DeepSpeech 2 repository. They provide pretrained models for English and Chinese based on their internal data. The quality is astonishingly good!
Edit: both of these models can be comfortably run in real-time on a desktop. At Arm I recently worked on a project to run <5% word error rate models in real-time on a mobile phone.
My impression is that a lot of the hardware costs from modern AI comes from training and updating the model.
So (speaking as a non-expert) it doesn't seem like there's any technological problem with running a model locally on a private network -- you just need access to the model and you need someone else to generate it. And that's exactly what Google isn't providing. It's like they found a way to turn modern AI into even more of a black box.
At the point where I'm OK consuming someone else's model and not being able to control how it gets updated, then I'm probably also OK with using a compressed, static model and needing to occasionally download and deploy new versions.
DeepSpeech offers trained models that are about 70% right, but none of them use the Common Voice corpus yet. I think that plus recent changes to the codebase should produce a much better transcription, but I don't have the GPU resources to go and train a model sadly, Mozilla will hopefully release another trained model soon though!
I am working on a basic web frontend and API for DeepSpeech: https://github.com/AccelerateNetworks/DeepSpeech_Frontend
Also, here is Common Voice: https://voice.mozilla.org/en
that's progress. how long did it take you to get these things running (including downloads and dependency installation)?
I really want to put DeepSpeech 2 with Baidu's model up against Mozilla's model and see which is better, seems like it could be quite interesting!
If you sell audiobooks, then a one off cost of a few dollars to convert the book to audio form is tiny compared to the authorship of the text.
If you are doing something like turn by turn navigation, clips are typically only a few seconds long, so very cheap, and again, many of your users will be needing the same clip, so no need to pay for it twice.
While I agree that Google maps price increase was a bit much, I'm sure Google realizes they can't offer services to businesses and viably keep doing this kind of 180 too ofren, especially if they offer it as part of their Google cloud. So it's not very likely they will pull the rug under this service again.
We want to make it possible to have embedded assistants in all your objects which preserve people privacy, and do this with open-source: https://medium.com/snips-ai/an-introduction-to-snips-nlu-the...
Take a look at our blog to get started in 1h: https://medium.com/snips-ai/voice-controlled-lights-with-a-r...
It also binds in popular Home automation platforms like Home Assistant and the Jeedom platform
But the approach used in various NLU services such as Snips and RASA is much more simpler. This can work fine for easy queries but once we start asking complicated questions using conjunctions and disjunctions these systems start becoming brittle. If they try to capture all the possible logical forms through intents they’ll need an exponential number of intents and also a huge dataset to capture all the intents.
I would like to know your take on using a grammar based semantic parsing.
Will they ship it with chrome to replace the existing speech synthesis api? (I believe right now it just uses whatever voices are available to the device or OS but chrome can fallback to a serverside voice)
[1] https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_...
[2] https://developer.mozilla.org/en-US/docs/Web/API/SpeechSynth...
I tried to get Google to fix this a long time ago and it seemed to work for a while after being offline for weeks.
OSX 10.13
Model Name: Mac mini Model Identifier: Macmini7,1 Processor Name: Intel Core i5 Processor Speed: 2.6 GHz Number of Processors: 1 Total Number of Cores: 2 L2 Cache (per Core): 256 KB L3 Cache: 3 MB Memory: 16 GB
When I click the "speak" button, it makes a network request (as shown in the Chrome devtools) and immediately gets back a response that looks big enough to be an audio blob.
But most of the time, instead of actually playing the audio, it just sits there with a spinner for anywhere between 30 seconds and 5 minutes before giving up, with no errors in the console.
We tried all the software we could find to turn the recording (Dutch) into text but there is nothing that gives a helpful result.
I know that a recording-to-text is different than speech-to-text but even when I use OK Google most of the time the results are horrible.
So after all those years I am still a little skeptical.
Just one example from the preface to Chollet's "Deep Learning with Python":
> If you’ve picked up this book, you’re probably aware of the extraordinary progress that deep learning has represented for the field of artificial intelligence in the recent past. In a mere five years, we’ve gone from near-unusable image recognition and speech transcription, to superhuman performance on these tasks.
Come on, speech to text is still far from usable unless in a very limited scenarios. Why pretend it's different?
Just wait - it'll come in a few years though.
I don't trust those cloud-based solutions.