Google’s new voice recognition system works instantly and offline (Pixel only)
techcrunch.com
techcrunch.com
Recurrence can help with robustness in some other very important ways as well.
Citations for this dates from the 80s and 90s. I don't know the best reference offhand. You could look at some old Hinton stuff if you're a fan. Lots published on this.
Nothing like this has been published AFAIK.
After you have the results of this experiment you can try to explain them with attractors and what not, but I would be surprised if there was much difference. Would make a good paper though!
Unless you're dealing with opinions, I disagree. The onus is on the person trying to give evidence to actually give evidence.
Also if you haven't looked into the properties of how exactly a RNN Transducer functions, I highly recommend doing so. They help resolve a great deal of problems that traditional RNNs and CNNs are unable to deal with.
The primary reason to be interested in convnets for speech is computational parallelism, not because they have especially strong results for accuracy.
I work in the field, a more accurate summary would be that there are a number of viable architectures that currently get fairly similar accuracy, but that have other pros/cons with respect to streaming, memory use, parallelism, model size, integration with external language models and context, complexity of the decoder, friendliness to different types of hardware etc.
- No comparison is given of number of model parameters. If optimizing strictly for model size, RNNs tend to be nice and compact.
- The computational advantage of the CNN at training time is throughput. The advantage of RNN at decoding time is streaming latency. Running the CNN frame by frame as they are received removes the ability to run frames in parallel and if the CNN is larger, it will run slower, and depending on its receptive fields it may not even stream well at all.
- That particular CNN system uses a strictly external LM that is not jointly trained and has an additional hyper parameter at decoding time to weight the LM that requires additional tuning.
- It is still autoregressive in the beam search, so the LM will still be run many times sequentially adding tokens just like an RNN LM, and is likely to be more expensive. The throughput advantage a conv lm has in scoring whole sentences is totally lost. In fact, there doesn't seem to be anything special about the choice of a conv lm for that paper except that it is fun to make all the parts convolutional.
- CNNs frequently require more total flops, but are high throughput on eg a GPU because they expose so much parallelism. On an embedded CPU this can be a bad tradeoff.
As a side note, there's no reason that CNN architecture, which in the paper is trained with a close relative of CTC and is decoded identically to a RNN CTC AM + external LM, couldn't be trained as an RNN transducer. Despite the name neither the am nor lm have to actually be RNNs.
The RNN-T is a nice idea though, if I understand it correctly it's another approach to the alignment problem. In CTC, you are generating sequence like TTTTHHHEE CCCAAATT, which mean that your language model must deal with these repetitions, and you can't train using text without repetitions. In RNN-T you are learning to advance the cursor on either audio sequence or text sequence so as to maintain alignment, kind of like you do when you merge-sort two sorted lists, therefore it outputs THE CAT, and you can use a standard language model.
Gated convolutions as LM is similar to RNN-T idea [1], but you have to deal with softmax, so I'm not sure how well this would work in practice, especially on a mobile processor.
Hoping paper details suffice and help with the parameter search, and happy to respond to questions over e-mail. Would love to see an open-source implementation with local or directed attention built out!
const mapTrans = fn => function* (x) {
yield fn(x);
};
const filterTrans = predicate => function* (x) {
if (predicate(x)) {
yield x;
}
};
const dupeTrans = n => function* (x) {
for (let i = 0; i < n; i++) {
yield x;
}
};
Clojure just observed that an isomorphism of these functions under Church encoding looked like f<B> -> f<A> for a special parametric type f, so could be composed with ordinary function composition, albeit backwards.An RNN transducer is fundamentally three functions:
One takes a list of recent As to some C1.
Another takes a list of recent Bs to some C2.
A final one takes a C1 and C2 to a B.
Rather than mapping over each input independently the RNN-T is allowed to learn something about the relationship of recent outputs to the next output, and the relationship of nearby inputs. Clojure transducers thus have an order insensitivity that RNN transducers are allowed to be sensitive to.
While offline, you might write email drafts, your blog, or even a book:
https://medium.com/@augustbirch/what-i-learned-writing-an-en...
What's missing is the ability to make edits using your phone. You can probably speak at over 100 words a minute but then you need to stop to bring up the software keyboard.
And then for someone on an expensive metered or slow connection - which is a lot of people in the developing countries - they might not want auto-update at all. So if they didn't notice expiration, they'll find out that their maps aren't there next time they try to navigate offline.
Seems pretty reasonable to me, after all roads do change over time.
My main motivation is to make nav less bandwidth intensive (I pay per GB with Google FI) and to ensure I have maps even if I don't have a good data connection.
It might have something to do with licensing. Either way, OsmAnd doesn't have that problem.
With that said I especially like the Google Maps offline features which have been added recently. You can even have it calculate driving directions completely offline if you have the starting and ending addresses.
Do you have a smartphone? Because that's most likely an Apple or Google surveillance device.
Well, on macOS offline voice recognition is actually much slower than online. Not to mention the choice of words and Vocab is quite limited. I love to get an offline version, but so far every online version seems to be better.
Interestingly may be only for English. In my experience Apple is doing far better in Japanese, Chinese ( Both Mandarin and Cantonese )
There are lots of ways to spin this, but I see it as a significant improvement for any app that could benefit from voice input. It's immediate and not susceptible to network glitches. The benefit for Google, IMHO, is primarily more sales of updated Android devices.
Gboard > Voice Typing > Faster voice typing
It says its an 85MB download for US-English
I dunno, Android and a lot of Google's mobile apps that aren't about online communication work fine offline. Actually, a lot of the online communications ones do too, as much as is even conceivable, they just don't transmit and receive offline, because, how would they?
This is translating what you said after the wake word from voice to text on the local [Pixel] hardware rather than sending it into Google's Cloud.
The biggest benefits here are speed and reliability. It could also handle some actions offline.
It could be used to improve privacy, I just don't know if it will be used that way.
However, as you said, if this is always a requirement then it doesn't affect privacy at all, which to me would be a real shame but this is Google after all. We just have to wait and see for now.
Typed from my s̶u̶r̶v̶e̶i̶l̶l̶a̶n̶c̶e̶ ̶d̶e̶v̶i̶c̶e̶ smartphone.
I generally think of Google the same way I think of the NSA. If they stop doing something invasive, either it didn't work, they found a better way of doing it, or it was transferred to a legally distinct category, and we only hear about it because of PR considerations.
If you have an android device with Google services and a firewall, you'll see that the device is constantly phoning home, which is also noted in the privacy policy.
This does nothing for privacy, rather than provide the illusion of privacy.
I'm mistrustful of Google's privacy stance, since they have a history of changing their privacy policy, then misleading users about it. Remember when they implemented personally-identifiable web tracking and sold it to users as "new features for your Google account"? Merging Doubleclick's tracking data with my Google account doesn't seem like a feature to me.
[0]: https://venturebeat.com/2017/04/06/following-apple-google-te...
[1]: https://www.propublica.org/article/google-has-quietly-droppe...
The Verge says it may reach other devices later.
It sounds like it's both better than the old dictation model, and significantly smaller.
The thought that every interaction with my phone is being streamed in realtime to a third party server freaks me out.
Kudos to Google for working on this.
You want an open source solution, not just an offline solution.
Yes, we want an open source solution, but I'm not going to work on it. So who's going to work on it? Are you?
In absence of resources working towards the ideal, I'll applaud any step in the right direction.
Nonetheless we all benefit from this progress
EDIT: Looks like something was added in Jelly Bean: https://stackoverflow.com/questions/17616994/offline-speech-...
This was before the cloudamagig, so I wonder it ran on.
Edit: found the link https://www.techrepublic.com/article/solutionbase-using-spee...
I'm hopeful that voice recognition assistants will help the burden during Christmas visits :D
The paper on arXiv goes into how they deal with this. Basically they run a traditional WFST decoder over the output of the RNN-T to take spelling context into account. Still, it's impressive how far the neural system can get with no explicit lexicon or acoustic modeling in general.
https://www.google.com/amp/s/www.howtogeek.com/340108/how-to...
Of course there will still be usage analytics, etc. But it does increase privacy to some degree, especially when compared to sending all audio after a certain phrase is mentioned.
We have to wait and see. I'm sure we all look forward to a completely offline solution.
But, enough of that. I'm holding out until decent voice dictation is standard everywhere and a well understood engineering problem with good open source implementations.
Mostly so I don't have to type address into my car's GPS.
I guess another assumption could be "they have more than enough voice data to train any future network improvements." While I feel they have a lot of voice data, I am skeptical to say that they wouldn't view more data as useful, or at least potentially useful.
I cannot think of a compelling reason for why they would stop collecting data all of a sudden. Can you? (Serious question, I don't mean to sound snarky)
And if you're uploading the text, then it doesn't really matter where the speech to text translation happens.
There is a machine that can work totally offline, listen to audio, transcribe it, have a basic understanding and blast me with ads everywhere I go in the digital universe.
It can then psycologically slowly manipulate my behavior via ads making us buy/do things that we don’t even realize it.
It’s gonna be a scary world for my kids.
I'm constantly dealing with the phone interpreting commands intended for a Google Home speaker, which sometimes results in both the speaker and the phone acting on the same command. To my dismay, there's no way to disable Hey Google recognition on the phone after it's been enabled.
Perhaps someone here has run into this issue as well? It's a huge pain point for me.
https://venturebeat.com/2017/10/19/how-googles-pixel-2-now-p...
it is quite good, and very fast. but it's still not there. it has trouble with nuances like "call" vs "called" -- can't hear that suffix very well in regular speech. for me, it also has a really hard time with pronouns.
many times I'll start off with regular speech, go to look at what was transcribed and notice a couple errors that would make me look like a fool, backspace the whole thing, and then repeating it all gain in a very robot like voice.
it's almost there.
Just a heads up, Google Voice is the name of a product that offers telephony service including SMS, and has been around for a decade or so.
Sure it might work better now, but that's expected when computers are much more powerful than a pentium 100 with 32MB of RAM. Uploading voice to google servers for processing was always just a data grab.
GBoard works somewhat reliably with 0 training and my Italian accent, that's an apple to oranges comparison.
The training was in part due to the hardware limitations. Some versions let you skip the training or have only a small training session and let it learn as it went, I'd be surprised if googles systems weren't doing some sort of continuous learning too.
> and then let another person speak, possibly someone with a different accent.
It had profiles to handle that, although it needed training for each profile. With google does it automatically switch profiles or does it mess up what it's learned about you're voice? It's very opaque.
> GBoard works somewhat reliably with 0 training and my Italian accent, that's an apple to oranges comparison.
I haven't tried the new system but all of google voice recognition has struggled with my Australian accent to the point of being unusable.
They were invented in the 80's, they weren't used because of limited processing power.
You had to spend an age just training it to your own specific voice and choose microphone carefully, neither of which are steps this requires. Not to mention the syntax for punctuation etc.
So no, I would not argue this is catching up the the 90s.