Google voice search: faster and more accurate
googleresearch.blogspot.com
googleresearch.blogspot.com
One step closer to conversational interfaces!
There's a reporter on NPR who sounds like he always introduces himself as "han zhi lu wong". I say that to google voice search, and it corrects it to "Hansi Lo Wang".
So google voice recognition works even if you don't know what words you're saying!
If you both have learned the word, who is better at recognizing it? And who is better at learning a new word they don't know? These aren't very exacting questions because the comparison between human and machine knowing and learning hasn't been defined. Maybe the machine needs more examples and more contexts to learn a word as well as a human and correctly map to the word as often as a human, but it can also examine more examples and contexts than a human can per unit time and can do a lot of scaling in this regard using increased energy that a human cannot.
e.g. with your example, the letter pair ‘si’ is pronounced differently in Mandarin that it is in English. So it's not surprising that you couldn't write it down properly in pinyin as you don't know how to transcribe Mandarin into pinyin. But Google does.
Another example - without knowing how French spelling works there is no way that as an English speaker you could work out how to correctly spell ‘peut’ (can) just from hearing it.
Now Google recognizes I'm talking about Minot but it says Minnow back to me.
They trained on 3 million "utterances" of average duration of 4 seconds. These were distorted by noise to get 20 variations (so the training set was 60 million utterances total).
I don't understand if these were labeled somehow There's a section on clustering into 9287 phones, but it isn't clear to me if these were used as labels.
I often listen to podcasts on my car bluetooth and on a bluetooth speaker at home. On my commute, I'll get at minimum 5-15 "Okay, Google" triggers in a 50 minute drive just from people on the podcast saying things like "and" or "okay" or phrases that sound nothing like "Okay Google". I have even done the voice training so it's only supposed to listen for my voice. On the other side of the coin, I'll sit in my car screaming "Okay Google!" over and over with no response.
Shortly after I retrained the voice model in a quiet room by myself and now it works flawlessly.
It's a shame they haven't done more with it really.
http://gizmodo.com/you-can-now-type-with-your-voice-in-googl...
Is this available offline or one must be connected?
I can't say how accurate it is, since I've used it very little so far. But adding my 2c:
I first tried it some ago (2 years+) on my mid-range (at the time) Android phone, and it was not really usable. Set it aside for a while. Then tried it recently - on the same phone, mind - which is 2 or more years older now, so not recent at all. Surprisingly, it worked a lot better than earlier (based on a small sample of tests, note.) Going to experiment with it more.
Something that might be known to many readers here, but mentioning it:
Peter Norvig, Director of Research at Google, has said in the past that by training the voice recognition software on huge amounts of data (at Google scale), they have managed to improve it a lot, by using statistical algorithms. (Similarly for spelling correction suggestions in Google Web Search.)
Related: A couple of simple experiments by me with voice recognition (speech-to-text) and speech synthesis (text-to-speech) using Python:
1:
https://code.activestate.com/recipes/578839-python-text-to-s...
http://jugad2.blogspot.in/2014/03/speech-synthesis-in-python...
2:
http://jugad2.blogspot.in/2014/03/speech-recognition-with-py...
Some time in September, Google made some server-side change to Voice Search which causes the Android Google Search client, at least some versions, to crash. Android handsets get a pop-up with "Unfortunately, Google Search has stopped."[1][2][3][4]. This also breaks voice dialing and texting. Some people who had voice input as the default found they could no longer text at all, until they disabled Google Voice Search. It's not a change on the client side; it's happening even for phones that don't have over the air updates enabled.
The usual suggestions, involving clearing caches and resetting various settings, have been made, and they're as useless as usual. The problem appeared a few weeks ago, and has been reported for at least T-Mobile and AT&T, and for at least ZTE and LG phones. So it's not carrier-specific or handset-maker specific.
Did this "faster and more accurate" change involve a change to the wire protocol? A recent change is clearly crashing the client side in the phone.
[1] https://productforums.google.com/forum/#!topic/websearch/0ZM...
[2] http://forums.androidcentral.com/general-help-how/582873-why...
[3] https://forums.att.com/t5/Android/Google-Search-has-stopped/...
It took the query and showed the same photos. lol
Collecting these sort of results into a larger data set could help refine the results.
I had the same happen when I asked it to play Bulerias. Which it kept understanding as some common variation of that word, Blue rays, etc..
As soon as I provided context and said flamenco bulerias that fixed it though.
Acoustic features are generated every 10ms, but are concatenated and downsampled for input to the network: 8 frames are stacked for unidirectional (top) and 3 for bidirectional models (bottom).
"…predicted word sequence where the word with highest prob- ability is taken ignoring repetitions and the blank label with no language model or decoding."
Which I took to mean: when the acoustic model emits a blank symbol, they don't run the decoder again until a non-blank symbol comes out.
If not, what problems are still left to be solved?
- Not yet Human level accuracy even for the clean speech.
- Vocabulary limit is still an issue (but they are much better than before)
- Recognition under very noisy environments.
- Recognition in the existence of multiple overlapping voices
- Heavy accents.
- Spontaneous speech.
- Languages other than English are also not in the same level.
So no, it's certanly not a solved problem, especially if you want to use that kind of functionality outside the few english speaking countries.
And this is close talk speech. The holy grail (20 feet away at a party with error rate below human) is decades away.