40+ voice searches thrown at Google on Jelly Bean
plus.google.com
plus.google.com
In many of these demonstrations I noticed one thing that was bugging me: Even though the voice recognition in Android seems really cool (I don't have JB yet), it doesn't give definite audible confirmation of the command in many cases and sometimes it even requires user interaction with the screen.
Now personally, I already believe that speech input is kind of a gimmick in itself (try using english voice recognition with my address book filled with german names...), I believe that to even have a chance to move from gimmick to useful feature, it must work without user interaction on the screen.
"Play <whatever band name>" followed by "beep" and the a button to press on the screen doesn't help me. A useful response to "play <insert band name here>" is "playing <insert band name here>" followed by actually playing it.
Or "call <some name>" - if you just get back "calling" or even just a beep - how would you know whether the recognition was successful or not and the correct name has been recognized?
Some commands on Android seem to be doing fine (the weather example), but others fail in one way (the play example seems to require user interaction on the screen) or another (the "turn on wifi" command doesn't produce any audible confirmation or error message - just the same beep sound as if it worked).
Siri, while it might not have as good a recognition as the Android solution, is much better in that regards: It always confirms your command. As such Siri seems moderately more useful as an additional input method whereas Android, by forcing you to look at the screen when inputting a voice command, reduces this to a gimmick and nothing else.
I don't have this either, but can tell you that Google's voice recognition does have a confidence rating on a per word basis. (See Google voice messages. Also use the recognition API and it provides alternatives.)
In their navigation product they directly do whatever was said if there is a high confidence. If not it shows the recognized speech with a "pie" based countdown to using the displayed recognition. You can press OK to go ahead (or wait), or cancel/try again.
They could obviously do something similar with this.
They also use context for their voice recognition. I grew up in a town named Piggs Peak (note two 'g's). If you say "piggs peak" you will get that spelling, but for example saying "peak pigs" gets you the spelling with only one g. This explains why wooster/worcestor doesn't confuse them. I don't have a siri capable device so I don't know what they do.
It doesn't require any input though - I just tested it. Once the progress bar reaches the end (it seems to take ~7 seconds) it will complete the action.
If it immediately confirmed like Siri, you would know right then and could re-issue the command.
You do not have to "wait ~7 seconds before you notice that it was wrong" so you can "re-issue the change". The result is displayed immediately. You can cancel the auto-action (which you first said didn't even exist), or force it through before the ~5 seconds (not ~7) elapses.
Go re-watch the entire video in the foreground, please.
"Call the Drake Hotel in Toronto."
*bling* "Calling..."
(wait seven seconds)
"Hey, this is Drake. What's up?"
Versus what Siri does: "Call the Drake Hotel in Toronto."
*bling* "Calling Drake Smith..."
"No, wait! Stop!"
Think about using the voice commands when you can't see the device. Like when you're driving or running. It's useful to have the audible feedback in addition to whatever's displayed on the screen.This is the correct solution, IMO. It would be quite frustrating to have the wrong phone number instantly begin to dial, for instance. One time when I said "call <name of restaurant>", it came up with "Call <name of restaurant>" with the address of the location I didn't want shown beneath. This gave me time to tap Cancel, which then showed me a list of the alternative results/locations.
Your phone understands this as "Call Bar Burgers" and shows on the screen "Calling Bar Burgers". The phone makes a "beep" sound and then proceeds to show a progress bar which you don't see because your phone is in your pocket.
Then the phone connects and you learn of your mistake as the person at the other end answers with "This is Bar burgers, Mr. Foobar speaking".
The only way around this is to enable the voice command, take the phone out of your pocket and then check what it says above the progress bar.
With siri, if you say "Call Foo Burgers", Siri would respond (in audio over your headphones) with "Calling Bar Burgers", giving you a chance to cancel before you annoy the person at the other end and without forcing you to take the phone out of your pocket to check (which is the point of voice commands)
I use Voice commands for pretty much all input to Google Maps and Navigation (and nowhere else). That it works flawlessly in my experience even with my mumbling is plenty good enough.
The point of voice commands, when I use them, isn't to avoid any visual interaction, it's just so I don't have to type something I don't want to type.
Even better, Siri gives confirmation by default, but you can disable audible feedback if you so prefer.
I don't understand how parent can write such a lengthy comment without watching the entire video and understanding what they're talking about. Even the parent's disclaimer states they are only going by the video, but it clearly shows they didn't watch it fully, since what they missed is contained within the very first minute of the video, multiple times.
"How much is Angelina Jolie worth?" is throwing an answer: https://www.google.com/search?q=how+much+is+angelina+jolie+w... But "How much is Brad Pitt worth?" isn't: https://www.google.com/search?q=how%20much%20is%20brad%20pit...?
"Fish species in Lake Tahoe": https://www.google.com/search?q=fish%20species%20in%20lake%2... But "fish species in mississippi river": https://www.google.com/search?q=fish%20species%20in%20missis...
Edit: changed ".es" by ".com" in all the links.
The other day, I told my phone "Call Nico Thornley." Instead, it searched for "call me-so-lonely." Not the first time Google Now has completely botched a friend search either.
Same for the race to having better maps or the better browser.
Is this what you mean? http://www.youtube.com/watch?v=0L_IhqGcRM8#t=8m36s
Here's the navigate by voice from nearly 4 years ago http://www.youtube.com/watch?v=jLXZ5BHeDFg
Here's the original voice actions video by Google from 2 years ago http://www.youtube.com/watch?v=gGbYVvU0Z5
Siri was introed 1 year ago?
I agree though, it's a great time to be a gadget consumer. I love Siri's conversational style. It actually doesn't seem too far off before I can actually start having a conversation with a computer though something like Siri which is both awesome and terrifying at the same time.
Here's what I misunderstood - I thought Jelly Bean was the same, but he would always just tap away the last item in the video so we don't see it. Maybe there is no scrollable backlog in Jelly Bean.
Rounded to 2 decimal places: 11.65600 is 11.66 not 11.65 80.4672 is 80.47 not 80.46
Sorry to be pedantic...
"One of the biggest issues with Siri is that it requires access to Apple’s servers in order to work. In Jelly Bean, however, Google will provide full offline voice dictation to users. Granted, that’s not a full Siri competitor, but the fact that the search company has been able to take it offline in a mobile setting is very important."
http://www.eweek.com/c/a/Mobile-and-Wireless/Android-41-Jell...
If it was then surely you'd do all the voice parsing locally so I'm guessing that it's not. Unless anyone can think of another reason you'd push it through the servers?
- online should have more training data and can be improved much more easily. Don't see that advantage going away.
- power efficiency and/or speed. Sending 5s of audio across the net can be less strain on the device (esp. older ones) versus parsing the audio locally.
Thinking about this, do systems like Android typically offer localization that specific?
Edit: or Vatican City?
Edit: Also, he says "Where is that museum with Egyptian stuff in San Jose?" and after he closes it, it shows "where is the tallest building in the world"
Disclaimer: the only edits I made were to cut time between each of my queries, as well as re-order some of the demos from the original order I recorded them in, so they would fit into categories. None of the queries themselves have been edited or cut down, and the sequences are intact. The processing time happened exactly as you see. This demo is made on the early build of Android 4.1 (JRN84D, takju build for Galaxy Nexus I/O edition), on a wifi connection. Consider this beta.
So what you're seeing are places where the queries were reordered.
That could explain why some of the results have other things shown when he closes them.