25 Years In Speech Technology and I still don’t talk to my computer
matthewkaras.medium.com
matthewkaras.medium.com
Last weekend I had the following failed queries:
"OK Google, what is the air quality like at Mt. Shasta today?"
"OK Google, add a waypoint for the last gas station before the mountain pass"
"OK Google, what percentage of people can you detect to be wearing masks on recent Instagram photos tagged at a location within a 10 mile radius of Mt. Shasta?"
These are all things I would expect a computer assistant to do really well. They have access to so much data, and so many APIs, that they should be able to break down these sentences into a SQL-like query and give me results. The third, for example:
"recent Instagram photos" -> Instagram has an API
"tagged within a 10 mile radius" -> parse the cities within a 10 mile radius and look for tags in all of them
"people" -> use your wonderful person detection networks your friends at Waymo developed
"wearing masks" -> I'm sure your internal datasets have this label, so run an object detector
Then compile and reduce the data to give me the number I want.
That's what I want an assistant to do. But it couldn't even do the first, which just involves a single API query to fetch air quality index information. Bleh. And as for the second, it has no idea what "last gas station" or "mountain pass" means; it's a query a human would know to be extremely commonplace.
It turns out that the current generation of "assistants" are mostly just template-matchers which really doesn't help me much at all. I can set my own alarms, thank you.
imho they're just used to give google/amazon a few more data points for their graphs
I sometimes feel that the trend to prefer NLP over formal query languages is comparable to the trend to prefer GUIs over consoles in the 80ies and 90ies.
I'd rather have an assistant type system with a fairly well defined query system that exposed its capabilities and limitations directly, rather than me having to guess at the corner cases and failure points.
Disclaimer: I work @ Google on display assistant devices, but I don't work on the actual assistant interaction pieces.
People like to reference Star Trek for stuff like NLP queries, but if you go back to TNG and pay close attention to the verbal queries to the computer, much of the time it isn't natural language. They seem to actually use some sort of formal query language that fits English a bit closer, but is still distinct from when the characters speak to each other.
"Computer, begin auto-destruct sequence, authorization Picard 4-7 Alpha Tango."
- Wake word. Command. Authorization stanza. (I bet the computer would prompt for authorization if missing.)
"Computer, Commander Beverly Crusher. Confirm auto-destruct sequence, authorization Crusher 2-2 Beta Charlie."
- Wake word (possibly superfluous). Identification stanza (probably superfluous for the usual crew, but I can see from an HCI perspective that you might want to make people provide it specifically for such a consequential protocol, and it may also be of merit if some random admiral usually halfway across the galaxy pops in to confirm). Command confirmation, authorization stanza.
"This is Captain Jean-Luc Picard. Destruct sequence Alpha-One. Fifteen minutes, silent countdown. Enable."
- Computer is very awake at this point, no wake word. Identification stanza. Sequence parameter. Time parameter. Verbosity parameter. Commit.
In any case I think this kind of speech is formulaic for the benefit of the audience, most of all, who are made aware through the formality that the speaker is addressing a machine. Additionally, we're watching navy men and women in space, so we expect them to speak to each other and to their computers in a formulaic manner ("Deck 5! Report!" etc, I can't think of good examples, brain's too tired).
Or perhaps the idea is that Trek AI is not really as advanced as to be able to understand natural language and that makes Data such a unique specimen.
Then again, there's the example of the Doctor in Voyager. I'm confused, I admit.
In any case, Trek isn't super internally consistent, anyhow.
In those aspects of UX, voice interfaces have the same drawbacks as console apps when compared to a good GUI.
Also, they have to work within the "bandwidth bottleneck" of audio - just imagine a phone system that tells you all the options you have, "Press 1 for something, Press 2 for another thing..." - they are so annoying because they are slow and inherently linear; a GUI can show the same options all at once, and you can read it much faster than listen to them.
So NLP as such is not dramatically more user-friendly unless it is at the "do what I mean" level which likely requries full human-level general artificial intelligence; before that it's just a voice equivalent of a console app, sharing all the problems of discoverability and needing to remember what the system can do and how should it be invoked.
They're even slower now because the brain trust decided adding voice control to the phone menu system was a great idea. So before, it said "For prescription refills, press 1." Now I have to wait for "For prescription refills, press 1, or say prescription refills." How on earth does that improve anything? I can just as easily press 1 as I can say a word, and when I press 1, there is a near 100% chance that the computer on the other end will understand my command.
Some phone menu voice automation are even worse. "Tell me what you want! <silence>" Then you say something, and it says "I didn't recognize that. Please tell me what you want!" Then it fails again and says "I didn't recognize that. For prescription refills, press 1, or say prescription refills..." Oh great so there was a menu? Why did you waste my time earlier?
Voice is just a terrible, low-fidelity, low-bandwidth way of commanding a computer. You might as well have handwriting input while you're at it: You write what you want on a piece of paper, and hold it up to the camera and the computer tries to figure out what you wrote. Just as silly.
So, I would rarely say something when I could push a button, but when using a smartphone on a call, it's not always easy or obvious how to push a button. Some people may have mobility issues making it hard to push a button, or be on speaker phone far away from the buttons. Or, maybe they haven't updated their telephone equipment in 50 years, and only have a rotary dial. Or, maybe on a terrible VoIP system that can't manage to get the tones through.
There's probably some way to clean up the script.
"(Please listen carefully, as our options have changed.) Please choose from the following options: Say prescription refills or press 1; say insurance denied or press 2; say referral to veterinary care or press 3"
I could get behind voice interfaces for more things if the commands words were documented, clear, and consistent, and the damn things worked. Until then, buttons seem good to me.
"Welcome to FedEx. [... blah blah blah ...] Tell me what I can help you with today."
"a package"
I mean, what the hell else can you help me with today anyway?
Ironically I used to love me some graffiti on PalmOS and Google Handwriting Input on my Droid 2, but I agree with the spirit of your comment.
foo --<tab>
You're now present with a list of options and depending on your config, the man page one liner descriptionsGUIs are for discovery, CLIs are for power via composability, voice/NLP assistants are for convenience.
Was it? Even if you discard stuff like tmux as already a GUI, you can still send whatever is running at the moment to the background with CTRL-Z and typing "bg" on any modern Unix system. "jobs" will then list all your processes, and "fg <ID>" will bring it to the foreground. I am sure this functionality predates most modern GUIs.
You may be thinking of DOS, which yes had almost no multitasking ability available.
However there were multiple timesharing operating systems that existed before the PC and GUIs, Unix being the most famous and still around.
Multitasking is quite possible on a Linux console for example. It has 5 or more consoles, each handling different users, each being able to be split via screen/tmux. Each shell can run jobs in the background as well.
1) I get the correct response; Assistant first asks me if I want to use "air check", I say yes, and get the correct response.
2) I get appropriate responses when I ask "What is the closest gas station to Mt. Shasta." Because the approaches are mostly from Rt. 5, you'll have no problem getting a usable response. Another approach from Rt 89 exists, so I don't know how you expect Google to know which "mountain pass" you mean. However, should you ask Google "What is the closest gas station to Mt Shasta on Rt 89", you will get an appropriate response. Should you be approaching from the opposite side of the mountain, unlikely though that is, Rt 5 or Rt 89 would still be the closest
3) I don't know why you expect a computer to answer this question well at all. I'm not aware of an API that would track this sort of thing. Google doing this itself, at scale, for any popular destination would be a very large project in ML with a very high error rate: Instagram photos are unlikely to be representative of a whole, the data set for a given location may be sparse, and especially for an outside location it is entirely likely (as I did at a local pumpkin farm) that people at an outside tourist destination will move away from a crowd, remove their mask, and take a nice photo safely, while people in close proximation still adhere to appropriate social distancing. Effective use would also need to be real-time: from day-to-day a given location could attract people that wear masks, and the next day a large group does not, especially in the > 300 square mile area you specified. It is not reasonable to expect this question to be meaningfully answered by either a human or a computer.
Yeah, I don't. I just get "Sorry. I could not understand that." if I remember correctly.
> 3) I don't know why you expect a computer to answer this question well at all. [...] Instagram photos are unlikely to be representative of a whole
Sure, but I defined the query pretty clearly, and I just want to answer the "general" question of "Are people generally wearing masks in that area or not" and I already babysat that question into a data query that could use the Instagram API plus some object detectors.
I understand it's hard for one person to write code that could formulate and piece together that graph, but I feel like it should be a tractable engineering problem for Google.
If you expect more, than your problem does not lie with voice assistants, it lies with search technology itself. "Expecting" these questions to be answerable is unrealistic given current capabilities. Working backwards from the UX would produce nothing better because your expectations are thwarted not by poor design, but by the limits of state of the art technology.
Actually, in the Boston area, it's "Route 128." Despite it never being indicated as such at exits or on maps.
You have omitted the rather important qualifier “Southern” from your description, as that usage is a key differentiator between Northern and Southern California.
I moved to Northern California (Santa Clara County) in the late 80s and I still use the same monikers. If I take a trip to Big Sur, I tell folks I am going to take the 17 to PCH and drive down. Old habits die hard.
https://www.kcet.org/shows/lost-la/the-5-the-101-the-405-why...
— https://www.kcbx.org/post/how-you-refer-us-101-says-lot-abou...
“Playing <super specific punk rock song>”
“Alexa stop. Alexa play punk rock playlist”
“Cannot find punk rock playlist”
“Alexa play punk rock 00’s playlist”
“Can’t find”
“Alexa play early 2000’s punk rock”
“Playing punk rock 00’s playlist”
It’s like she’s trying to mock me.
Me: "Alright man, I'll tell you where is money ... ALEXA CALL THE POLICE!"
Alexa: "Shuffling songs by The Police"
* EVERY BREATH YOU TAKE plays as I get punched 24 times *
from https://twitter.com/ppathole/status/1092034892249079813
Drug addiction makes people do crazy things, but there's a reason why most burglars flee as soon as they are noticed
My fav is when she decides to interpret your words instead of playing the exact playlist you’re telling her to play.
“OK Google, what's the speed limit?” → “The speed limit is defined as the maximum rate at which a vehicle can legally travel on a given stretch of road.”
And not only that, but a specialized user interface is often preferable to voice even when the assistant passes the Turing test. There's a reason that people use apps to get food delivery rather than calling.
Your alarm example is frustrating. Of course, you could say "set alarm for tomorrow at 6:30 am" which would work, but then you're back in the realm not of natural language commands, but formal commands that exist in the uncanny valley, just similar enough to natural language to be irritating when they fail.
Edit: also, your "tomorrow at 6:30am" example is also open to interpretation if you're saying it two minutes after midnight. I'd really like it to recognize these sort of ambiguities and prompt me for clarification. "It's after midnight. By 'tomorrow at 6:30am', do you mean you want your alarm to go off about 6 hours from now or the next day?"
It doesn't even support 24 hour time - at least in English.
They use all kinds of languages with all kinds of accents and and ticks.
In my experience, and from the time I did a bit of NLP the situations is often along the line of. It works for mostly accent free simple English. Fails to get anywhere usable on most other languages and or accents. Sure that is to some degree because of missing training data. But for a consumer this doesn't change that for very many consumers this features work terrible bad.
Just out interest I tried out the youtube auto generated subtitle for a german video, but even at the parts where the text was super clear and well pronounced the result was hardly differentiable from randomly picking arbitrary words. It wasn't even that the algorithm choose similar sounding word. They where completely different words in many case. I think in a sentence of ~10 words in average 1 or 2 where correct. And that was at the parts where the text was unusually clear understandable. At other parts it wasn't even able to recognize that there where words...
Just imagine all the idioms they have to account for in these voice algorithms. I know in German they also handle "halb neun" correctly as 8:30.
Although it does have a problem with hours over 20 (“22 o’clock” becomes 20:02 for some reason, possibly because of my thick Italian accent).
Again, whether I would understand them or not, no one to my knowledge speaks like that in U. S. English. It is a great example to use to show the quirks of language. It is a bad example to use to show that Siri "doesn't even support 24 hour time - at least in English".
And in any case i said eighteen not eight a m.
Well, if I had some type of personal assistant who worked for me (as in, a real human), I would just call out "Hey, Sam, please order pizza for me" and continue with whatever I was doing.
The reason people don't make phone calls is they add a lot of additional friction. You have to dial the number, and wait for someone to pick up, and give them your address, and deal with frequently-questionable voice quality.
However, if you frequently order the same thing from a given restaurant, having a way to order "the usual" from that place via voice might be convenient. But I'm not sure it's much more convenient than being able to do the same thing with a button click or something.
For delivery, I agree. It is easier and faster to use a service, since my address and payment information are already saved.
For me, the ideal interface to engage (order something, create a note etc) when you know what you want is through speech. It just feels so effortless.
You can't get that with voice no matter how you do it. It's like if a restaurant had no menu and the wait staff just recited the menu in your face and asked you to pick. Even with real humans, voice isn't the best interface for presenting food choices.
That's exactly how food is presented to you at lots of high-end restaurants. Once you get above the fast-casual tier, menus never have pictures. If you know a lot about food and you trust the chef, voice is a perfect medium for explaining the choices. Pictures are just a discovery aid.
I understand it's a hard problem, but with current software capabilities and given Google's compute infrastructure I honestly think some of these things are well within the realm of what a team of several hundred Google engineers can do.
I'm not asking it to write Shakespeare, I'm asking it to crunch data with a sentence that could reasonably easily be parsed into a graph and be turned into a MapReduce query. I thought they were good at that stuff.
Context? I know it's hard, but I thought Google has been working on that. I'm very much adjusting my expectations to what I think Google should be able to accomplish in a decade. I expect some basic context capability now, at the very least these data-crunching type use cases.
There are two ways to look at this problem: (1) what is hard and what is easy for our current tech to do, and (2) what are the things that humans actually want to pawn off to assistants?
The problem is that those are two very different answers. I don't think I agree with the grandparent comment that some of those things should be easy, but I do think that comment contains good examples of the level of sophistication that would make an AI assistant more than just a curiosity for the couple times each day you prefer not to hit the button yourself (or if you listen to the radio while you shower).
The rest is easy, it could even be done completely on-device if you have a recent high end chipset. That's how powerful phones are nowadays.
Tbf, it is easier to just run your own server, and pre-program everything for yourself, a truly personalized experience.
Though Google could easily do it for the millions of people using Android. I really wish they allowed custom modules or something for their Assistant, the voice recognition is unmatched.
“Here are the results I found for the last week...”
“Show the same for the last 10 days”
Maybe the difficult part is whether or not you look for the 15 most recent photos containing people, or the 15 most recent photos of anything.
If I ask for recent wildfire news and I'm in a state that doesn't experience wildfires often, are you going to return 15 news articles about wildfires spread out over the 200 year history of the state? I almost certainly want 15 news articles about the current wildfires in some other parts of the country. Your algorithm doesn't really say what to do here.
If I ask for "recent relatively rare astronomical event" recent might mean hundreds of years or more. If I ask for "recent PC game releases" it might mean a month or the current year. If I ask for "recent public events in my town" it might mean over the last week.
In many cases, "there are no recent events" is a better answer than "here are the last 15 events."
The second query was hard for me to test but it seems like a reasonable one to make.
The third one is probably impossible due to Instagram's TOS (among other things.) Their TOS states: "You must not crawl, scrape, or otherwise cache any content from Instagram including but not limited to user profiles and photos." [0] After Cambridge Analytica I'm not surprised that this is the case. Even if Instagram allowed scraping of their data, this feels like a somewhat specialized (though currently very relevant) request. I tried reformulating this request as "What is the mask compliance rate in Shasta county?" which also didn't turn up results, but I'm not surprised by that either. I suspect that data doesn't exist anywhere, so it's hard to fault the assistant for not pulling it up.
[0] https://www.instagram.com/about/legal/terms/before-january-1....
I don't know how much of FAANG's budget goes towards improving voice assistants, but considering how much cash these companies have on hand and their operating budgets (the size of some smaller European countries), the progress in that area is just super disappointing.
The 3 most common use cases were refined a long time ago (Directions, Alarms and Play Music) and everything else ran into a hard wall.
- "Take me on the most scenic route to X." Can't you figure that out from social media tags? Simple first order solution: routes that have more photos with more likes = more scenic. Took me 1 minute to think of that. And 1000 engineers at Google couldn't implement that? These data crunching tasks are the kind of stuff computers are supposed to be good at.
- "Navigate to X but make sure you stay on paved roads." An actual problem if you are trying to use Google Maps in the back roads of California and don't have a 4WD vehicle. Google Maps loves taking you on 4WD dirt road detours. Don't you have satellite maps and street view? Can't you differentiate paved roads from unpaved ones? What the hell does your machine learning department do?
- "Stop at the last grocery store before the highway 120 junction." Yeah, it doesn't even begin to understand this type of query.
- "OK Google zoom out the map slightly." Nope. Sorry.
Paved/unpaved: that's one suck-ass assistant you've got there. I know the Garmin on the dash of my motorcycle will give that option. The Garmin RV-specific GPS has loads of other options, such as avoid any low overpasses that will rip the solar panels off my RV. Though it doesn't look like Apple is any better than Google, only because Apple Maps doesn't offer the option. Anyway, satellite and street view? Rand McNally had this information since dirt was first created, no need to get the satellites out.
The real frustration is when it fails on tasks I swear it has done before easily.
Like 'ok google' ends up googling ... the ok google command word for word sometimes. 'ok google set timer 10 minutes'
As soon as I have to babysit an assistant like that I'd rather just make sure the clock is on my home screen so I set a timer.
Raise your hand if you remember the early days of Xbox "recognition", as you listen to other people in your party:
"Xbox, record. XBOX, RECORD. X...BOX, RE...CORD."
It's the reason that "oh, we can't sell an Xbox without a Kinect" Kinect went into the garage after a few months.
It's really too bad that MS went with a windows only sdk for the 2.0, the 1.0 being external tech had a multiplatform sdk.
It could even be done on the phone itself, although battery life and heat would likely be a problem.
The fact that it doesn't work just shows that Google isn't that interested in their Assistant.
Which is rather strange, but maybe they're looking for better ways to monetize it?
Or maybe it's going the way of Reader once AR gains momentum or something... then again Google Lens is a really nice product and no one I know even heard of it.
I'm amazed at how well it can recognize any kind of writing, even my chicken scratch, and how it can look up any products/labels with decent results.
- you doing a manual search, and being tracked/targeted - them displaying ads as a result - them displaying contextual, higher paying ads
I've noticed that some things over the years, many things in fact, have been removed from Google Maps. I hypothesize that the entire reason is, these things reduce profit.
For example, you used to be able to pause Google Maps. You can't now, and therefore, if you have 'history' off, you have to stop, and manually re-start your destination by typing it in.
Well, history profits them.
And pausing is not the best, because then, you're not active 100% of the trip. Having you active lets them determine all sorts of things, like traffic patterns, where you go, who you're near, and so on.
There are lots of little things like this, which seem to be gone from earlier versions of Maps. Again, I presume, all to make more $.
Which is logical, and fine, but it gets a bit tiresome and sad at times.
While an artificial general intelligence would be capable of fluent speech, a perfectly good speech system does not necessarily need to be an AGI.
To give a more concrete example, here's a UPenn demo (https://cogcomp.seas.upenn.edu/page/demo_view/ShallowParse) for your instagram query:
> NP What percentage PP of NP people can NP you VP detect to be wearing NP masks PP on NP recent Instagram photos VP tagged PP at NP a location PP within NP a 10 mile radius PP of NP Mt. Shasta ?
Part of the reason we don't progress beyond that is that speech recognition like in the OP article is quite bad: 95 percent accuracy is considered "good." But it means we expect 1-2 words of your query to be misrecognized, so even if it did parse the query as you proposed, it would probably be answering the wrong question!
But, on the other hand, I have dictated these paragraphs, (pretty much just adding punctuation and minor edits at the end). I think the most useful feature I've found related to voice is dictation (speech to text). It is almost perfect, at least in English (and I am not a native speaker).
The thing we're discovering is how to marry NLP (which works pretty well) with structured databases and automated tools (which work really really well).
A pure AI play probably won't get this done. But if your AI starts knowing how to use information tools then I think we'll see a lot of near term progress.
https://arxiv.org/abs/2010.05243v1
more
https://scholar.google.com/scholar?as_ylo=2020&q=text-to-sql...
That's it.
People want "[wake word], [action] [modifiers]".
Companies want "[trademark product] [trademark product] [modifiers]" (eg the parents example "Hey Google (RTM), Live Transcribe (RTM) 'words to transcribe'").
I've only used Alexa (and only for fun) but the replacement of verbs with companies/products, and of nouns with proper nouns, is really annoying to me and gets in the way IMO.
"What's the weather in Yosemite" and "Yosemite weather" can get you two drastically different results. As of right now (10/26/2020 10:35 AM) the two results I get are 43F and 32F.
I made an "and finally ..." to add funny news to the end of the daily briefing and was super impressed how easy it was.
We're humans - we can deal with ambiguity. Systems should trust that we'll respect them more if they tell us they're unsure, rather than either jumping to the wrong conclusion or simply being unwilling to guess!
I have an admin. I can text, call, talk, email her and say “let’s get bob, mary and finance on the phone at 4” and she’ll move stuff around to make sure that happens. I can also delegate complex administrative tasks that need to be done in my name.
When I’m at the dentist, I usually do a “hey Siri” to book it as it saves time. At home, “hey Siri turn on the lights” is helpful. It’s magic, but not the same.
The tech companies try to frame this stuff “bigger”, but in doing so they create an unreasonable expectation. Google home, Siri, Alexa are amazing, but we bitch and moan about them as a result.
I don't have a human assistant, but if I did, the biggest things I would expect from them isn't about scheduling a meeting or scheduling a barber appointment. I can handle that stuff myself.
What I would REALLY expect of a human assistant:
"Can you call up my health insurance and fight this stupid bill of $400 for COVID testing that should have been $0. Please escalate if necessary. Thanks."
"Can you figure out how to fight this red light ticket that I got due to a malfunctioning red light camera? Thanks."
"Can you fight this parking ticket? I had a valid permit to park there, and here's documentation of that. Thanks."
"Can you register me to vote? Here's my ID. Thanks. Make sure they don't sign me up for spam."
"Can you call up Comcast and fight this bill increase? Threaten to switch to another provider if necessary, I heard that works."
"Can you call up this company that posted my personal information and ask them to remove it? If they refuse, threaten legal action."
"Can you dispute this electric bill for me? My heating shouldn't have been $400 a month. Something must be wrong with my meter or someone is leeching power from my line."
"Can you call up the 10 different grocery stores in the area and figure out which one has X in stock?"
In all honesty I do wish the Google automated assistant could do all of the above. Sorry Pichai, I don't care for scheduling haircuts automatically. I want your assistant to use its hundreds of thousands of hours of human conversations and use machine learning to craft and engineer responses to humans to know EXACTLY when to escalate, EXACTLY when to ask for a manager, HOW to threaten legal action, and basically fight tooth-and-nail with language to get me what I want against the companies and institutions I need to fight with on the phone. The job of an "assistant" should be to get sh*t done and get me what I want. Use machine learning and lots of data to master the art of negotiation with customer service reps.
All the cases you've indicated as "so simple" are only so simple as template matching.
When generalized, they are hard problems.
Sure, you could have humans sketch out a few thousand templates enabling "Ok Google" to support such simple things like "What is the air quality at <location> <Time>" which looks up from an API. But that doesn't generalize "for free" outside of your templates.
Google seems to want to avoid handcrafted templates and go straight for the generalized solution.
Ok, so you need the assistant to:
* Already have a trained dataset of people wearing masks
* Fetch ALL instagram pictures it can find
* Not only detect if there's a mask in the picture, but count them
* Fetch the location of Mount Shasta
* Calculate a 10 mile radius around it
* Apply that calculation to the people counted in the first step
* Calculate a percentage of people wearing masks which needs:
* Count people not wearing masks.
Those steps are sub-optimal. So you need not only understand that, but to run some sort of 'query planner' in order to get the results you are looking for.
You are thinking of Jarvis from Iron Man. Forget assistants. Allow someone to construct a query like this in a couple of minutes and you have a great product you can sell. Existing ones require quite a bit of domain knowledge and setup. Not even things like Wolfram Alpha would be able to parse this query.
This reminds me of: https://xkcd.com/1425/
Sure, but we're talking about Google here, not Wolfram. The masters at query optimization and MapReduce. I would have expected them to be able to parse this query, or at least fetch the AQI near Mt. Shasta but can't even do that.
"There is no 'Grocery' list. Would you like to make one?"
An example of how Apple's ecosystem breaks down if you want to use something outside it. Saying no should offer to add a reminder as requested, but it doesn't.
It's the only thing I use my Alexa for.
Siri play song “song-name” I’m sorry there was an error with Apple Misic Siri play song “exact-same-song-name” [song plays]
I’ve figured out what queries almost never seem to fail and use those almost exclusively. I don’t get creative.
> (1st attempt) "Computer! coffee please"
> (computer dumps coffee on Kirk) "argh!"
> (2nd attempt) "Computer! coffee in a mug please"
> (3rd attempt) "Computer! hot coffee in a mug please"
....
> (25th attempt) "Computer! 10cl of coffee at 50C with 3cl of fresh milk at 6C in a bottoms down ceramic mug of 15cl"Google Home can do some clever things however, it also does not have the ability to do some very basic stuff. As a user, how do you know what Google Home can do and cannot do?
It is just trial and error. And if Google Home introduces a new feature to be able to complete new types of queries. What now? How does a user know that last month it wasn't able to do something and this month it is able to do that thing.
And lastly, the interface of voice is very clunky. It has no concept of temporal memory like
Me: "Ok Google, navigate to the nearest safeway" Google: "navigating you to the nearest gas station"
The natural thing to say is "no, I meant safeway nott gas station" however, I now have to say "Ok Google, navigate to the nearest safeway".
This is analogous to if keyboard had no backspace and you have to retype everything everytime you have a typo. Well that's the state of speech technology right now
Amazon's workaround for this problem is to have it tell you when new "options" are available for a command (ie: if you set an alarm it would confirm setting the alarm then tell you it can wake you up to the sound of birds then it provides an example command) and Amazon sends out a "what's new with Alexa" email every so often that's 90% example commands.
In some ways a voice UI has bigger problems to deal with than PC GUIs or iOS. Those UIs were replacing pre-existing UIs (eg blackberry, dos, unix, norton) and they could target whatever tasks a smartphone/PC needed to do. For voice UIs, it's a cold start. It's not even obvious what an audio only computer should do. Our mental model for a "virtual assistant" is a person-2-person exchange, and computers still aren't great at communicating like people.
FWIW, I think slipping into existing niches is the way to go. That's where a useful voice ui will be discovered. Car stuff, accessibility software, living room controls.. At least these have clear goals. Voice operating spotify, netflix or just an iphone is something people actually need and will use if its useful.
1. "Who is the president of the United States?" 2. "What is his wife's name?"
And it will resolve the deictic pronoun.
I haven't tried this feature out extensively, but it has worked for a few years now.
Now for long-form typing, I'd love to use dictation and sometimes do for taking down short thoughts I e-mail myself from my iPhone.
But the problem is not just that it still makes tons of mistakes. (Probably a quarter of my notes-to-self involve errors so big it's even impossible for me to later figure out what I even meant by trying to sound it out phonetically.)
The problem is that I can't correct those mistakes using voice. There's no way to say "pause, correct affect to effect" or anything like that.
Even more maddeningly, the words keep changing in real time. Sometimes I'll utter a sentence it gets right, then it "re-analyzes" it and completely messes half of it up.
I just wish there were a kind of dictation where I could say a phrase, pause, see if it's right (and it wouldn't change after), say the wrong part with a kind of emphasis that lets the system know I'm issuing a correction, the system would look for the next most probable alternative, repeat as desired. Then I could actually dictate successfully.
This UX where the words are always changing back and forth according to updated statistical probabilities, even as long as 15 seconds after I said them, and where there's no ability to go back and correct them with voice... it's just so so dumb.
The problem isn't voice recognition anymore. It's voice correction.
Classic scene: https://youtu.be/xaVgRj2e5_s?t=171
Star Trek isn't my favorite sci-fi show/movie, but they really nailed some futurism and in this case, lack therof.
imagine if your keyboard had about a 2nd latency and every couple words got messed up in some way. Not only that, but those same words that got messed up are probably getting be messed up again when you try to go back and fix it with the same broken keyboard. you wouldn't say that typing is nearly solved, you'd say that typing absolutely sucks and keyboards just don't work.
I firmly believe that speech is going to be the main interface-actually want to scratch that last sentence, but correcting it is can the pain with voice. I firmly believe that speech is going to be a game changer of an interface, especially for coding, eventually. But until it stops sucking, it's only can be used where where there is no other choice.
(By the way, this is me, a native English speaker with a standard salmon Cisco accent-that's a standard San Francisco accent- dictating it on a many hundred dollar microphone into a many hundred dollar speech engine, speaking relatively slowly and enunciating the hell out of out of things.when I first started with that, it was even worse. and yet somehow a lot of people treat speech recognition as if it is in some way solved, or the error rate is better than human.or loader ship-that's supposedd to be what a load of ship-whatever you see what I mean)
-----
(edit: After the fact, I decided to calculate the word error rate dictating thhis. About 7%. If you ignore the non-speech-related bugs, it's about 4%, which is supposedly "superhuman." take that as you will. my take is that "human level" is not the same as a human trying to work as fast as they can, and maybe not paying particularly close attention. And that 4% is ridiculously far from 0% in terms of usability)
It's super unnatural to correct things by voice and until we re-imagine what it looks like if we HAD to type via voice, it's gonna be painful.
(Speaking as someone who types via voice as well). Have used Dragon and now currently Talon
At the start of my morning commute, I would say, "Ok Google, navigate to work".
Often, this would fail because I was in the network limbo area outside my house, where my phone struggles to transition from home WiFi to data.
Worst of all: The failure would be horribly slow. I would have to drive for another 30s before my phone realized, yes, we are really out of WiFi range now. And the voice command wouldn't be auto-retried. I would have to tell my phone again. It didn't remember.
I added one-touch "Home" and "Work" Google Maps widgets to my home screen and never looked back.
As an engineer, I realize why this is a tricky problem. As a consumer, I want it to "just work".
Nowadays it might work via a location-based reminder, but I can't trust that to work within the 15 second window I have to get off the train.
You can speak to these assistants, but the language is still restricted. They show little to no common sense. It's a lot of party tricks bundled together.
You can't interrupt them and it's hard to correct them.
On the speech recognition side, an issue I've found (although is a rather niche one) is triggered because I'm bilingual (I'm fluent in spanish and english).
Speech recognition only works well on a single language.
I have Alexa set to speak english, for example when I'm searching for a song with a spanish title, I have to try to fudge the name into a fake english pronunciation for it to produce the right phonemes that will match the song title, rather than just say the name properly.
Also, if it misses the match, there's no easy way to stop it and say "No, not that one", and be presented with a list of similar matches.
Just my 2c, I'm sure other people have uses for it. The most interesting one (to me) has popped up a few times on HN, which is voice-based programming. I would love to see that mature and become more widespread, there are a few things that are annoying enough to do that if I had a voice shortcut or eye tracking it would be pretty cool.
Speech is competing against every other tech item trying to be convenient, from laptops to phones to watches... the only space where I'd want it is something like baking or cooking when I can't interface with a computer.
(edit) and what about the response back? A visual interface I can confirm at a glance whether my input was accurate, a voice one I have to listen to the whole thing. I even don't like maps directions in the car as half the time I know the next direction already and don't want the interruption to music or whatever I'm listening to.
This is the main reason I rarely use voice controls. Even if voice is faster than typing for most cases, typing never fails catastrophically like voice does.
I could use voice only when I'm confident it will work, but then there's the mental load of making that prediction. It's easier to just always go with typing.
Siri created a reminder (good) called "Take my meds at lunch" - no time (bad), just a simple reminder.
I saw this post, looked at the time and wondered why Siri didn't remind me, given it was 1:00 - now I know to be more specific but in reality I'll just stop trying like most people.
As far as I can tell, the assistant APIs seem to be like plugins? On Android for example, it appears custom assistants still run through Google assistant.
I want to be able to say "TriggerWord, do X and Y" and the OS activates my app, passes the voice sample and I take care of all the language processing from there. Which doesn't seem possible...
I did not want the home asst listening all the time - as it's highly likely to falsely trigger just listening to radio or tv it... so instead I use python+opencv to detect if the assistant (an old rooted samsung note) is being directly looked at - then it wakes up to listen for the trigger word and the command. Of course, I can also manually trigger it via any device in the house.
Imo, Assistant is limited because of privacy, see how it asks you to opt in for a more personalized experience and it still won't unlock your phone for you.
"Turn on the lights" isn't exciting. "Make dinner and do the laundry" would be exciting, but that's at least 25 years and some major advances in robotics away.
IoT is very crude and contrived compared to what would be possible with an active technology that could do useful physical things of all kinds.
Where it could really shine is in more rare commands. "Make me some Pad Thai". Unless you're a huge fan, you wont want a button for this. And it's faster to say than type into your phone.
I also never talk to my phone.
1. When these systems go offline, they’re nearly useless. One WiFi glitch and the system gives me the audio equivalent of a blank stare. This directly reduces the chance I’ll casually use voice and instead prefer my more-reliable phone.
2. Most systems haven’t figured out reasonable responses and are basically chatty and full of sounds. Make it Unix-like (silence on success = golden)!!! If I say “turn on the light” and you turn it on, I CAN SEE THE LIGHT ON so I don’t also need to hear a loud chime and some voice confirmation like “sure, no problem”! The fact that I can silently do things quickly in other ways (e.g. phone) is another strike against voice. Yet this is something they could easily fix.
3. Voice systems are not 100% perfect at comprehension yet they tend to babble out long responses. This puts me in the situation of trying to shut them up for long enough to listen for my intended query. Maybe they can improve this by erring on the side of fewer words in replies, with more pauses? Not sure.
Otherwise, I have absolutely no desire to 'talk' to my computer and have it understand me, unless that tech comes packaged with an empathic AI module so I can tell it off and repay it a small percentage of the emotional pain computers have inflicted upon me over the decades.
The critical benefit would be to be able to do this whilst at the same time 'controlling' the focus with a mouse and keyboard at the same time.
It's honestly quite shocking how sparse the research and implementations for everything is once you go beyond a single sentence/command that you shout at your personal assistants.
Frames mean that words and sentences have a context, and you can't understand conversations unless you understand the context.
This starts from simple and obvious distinctions. E.g. - as a silly example - "make dinner" usually means "Prepare and cook an evening meal". But if you have a project called "dinner" it might mean "build and compile 'dinner'" An AGI should be able to understand the difference, and ask for clarification if it doesn't.
Eventually you end up with subtextual and implied communication - e.g. "I'm fine" can mean two completely opposite things depending on tone of voice and the contents of minutes-to-years of previous conversations.
All of this is many orders of magnitude harder to handle than "Bedroom lights off."
But there's a huge swathe of use cases where it does make sense, and I think we should be focusing on those - situations when you can't use your hands. Voice assistant technology doesn't even need to be great for this, just good enough that you can look up unit conversions with your hands covered in bread dough, or navigate while driving or whatever.
The author notes that they only get about 50% of regular speed with this approach, and that may be a significant part of the challenge---speech can encode complex concepts into a few words (especially given context), but the actual baud rate isn't particularly impressive. Keyboard interface, where possible, seems to still win out.
I can open Microsoft Word faster than I can say "open microsoft word". It would have to be smart enough to short circuit the entire process of doing something useful with Word.
It might pass in a coffee shop, it probably won't in an airplane, it will almost certainly never pass in a library. People will try it regardless of the appropriateness of the location, because some people aren't aware or are assholes, and as a result the entire technology will get a bad rap. See google glass as an example.
I agree if it was already the social norm it wouldn't be a problem (same with Google glass), but it turns out that the technology being ready isn't always enough to make the social norms change.
But beyond that, I don't really use these "digital assistant" because you can't teach them when they're wrong, so after 2-3 failed attempt at a request, you lose interest because you can't trust the system.
"Start the dust extractor Hal" = start the workshop dust extractor, but I no longer use this since upgrading to remote start/power line detection on my dust extractor.
"Hey producer, <various commands to control video production software and PTZ cameras>" = start recording/cut cameras/re read the prompt/extreme close-up on keyboard/pan the camera to me/etc for video production in a one-man multi-camera recording setup.
Everything else, when it comes to NLP, chat bots and voice recognition is "too much trouble" and I find hitting a dedicated button or punching in to a few numbers on a physical keypad to be easier and more reliable than any voice interface.
The interaction is very fluid, low-latency, and accurate, and the system doesn't force itself on you: there are still plenty of non-speech-based user interfaces to be found all over a starship.
In general, I find dictation useful on my phone to take notes to myself. And maybe to send a message in a chat situation. But it just doesn't work for me for anything slightly more formal / long.
At this point I'm dictating even my (work) design docs and (personal) blog posts. I do think I'm a bit less fluent when using speech to text, because it's harder for me to jump around and make slight word choice improvements, but not by much?
(For example, I dictated this reply)
I think a lot of people are counting on speech to bring us into a sort of Star Trek future.
The real game changer for input is along the lines of what the neural lace is supposed to be. Cognitive input. Silent, fast, efficient. In many cases once the technology is mature people won't even have to internally verbalize commands. Just look at a light and desire it to be dimmer.. it dims. "typing" at the speed of internalized though will also be amazing.
Every time I hear someone (including myself) tripping over "OK Google" I cringe.
Loud is relative, privacy is an aimless indictment that's orthogonal, and for brevity?
Speech is a fantastic tool for communication, which includes input and output. It's part of why most large animals, for which quick communication is imperative, use it. It's imprecise, which is the problem that machines are not good at dealing with. It's was a good direction, when we used to have machines that were initially trained with speech for better accuracy, but now passive listening of devices isn't even used for that!
> The real game changer for input is along the lines of what the neural lace is supposed to be. Cognitive input. Silent, fast, efficient.
Silent sure. The human mind is rather random, highly variable between individuals and ages. I would not call it fast or efficient. Then again, speech to text is contextual cognitive input. Without drugs or intentional damage (minor) to the brain, I don't expect neural implants to be very effective, even in the next 100 years.
https://arxiv.org/abs/2010.12096
It improves much faster than you might think.
1. Voice based operations in factories, construction workers who want to have both the hands free but want to navigate via a device
2. Use cases while Driving. E.g. A Driver who is delivering goods.
3. Call centre - Analytics of audio calls etc. Can have many use cases
4. Voice assistant like Alexa, Siri. Mind that Alexa, Siri have vision to do more than just Music.
5. Any use case where visual interface is either not there or visual is not an easy option for user.
Speech tech is challenging when you have to deal with noise or want to do Speech to text on lw profile devices (on the edge).
They can't understand speech properly in a quiet environment, and you want them to get your commands on a factory floor or construction site? :)
There is just no way I will let google device that listens my conversations let at my home (let me be faster - no, my phone doesnt listen, and yes I am sure, its maker wouldnt recognize it any more).
Or amazon. Or facebook. Sure speech technology could be usefull but not from those companies.
On the other side, there is just no usable way found by anyone else, how to monetize the technology (i.e. except for spying on people).
I think that this is THE problem of speech recognition.
The Pentium processor's power over the 486 was supposed to be the missing link to 'working' text to speech.
AFAIK it was at the level of super funky chording keyboards; with a lot of investment it could be a cool and impressive input method, but the personal investment side was too much for most users. There was an "assistive technology" user base that i dont think they ever realized they had.
When people demo Speech Technology they have the computer doing general AI in response.
Things that are impossible with text or any computer input or anything that can happen for decades.
Reality is, no one wants Speech Technology, we want a computer that can order a pizza with the simple typed words, get me a pizza.
Heck, you pull that off people would even do the way more annoying 'talking' to the computer to access that tech.
What people really want is text to speech. But that's to hard to do the AI scam on.
Maybe if there were a better way to pronounce "open paren" and "make-vector" we could just yell scheme at the nearest computer.
IMO you're never going to get anywhere trying to make shells speak natural language. Even if you crammed a person into the box there's just way too much you could ask to narrow it down without something formal. At the end of the day you'll have to have something that looks like sh or any other formal language.
It's kinda funny watching someone without a background in CS repeatedly ask google/alexa/siri a question in different ways that you know isn't going to elicit any kind of useful response.
Dictation is widely used in medical transcription.
Dictation is a killer way to write a first draft quickly, transcribe rough written notes after a meeting, etc. Also, about half of my emails are dictated, and I know I'm not the only one. It takes some time to get used to, but once you're there (like touch typing!), you can't go back.
..etc..
Your computer has yet to process meaning.
That is where we are stuck right now.
"Hey Siri, remind me in 2 hours to do X"
"Hey Siri, remind me when I get home to do Y"
"Hey Siri, remind me next time I go to Costco to buy Z"
It's pretty quick to set up reminders by voice, especially location-based reminders. Almost anything else I'd rather do via a UI.
I was surprised by google meet transcription system, which to me, it's the most accurate I've used so far. Same with the google docs dictation.
That was fun for about 10 minutes. It's been disabled ever since.
like imagine saying " Hey open a reddit/meme subreddit on the side " while you are doing your watching youtube. or imagine saying "hey Wikipedia that *"
Google's voice capture often astonishes me with its accuracy, and nearly as often makes me laugh at how bad it can get things.
Hands free.
"send message to john"
"how far is it to"
"get directions to"
"play podcast"
"play audiobook"
I can imagine people using it exponentially more when you no longer get weird looks when you say “Ok google” or whatever in a supermarket.
Modular so I can switch out the speech recognition and the NLP and fulfillment
it may mean that if you aren't comfortable sharing rich context about your life with a cloud platform, you'll get left behind by technology
My wife always talks to google. It seems backwards to me. Next generation should skip speaking and wireless read thoughts. Language/speaking is a bottleneck.
Do you suppose it's possible that speed isn't really what counts? In the era of typewriters, authors like C. S. Lewis opined that typing obscured their thinking because it was too fast. It didn't let them savor the words effectively. Maybe what we really need is to slow down?
In some cases, we should slow down. But in general, I disagree. Especially because of the way in which we use technology now. Our smart phones are becoming a bit of an external brain to us. It holds contacts, conversations, searches, notes, musings, etc. We use it to recall a fact or answer a question without really even thinking about it. Maybe what we really need is to slow don’t think it is only a problem now because smart phones and the like aren’t really designed to improve our lives, but as ad delivery platforms to manipulate us into buying products and services we probably don’t need. I would love better tech that helped offload mental tasks my human brain isn’t great at, but computer brains are, and let me focus on things my human brain is good at as well as things I just would rather think about.