[0] http://www.acquired.fm/episodes/2015/12/14/episode-5-siri
[1] http://9to5mac.com/2014/06/30/why-is-apple-hiring-nuance-eng...
[0] http://www.acquired.fm/episodes/2015/12/14/episode-5-siri
[1] http://9to5mac.com/2014/06/30/why-is-apple-hiring-nuance-eng...
This is super hard. Not only do you have hidden functionality (what a great insight!) but they are trying to do something better than the current state of the art (which is human-to-human). What do I mean? Just listen to yourself when you get something over the phone. You don't call the restaurant and say "give me a reservation for four at 8." You typically use an interactive process of query and response ("can you fit four in at 8pm?" "Yes, today." "No, not five, four". Then they repeat it back to you just in case.
We get annoyed at how crappy the voice recognition systems are, and they are crappy, but human voice recognition reminds me of a TCP session negotiation, or maybe the baud rate negotiation of a bell-compatible Hayes-style modem. It's hardly a one-shot process, yet we want our machines to be.
In the next ten years there will be a large battle - which will play out on HN daily - over the AI to AI standards; it'll remind older developers of the XML / JSON days.
Once you know what the user wants it's "easy" (well, somewhat easy anyway :-). It's the iterative process of understanding the user in an unconstrained opaque ad hoc UI that is big-H Hard.
The problem with these spoken interfaces is not just that we don't know what we don't know, we don't even know what we want to accomplish!
Open Siri (hands free Hey Siri works for me 5% of the time at best)
"timer 20 minutes"
"Okay, 20 minutes and counting"
A little later, oops that'll need 5 minutes longer than I thought
"add 5 minutes to the timer"
"are you sure you want to change the timer?"
"yes, change it"
"Okay, 5 minutes and counting" Great...
Other than simple keyword commands like that its voice recognition doesn't seem to be good enough for things like transcribing messages anyway.
It's sometimes quicker than spotlight for a web search because of that weird way spotlight hides the search web button for a few seconds when it can't find anything to show.
Plus, even if it did work, it's just a really inexact and inefficient way to do anything. You always end up trying to see through the abstraction layer (natural language in this case) to the much more well-defined hidden system underneath. It's like reverse engineering an expert system, as a UI paradigm.
Programming languages made to resemble natural language face a similar poison pill, AppleScript being a prime example.
I totally get the appeal of trying to make software understand people better instead of forcing people to understand software, but in practice it just ends up being one more abstraction layer users have to struggle to get past in order to unlock the functionality of the underlying system.
The problem now is Android randomly kills the "OK Google" background listening process, or it fails because it can't handle a handover between wifi and cell data, or it "can't open microphone" even though it just heard me say OK Google, or any one of many other Android problems rears its ugly head.
The reliability of voice transcription now is far better than the general reliability of Android on my Nexus 5X, and the Android team ought to feel pretty bad about that.
It's also incredibly easy to fall off the blessed path where you can interact solely with voice - when setting a reminder, for example, if it doesn't hear 'Yes' when it's expecting to it'll just sit there and you need to touch the screen to continue - defeating the whole point of voice interaction.
Google's voice search has incredible voice recognition and text to speech and can tap into an amazing amount of information through the knowledge graph, but they don't seem to be capable of fixing the basic bugs and UX issues preventing all that technology from actually being usable.
I had this happen with Google Maps the other day. I was driving into San Francisco with the turn-by-turn directions when it said something like, "There is a faster route available in two miles. It will save 5 minutes. Tap 'Accept' to take this route."
So I had to fiddle with my phone in a hurry, in traffic (after all, traffic was why there was a faster route coming up), and hope I don't tap the wrong button or crash into anyone while looking for it.
Why couldn't I just have said "Yes! Please and thank you. Of course I want the faster route, why wouldn't I? Oh, sorry, I mean 'Accept'!"
The funny part was that the "faster route" was the way I usually go into that part of town anyway, but Maps had been sending me a different way because of congestion on the usually-faster route. Why did it even ask me to "accept" the faster route instead of just redirecting me the way it does automatically on most occasions?
Like the time I was heading south on 101 through Morgan Hill and Gilroy and Maps had me get off the freeway and take a side street for a couple of miles because traffic was stopped on the freeway due to a car fire. It didn't ask me to tap Accept then, it just gave me some very practical directions on the fly. That's how it should work.
We're not suffering from too stupid AI, we're suffering from stupid app design and laziness when it comes to adding more variations to conversation. Just put a team of script writers to imagine thousands of replies, it is not that hard, people have been doing this kind of rudimentary chat bot scripts for decades now.
I'd like to be able to install an App and suddenly, new commands are available to me.
Car navigation systems would be a good place to start.
I can also use it dictate an email or search the web, and occasionally it fails (maybe 10% of the time?), but it almost never gets my wife's name.
What's your wife's name?
In a somewhat related note, I've noticed people search on some websites fail with Vietnamese names because surnames are interpreted as prepositions, and thus dropped as a stop word.
The rate of adoption will skyrocket when we reach a certain threshold, people will see other people speaking to their phones and they will learn to form useful voice commands as well. It has a big social factor.
I don't understand why they didn't at least add a few commands for YouTube. At least "play Band", "play Song by Band", "play Music-Genre", "play something like X", "I feel sad/happy/lonely/bored play some music".
If you want to understand the extent to which Siri is useful, you just have to look at the integration it has with the main apps. It's just Alarms, Calendar, Phone, SMS, Music, Weather and similar. It's not really groundbreaking AI, just a voiced command list to the main apps.
I am eagerly waiting for the moment when we could have more in-depth chats with our bots, but the current crop is dragging its feet when it comes even to simple commands sent to apps. Maybe they should ask the app creators to add hooks for voice commands, to make Siri more useful.
For example, if I can disgress a little, Google Now can't start playing the first video in the YouTube app on Android. If I am already in the YT app and say "OK Google, play the first track", it goes to web search instead and searches for "the first track" on the web! The same company made the chatbot, the app and the OS and they don't work well enough together. At least let users add new functions to the bot, if they can't be bothered to do it themselves.
Because YouTube ain't a music player? [1] And besides they have their own music app.
[1] Seriously, even though some people use it as such, this is far from most. According to YouTube itself, the average user only spends 1 hour per month listening to music on YouTube -- consumption is mostly non-music videos.
Except when it doesn't. I had number of times when siri would fail to set reminders. I do not have a need to set alarms that often, but people seem to have problems with it too: https://discussions.apple.com/thread/7280346?start=0&tstart=...
It's not just about inferring intent. Most discourse systems have little or no model of the consequences of their actions. Systems that can act on the user's behalf need to know if they're potentially making a big mistake or a little one, as guidance of when to ask for clarification. This requires some degree of "common sense". AI systems still suck at that.
Systems that actually do something need to be much better at this than ones that just answer questions. There needs to be some cost/benefit analysis of whether it's necessary to repeat back info and ask for confirmation. Without that, the system will either ask for confirmation too much, or screw up too much. The system needs to have a model of how clear the user is in their intent.
Pizza ordering is relatively low-risk. Travel reservations are a higher-stakes item. Medical is a long way off.
As I wrote last time, I recommend watching "The Devil Wears Prada", the scenes of Andy's first day as Miranda's personal assistant. That's the level of performance you want.
I need 10 or 15 skirts from Calvin Klein.
-What kind of skirts do you… -Please bore someone else with your questions.
And make sure we have Pier 59 at 8:00 a.m. Tomorrow.
Remind Jocelyn I need to see a few of those satchels that Marc is doing in the pony.
And then tell Simone I'll take Jackie if Maggie isn't available.
-Did Demarchelier confirm? -D-Did D-Demarchel…
Demarchelier. Did he… Get him on the phone.
Uh, o… okay.
-And, Emily? -Yes?
That's all.
This is why I think that eventually people will bring this technology up to a useful level it's only going to be the kinds of companies that can make huge time and money investments and have access to huge datasets. Not startup territory really.
So the obvious solution is either to have duplicate contact entries or only associate with people who sound like those you already know, so that Siri has to go to the "do you mean Person A or Person B?" state :)