I ask 100 information questions to four digital assistants
vlad.d2dx.com
vlad.d2dx.com
We are moving past expecting the most word matches into expecting the server to understand what we really want. Voice search is more like the "I'm feeling lucky" button, because it takes longer; you only have time for one answer and the first answer has to be right. It comes without the expectation that you're lucky if the answer happens to be right, now we need the first result to be the rightest result there is.
So the glass is half empty. I personally prefer to see it half-full, but it's also true and the critique is more valuable and interesting than optimism.
Perhaps Amazon, Apple, Google and Microsoft all initially thought that the revolutionary part was being able to speak a query and have the query match what you said, and that the search part was already good enough.
List decoding was invented by 1955. This is a set of hard problems, but very well studied ones.
"Set an appointment to play Lost Boys tomorrow." "What time?"
Absolutely. In fact sometimes I'm searching for something, whether in a search engine, at an ecommerce site, or whatever and I'm not getting what I want immediately. If I'm on a phone or tablet, I'll often grab a nearby laptop because it's just faster and easier to do a lot of typing and clicking on. (Less true with more recent tablets but my basic point is that there's sometimes a lot of fast iteration when I'm trying to find the answer to something non-obvious.)
I think @seiferteric's point was that the expectation may not be there because it is a voice search. That expectation is there because that's how it was marketed.
If the marketing for these things was: "Ask a question, and get search results by voice" I don't think I'd have the expectation that it find and deliver the correct answer to me.
But the marketing for all of these devices is: "It's a personal assistant! Ask it a question and you'll get an answer!"
I'm personally not convinced that the high expectations are because it's a voice interaction, but rather that the technology simply can't live up to the marketing pitch.
"Claim: Comedian Bill Murray is running for president and proclaimed religion to be "the worst enemy of mankind." Claimed by: Internet Fact check by Snopes.com: FALSE"
It would seem they are reasonably close but this is more of a product integration failure than a recognition failure.
Bigger picture though, Bill Murray is famous making it easier to answer questions like this. In general, does the wording of the author's question truly imply he's searching for a fact check, and do you expect Google to know that even if there are no articles that match the wording of the question? The snopes articles does contain the terms "did", and "Bill Murray", and "run for president", so we don't have any evidence that Google understands the question, we just have some content that matches the query.
The issue I see is that the computational question of search has long been trying to measure relevance by matching the query against the corpus. This Bill Murray question is an example of how that can break down. I might actually want the fake articles... and I might not. There's no way for the search engine to know without making an inference, and the expectation that mass market search engines make inferences seems pretty new to me - and I don't expect that when I do text searching. I guess I just expect voice search to push the need for question understanding and inference making even faster than text search has.
My first Hacker News inclusion. I feel like there should be some rite of passage. Well, other than the sudden and unanticipated login attempts.
My testing device for Google was the Google Home speaker, which appears to have a different tolerance for reading search results. I've had it rattle off several sentences from web pages for other keywords in the list (see, for example, the boiling point of water), but for the Bill Murray question there seems to be some kind of limiter. I just re-checked, using the exact phrasing I had before, and it still says that it doesn't know, but it's learning all the time.
I'm guessing there is some kind of a relevance check for the speaker version compared to the phone version. The phone is probably happier to return any result (a la Siri), whereas the speaker appears to be making some attempt to understand what I'm asking for before reading search results.
This particular question appears to trigger the speaker not to read the search results. We can only speculate as to why: does it not find it relevant enough? Is there a reserved path on "Did xyz" questions when sent to Google Home? Am I unknowingly in the A/B testing group that doesn't get the answer? There's few ways of knowing black-box without massive data testing, but it is curious.
> I'm guessing there is some kind of a relevance check for the speaker version compared to the phone version.
I would bet on that & expect it too... I'm sure all these voice search products are experimenting with how voice search needs to be tuned differently than text search.
There's a really interesting discussion to be had about how UI decisions can make that process smoother — I really liked https://bigmedium.com/speaking/design-in-the-era-of-the-algo... as a call for how you can make the failure modes of the system more graceful. I think a lot of the success in the next decade or so is going to come from the places which figure out good answers for not making a system which seems to promise more than it can deliver.
Very nice article, I only skimmed so far, but I think I agree with all of it. It has a definitely pro-consumer bent that I wish would come true, but seems like trends are in the other direction. I suspect there's too much money in search and improving query understanding at scale for companies that get there to be as transparent and open and sharing as this author is asking.
It's not the voice UI that does this, it's the marketing.
If they sold it as "speak your google search terms", it'd work a lot better. It'd also be a lot less sexy, but that's still mostly what it is IMHO. Not to say that it isn't impressive stuff, it is! But it's highly oversold, still.
Currently there is one way for a computer to interact with humans via voice: direct and unquestioning answers.
I think you can. Alexa can read you the titles and you can ask it more information about a specific title. It’s like asking the waiter which desserts they have then interrupting him because you don’t know what a pannacotta is.
If I were to ask a question that had a long list of potential responses, I'd expect them to ask me to clarify or narrow down what I'm looking for or at least explicitly ask me if I really wanted them to read the whole list.
That said, for those of us who don't have personal admins, there's a lot of opportunity for digital services that fall between purely self-service travel booking as it exists today and and having an assistant.
https://www.microsoft.com/en-us/design/inclusive
As an example, a coworker mentioned that his use of Alexa went from casual to heavy when they had a child and the ability to do things while carrying a baby suddenly became really important. I suspect there are more situations like that than we might think at first.
1. I know, 90s me is still getting used to saying that too
Also the query "did Bill Murray run for president" returns a Snopes article debunking the myth as the first result. This should totally be something a Siri thing could parse and tell you about.
https://www.google.com/search?q=did+bill+murray+run+for+pres...
It's a bit like pressing I'm Feeling Lucky for your result. I'd hope that it was more optimized for always good results rather than often great, but occasionally lousy.
That is, the things you need from a digital assistant -- you really need, and you need them right now. Otherwise you wouldn't be using a digital assistant to look them up, you'd look them up on your laptop like everyone else. I know, I know, in Africa no one has a laptop and their only computing device is a phone, etc, but in the first world, access to fully powered machines with full keyboard interfaces are ubiquitous, people prefer to use these types of interfaces for general research rather then their phones, and we turn to voice assistants when we are in the middle of doing some task and need specific information to assist us in completing that task. Therefore the expectations are pretty high. The frustration level of say, getting bad information about a store being open that you are on your way to is way higher than getting bad information about the height of the Eiffel tower.
To this day, I can't find out the hours of the stores I am going to. The whole experience is really frustrating:
"Siri! When does cup-o-Java close?" shows directions to random stores "Siri! What are the hours of cup-o-Java" shows directions to other random stores "Siri! Is Cup-o-Java closed or open?" shows directions to more random stores
On the other hand, I was stunned that I could walk the Byzantine alleys of Venice and get precise information about turning left in 100 feet to get to my favorite campo. These are tiny mazes of backalleys some wide enough for just a single person to walk down, and in which no cars are allowed, but they're all precisely mapped out so that I'm never lost any more. You can drop me pretty much anywhere and google maps will guide me out of there. But it can't tell me whether the museum in my destination is open today, or how much admission costs today, or whether they accept credit cards.
I'm increasingly believing that the answer is: "No." These machines (especially Alexa) are rapidly gaining popularity while still providing pretty ragged answers to search queries. So we should start asking: "Are they taking on a different function that didn't match our early expectations?"
In a word, yeah. Alexa is a really nifty jukebox for those of us that don't have the good sense to create formal playlists. It's a handy kitchen timer, especially if you've got multiple pots doing different things. It's a better alarm clock and a better purveyor of soothing bedtime sounds. (If you're asking: Good god, how many people really want or need that, think: Fussing infants.)
Smartphones already provide pretty excellent search results on the fly. I'm not sure voice-powered assistants will re-solve that problem with great success. But there are a surprising number of rudimentary needs around the house for which a voice-enabled device becomes quite handy.
I agree that voice interfaces aren't good for a lot of things. How do I cook XYZ? probably isn't suited. But overall performance just isn't that great.
That is perfectly suited to voice if voice worked. When I call my mom for the recipe for cake it would be a whole lot easier if my mom would say "beat the eggs for 1 minute", listen for the beater to start and then say stop after one minute. My mom has better things to do with her time than walk me through the recipe, but an assistant should be able to do this.
Of course I have just transformed the problem into something that technology isn't able to do. However the problem isn't with the voice interface it is our AI isn't yet up to all that. (poor AI, every time they do something useful we rename it and move the goal posts)
Tasks that require a high amount of breadth, like search, don't scale well to a voice interface.
http://stevenhickson.blogspot.com/2013/06/installing-and-upd...
[1] https://github.com/StevenHickson/PiAUISuite/blob/master/Voic...
[2] http://jasperproject.github.io/documentation/installation/#i...
[3] https://cmusphinx.github.io/
[4] https://github.com/julius-speech/julius
Edit: I dug a bit deeper and found this comparison [5] which suggests that Kaldi [6] is the best toolkit to use.
Could Google reasonably put their current voice recognition in a small device?
Software that limits itself to specific keywords can absolutely work well. After all, Google and Alexa do it with their wake words. A voice activated timer could be built fairly easily if it hasn't been done already.
General purpose is a lot harder--and then you need the Internet connection for a lot of the queries anyway so there's no real reason to build in local voice recognition if you then can't really do anything useful with it.
TIL I am a fussing infant. I love that "sounds of the rain forest" bs (though I don't use an echo).
"Should I take an umbrella tomorrow?"
"No..." and shows me tomorrows forecast.
"What about the day after?"
"No..." and shows a forecast that when I look more closely at I notice is for today.
Neither of these also spotted that I'm heading to another city tomorrow, which is in my calendar. If I changed it to "do I need to take an umbrella for my trip tomorrow" it just searches google and gives me a search result suggesting I take a small folding umbrella... for a trip to Thailand.
Also, it might be possible to create a much better assistant today, but it would be too expensive to offer to the public for free. What if it requires 100 TPUs to run?
The other main mistake was (despite it getting the conversational part right) thinking "the day after" means today.
I don't think these things require TPUs or new NLP.
For example, Google Assistant doesn't play nice with YouTube. I guess it isn't in Google's interest to have a free agent software serve a free music site when the same service makes money for Amazon (Alexa). They wanted to make their agent profitable, so didn't let it be extended for free. Just my guess.
What you are seeing here is a much simpler CS problem, but much harder social problem. It is "why can't my applications talk to each other?". I really doubt it will be solved in 10 years.
1 - Humans don't have a perfect solution for it, and machines are still worse, but not that worse.
That means the end to end performance of the system is bad, and it isn't clear how to systematically fix it.
In this particular case there is the added system integration too, which is "just programming".
1. There is an in-built assumption I am always where I currently am. This part isn't anything to do with the NLP.
2. "the day after" is translated to "today", possibly. [edit - see lower, it is in only one case]
3. This is more of an NLP one, it understands that the context of "weather" carries from one question to the next, but not the timing. So asking for the weather tomorrow and then "one day later" gets tomorrows as well. [edit - more complex than this, it's actually working in some areas and not in others]
I'd like to see what user-stories it's trying to solve, because apart from setting timers and alarms it's been massively hit and miss for me.
I tried to repeat what I'd put in and this time I had:
"What's the weather the day after" - translated to tomorrow, with no context, that makes sense.
"What's the weather today?" - weather today, followed by "What about the day after?" which gave me results about the film.
"What's the weather tomorrow?" - weather tomorrow, followed by "what about the day after?" which then worked.
So it works just fine for "weather" but not for asking if I need to take an umbrella. And it doesn't work if I ask for today then the day after, but does for tomorrow and the day after.
Why does the context get passed on for the day correctly for "tomorrow" but not for "today"? Why can it get that "the day after" means tomorrow, unless I've asked about an umbrella in which case it means today? At the core of my question is how is it this inconsistent?
What tests have they got around this?
Of course if I ask for a fact like "How many inches in a meter" I just want the answer. But, if I ask "what's the weather going to be tomorrow" I might prefer an answer like "badweather.com says it's going to rain tomorrow" so I can then think (ugh, badweather.com is always wrong) and ask "What does goodweather.com say about tomorrow's weather". Ideally I could ask the assistant to use a particular site by default. This is specially true for me because Siri's default doesn't seem very accurate to me being that I live on the other side of the world from the offices of the company they use for weather info.
I'm sorry for the confusion. My issue with quoting webpages is not that it quotes them -- that's fine -- but that it does so in a very verbose manner. This leads to information overload. For example:
Me: What is the boiling point of water at an altitude of 1km?
Google: At sea level, water boils at 212 °F. With each 500-feet increase in elevation, the boiling point of water is lowered by just under 1 °F. At 7,500 feet, for example, water boils at about 198 °F. Because water boils at a lower temperature at higher elevations, foods that are prepared by boiling or simmering will cook at
The problem is with the vast quantity of information, and the fact that some is both irrelevant and truncated. The last sentence is incomplete and cut off, yet as a listener I have no way of knowing this. I will thus try to remember it, at the expense of the facts that came previously.
When reading a webpage, the important part is to read the specific parts of interest, and not overload the user. If it can't do that, it risks providing irrelevant or, quite frankly, confusing data (such as the odd answer to how much a Dreamliner weighs). I don't know if that's better than not providing an answer at all.
Me: "Okay Google, What's the latest album by Death Cab For Cutie?" Google: "The latest ablum by Death Cab For Cutie is Kintsugi"
Me: "Okay Google, Play the album Kintsugi on Spotify" Google: "Okay, asking to play Kintsugi. [album starts playing]"
Or, at least, for you to be able to say "Play it!" rather than the unnatural "Okay Google, play the album Kintsugi on Spotify"...
Maybe not a bad idea for a gameshow.
Give it some parameters for a trip you're taking. It comes back with some options and follow-up questions. We are a long way from that point. Even that not-so-sharp intern has a huge amount of internalized knowledge about general preferences, cities, airports, etc. and probably knows questions to ask to narrow things down.
I strongly suspect there are other domains where a lot of people are assuming we're 90% there and we're not.
[1] Not to insult interns or any other group. I just mean you don't need to be at experienced executive assistant level to be really useful.
Also, I can accurately get the weather from Siri most of the time.
Edit: And "What's the height in meters of the Empire State Building?" gets me "381 meters, 443 meters to tip"
Both Google Home and Alexa have very little problem recognising my speech whilst friends often struggle even when they seem to say the precise same phrase, to the point it's mildly entertaining. With Google I suspect they've tailored to my voice (I've used voice commands extensively for several years) but I've only had a Dot briefly and it worked well from the start. Another surprise is that they cope well with my perculiarly English English phrasing and pronunciation, but I'm sure there are lots of less widely spoken dialects that would throw them.
I suspect that the microphone array has a lot to do with it. Anecdotally, I've read pieces by people saying that homebrew "Echos" together with the Alexa APIs aren't as good as an actual Alexa.
My wife who, unlike me, is not a native English speaker has probably a 10% success rate. This is why any kind of forthcoming voice-response Apple device is completely a nonstarter to me.
It might actually be better of Apple were to license the speech recognition or use the Cloud Speech API.
Didn't Apple ditch the WolframAlpha integration pretty quickly after Siri was released? I remember a lot of the Wolfram type queries stopped working shortly after release.
Where does the Jackfruit grow?
What is the boiling point of water at an altitude of 1km?
What is 1km in feet?
How far away is Disneyland?
For "km" I said "kilometer" and not "km".
I think these digital assistants are nerfed by the real-time response requirement. I'd be happy to ask some of those questions, and get a pop up in a few minutes. And they could be of much higher quality as they can be processed and better researched.
I want to be able to ask basic factual questions while driving, get the answers and dig deeper.
Until then I would like an audio service like Google Helpouts used to be, on demand. Like Magic service.
Why don't apps aggregate? The "can you handle this?" api endpoint is frequently returns false positives, and the proper API is really slow (multiple seconds) for negatives . If we get a false positive, or something hard to detect as a negative, that's the only answer we can show. And since a voice assistant is expected to return one answer quickly, this is straight out.
Google "Siri getting dumber" for reference.
This idea, while simple in principle, might be kinda annoying in practice. You're still left with similar issues - how do you decide which talking cylinder service answered the question best? Do you play all of the answers? For me I'm fairly sure listening to all of them in a row would frustrate me even further - just waiting for Alexa to finish telling me the news headlines is sometimes kinda annoying, especially when that information in visual form can be grokked almost instantly. Many of these devices, especially the Google one, are getting better at context based followup questions - managing who to send your follow up question to could be kinda crappy as well. I suppose you could do one device that could ask each service individually ("Alexa...", "Ok Google..."), but in my experience as soon as I get one bad answer, I inevitably just use google.com to find what I need rather than risk wasting my time on another failed conversation.
The main part that I've found hard to do in home rolled voice assistants is microphone arrays. Almost all these devices use pretty sophisticated microphone technologies for things like noise cancelling, subject isolation etc, which so far has been non-trivial to do to a similar standard in homemade versions of them. It also certainly used to be the case that creating your own "hotword" system to call the Alexa API was technically against the ToS (it allowed you to use a button press to call Alexa instead), as naturally Amazon would rather you buy a real Echo. No idea if this is still the case, and at any rate Amazon can't really enforce this either, but worth mentioning.
Here's a short highlight: https://www.youtube.com/watch?v=WoI6_z2mfdY Some implementation details in AMA: https://redd.it/5nz3eb
Not sure how other languages cope, I suspect the simpler ones cope much better. We almost need a spoken equivalent of SQL
- OK Google, please download an offline maps area for Yosemite National Park and about 80 kilometers around it.
- OK Google, navigate to Yosemite National Park. Highway 120 is closed so please remember that when we get to an area without reception, do not route me via 120 in the offline maps.
- OK Google, let me know when we reach a place along our route where I can buy an SD card.
- OK Google, let me know when we reach the last Trader Joe's or Safeway along our route that is still open.
- OK Google, when is the next Caltrain arriving in Palo Alto that stops in San Bruno?
- OK Google, if I miss that train, when is the next train?
- OK Google, get me an UberPool for 2 people to Castro St. in Mountain View as long as it's under $10 and would arrive within 30 minutes.
- OK Google, what is the name of the driver and what is the license plate?
- OK Google, navigate to my friend's party on Facebook.
- OK Google, call my friend 5 minutes before we arrive. Their number is in the wall on the event page for my friend's party.
- OK Google, find me a restaurant that has non-americanized Chinese food, recommended with higher frequency on Chinese language websites than English websites, and has vegetarian options.
- OK Google, connect to my Nest thermostat. My username is XXX and my password is YYY. Turn on the air conditioner 30 minutes before we arrive home based on the navigation.
- OK Google, close that Java error that popped up and is covering the navigation.
- OK Google, please zoom out the map slightly so I can see how far we are from the destination.
- OK Google, please go back to the normal navigation view.
- OK Google, navigate to my next calendar event's location.
- OK Google, install Facebook Messenger, login as XXX with password YYY, and message ZZZ saying that I'll be late by however much the Uber app estimates.
- OK Google, turn off my alarm clock whenever I am biking or driving.
- OK Google, let me know if I get an e-mail from XXX in the next 2 hours.
- OK Google, please block all calls except from XXX for the next 1 hour. XXX's phone number is in their e-mail signature.
Yeah. We're a LONG way before assistants are useful. The pieces are all there. It's not a machine learning problem anymore. It's just that there are just way too many walled gardens between the various parties that hold the data necessary to be useful.
This is changing. If Google had their way, they would be assigning IPV6 addresses to bits of dust lying around in your house and trying to assign semantic meaning to them. If they had their way.
These 4 digital assistants should partner with us to help their users find the prices for a pack of Lays chips, iphone, etc. ;)