Google Now vs. Siri vs. Cortana – The Great Knowledge Box Showdown
stonetemple.com
stonetemple.com
Some of the killer features of Siri for me are being able to to write emails, send messages and get quick answers to general questions throughout the day. She is much better at the first two than she is at the last one though, but even then she is not too bad.
And now that Apple has hooked her into much more of iOS in iOS 8, it means that more of the OS is open to me than before just by simply using my voice. And the "Hey Siri!" feature is one of the best accessibility features they've added so far.
In what way? The OP seems to use it in the most basic way possible--dictation and simple searches.
I wouldn't want to put words into his mouth, but I would suggest he meant using Siri as a disability aid. Might be wrong though! :-)
Fun specific example at the lock screen after holding down the iPhone's button for a couple seconds: "Read the latest message from my wife to me." Then, still at the lock screen, "Reply to my wife I love you".
Maybe it's just my experience, but I did not care much for Siri until she allowed me greater control over my device. When she first came out, you could not ask her to read the screen; later, the program was able to toggle assisting configuration like VoiceOver by command.
Siri was essentially a curiosity for me until they allowed me greater control over the device, but I use my iPhone in exactly the way you've just described.
Along with "am I meeting with my physiotherapist today?", and "arrange an appointment with my doctor, physiotherapist and partner" at which point Siri creates the appointment in the calendar and sends an email invite to the parties mentioned. Assuming those people are in your contacts obviously.
I'm using a device called a TrackerPro[0] which tracks a small reflective dot on my glasses, this translates my head movements into cursor movements. To initiate a left mouse click I use a [buddy button[1] under my right index finger, this means that using just my head and right index finger I can control the mouse just like anybody else.
To input any kind of text into the computer I use DragonDictate for Mac[2], this has the dual function of enabling me to do speech to text really well but also to trigger little shell scripts and AppleScript's I've written.
So, if I say "Xylophone Google" Dragon will recognise that as a command and jump to the Google homepage. That's just a really simple example of what you can do, but as you might imagine it's one I use a lot each day. :-)
[0]:http://www.ablenetinc.com/Assistive-Technology/Computer-Acce... [1]:http://www.ablenetinc.com/Assistive-Technology/Switches/Budd... [2]:http://www.nuance.com/for-individuals/by-product/dragon-for-...
How do you handle typos in iOS as a quadriplegic? They're a pain in the ass to fix by hand and I didn't even know you could do it by voice.
I also use the Apple Switch Control[0] feature which enables me to navigate text and correct typos, this is done through a series of sequential button presses which enables me to use menus to access all of the features of the phone. I use Switch Control in conjunction with the Tecla Shield[1] which uses Bluetooth to connect up a button by my cheek and the iPhone.
So yes, fixing typos when you're quadriplegic takes a little while but compared to the alternative it's awesome!
[0]:http://support.apple.com/kb/HT5886 [1]:http://gettecla.com/pages/tecla
Edit: Stupid voice dictation mistake!
Also, does the CPU get hot whilst it's plugged in and constantly listening every few milliseconds?
For the iPhone 6, Apple claims 10 days standby time, 50 hours audio playback time (which uses dedicated hardware for most of the work) and 10-11 hours of internet use. Which doesn't isolate the CPU, of course, what with the screen and radios, but should give some idea of what's going on.
I wonder if it would be possible to use that dedicated audio playback hardware for this purpose, I would definitely pay a relatively small sum towards RND if anybody is interested!
I also wonder if there's a worry that average users might turn this on accidentally, running their battery out and then shouting at Apple? If that's the case, it could be happily buried in the Accessibility menu I think. Seriously, this feature would enable me to keep my iPhone in my top pocket while I go out and be able to access Siri rather than cart about all the other equipment. It would be great!
Edit: Why isn't accessibility in Dragon's menu?!
I'd bet that allowing "Hey, Siri" on battery will be one of the first tweaks available once an iOS 8 jailbreak is released. Might be something to keep an eye out for if you're brave, jailbreak-wise.
Also, have you considered an external battery providing power over USB? The phone will be in a charging state that way. Sort of a silly workaround, but maybe worth it.
I think an add-on for this would be fairly easy. All you need to do to activate Siri is long-press the home button, which you can do externally through the microphone port or over Bluetooth. A widget that listens for the magic phrase and then simulates a button press would be doable. No idea how much it would cost or how hard it would be to build....
He was a smart guy and was really interested in all sorts of things, but he was sadly limited. Posts like yours remind me of him and make me wonder how he'd get on with modern technology. I do recall investigating speech recognition back in the day, I think it was an early version of Dragon, but the cost was just too high and the uses, with 80s-era computers and no internet or even nearby BBSs to connect to, were just too little. Anyway, thanks for sharing your experience.
I can only imagine what I would have done to indulge my intellectual curiosity without the Internet, sure I could have waited for people to bring me books, but that just makes me shudder. I also tried the early versions of Dragon and they were great, as long as you could also use a mouse to correct little typos and similar. So yes, not much use to a quadriplegic.
Edit: I forgot the: shudder
I haven't used Cortana, and have only had disappointing experiences with Siri (after coming from Google Voice Search); how does the performance on Google Voice Search Compare? I wonder if the performance rankings are just a complete reversal of the quality rankings; If so, I wonder who's closest to the sweet spot?
Have you tried it on iOS 8 since it came out? They finally added streaming voice recognition, which speeds it up a lot in my experience. Previously, it waited until you were done speaking to initiate the upload, which added a pretty big delay to any response.
There are some things I feel Cortana does that Siri and Google Now does.
Every morning Cortana gives me the weather, appointments on my calendar, news stories and how long my commute will take given the current traffic conditions. It also learns when I normally leave for work, so about 10 minutes before I usually leave, I get an update on how long it will take to get home given the current conditions.
To me, Cortana is more like an assistant. You can set stuff so its off limits and she will ignore, and likewise, tell her to remember certain reminders like you pointed out.
Do either Google Now or Siri have these features that actively learn stuff about your preferences?
Things like these are a much more useful part of it than voice detection, I very rarely interact with it via voice, and tend to just do searches through my normal browser. It also hooks into Gmail and can give you reminders about things like flights you've booked, although that raises questions about privacy, and whether you really want Google to be trawling through your emails.
Eh, it also does it at really inappropriate times.
About 30 minutes after I get into the office it starts telling me how long it'd take me to get home, and that card stays up in Google Now until I actually do go home!
For a good believable test your sampling methodology is very important. In fact, you want to have a random sample that is free of biases instead of somebody cherry picking queries. Perhaps you may do random weighted sample of queries with weights = frequency which represents usage pattern more closely. In any case, it's very important to describe your sampling methodology or otherwise this kind of testing has little value.
Also commands that already sort of work lead to more queries issues lead to better quality.
Her location-based and people-based reminders have been a killer feature for me. I can open Cortana and say "Next time I talk to my mother, remind me to ask her for her cheesey corn recipe". Then, next time I open the messaging app to text my mother, or open the email app to email her, or when the phone's GPS realizes I'm over at her house, Cortana will show that reminder[3].
It also works for location-based reminders: "Next time I'm at the hardware store, remind me to buy softener salt". It's been about six months since I've last used Android, but I don't think Google Now could do these types of reminders when I switched.
[1]. Two other reasons: I much prefer "Live Tiles" to widgets or app buttons, and the Lumia line's cameras blow most other phones out of the water.
[2]. According to the /r/WindowsPhone subreddit, Microsoft tends to update their apps on Android and iOS long before they update them on Windows Phone. Additionally, apps on the other two platforms tend to be more feature-complete. This is all anecdotal though, so take it with a grain of salt.
[1]. https://support.google.com/websearch/answer/3122344?hl=en
1. http://www.androidpolice.com/2014/03/17/rumor-remind-me-when... 2. http://www.androidpolice.com/2014/06/06/exclusive-google-wil...
It has proven to be an incredibly powerful feature at least for me. I could repeat the list of things others have listed (e.g. https://news.ycombinator.com/item?id=8435239 ) but why be redundant.
"Especially when you're not otherwise bought into the Google ecosystem"
Well yeah. I have to admit that at work I'm a heavy user of google calendar and email and in personal life maps, navigation, email, calendar and google search. Honestly, that's what I see most iPhone users around me use too. The point here is that if google cloud services are already in your life (as they are in 100's of millions of people's lives), Google Now on Android puts a nice face to it and brings them together in a very user-friendly and intuitive way.
"I'm in the mode of dictating things manually. And Siri shines at that"
Perhaps you didn't read the original posting that we're discussing here. It shows how far Siri has to go to catch up with Google's Now's voice search.
Some examples of things I've not done anything to explictly be told about, but have been relevant when shown:
- Directions to a restaurant I'm about to go to
- Flight delayed notification
- Traffic time abnormalities going to/from work
- Package shipping updates
- Sports scores for teams I'm interested in
Edit: some data at http://www.quora.com/Is-anyone-working-on-an-open-source-ver... & https://news.ycombinator.com/item?id=4987875
Move along, nothing to see here.
The big problem for open source speech recognition is training data.
Is training really the issue for voice recognition? It has been a problem that has almost been solved for over a decade. Last year I saw this impressive use of Dragon Naturally Speaking for the PC, running in a VM on a Mac, that pretty much worked to code by voice.
https://www.youtube.com/watch?v=8SkdfdXWYaI
The developer mentioned that he didn't have any luck with Sphinx.
Xah Lee summarized the talk here: http://ergoemacs.org/emacs/using_voice_to_code.html
The system in the video you link is single speaker, closed vocabulary. You need massive training data for multi-speaker, open vocabulary.
Training is a huge issue for voice recognition. It's the only way Google and Apple have managed to take voice recognition from "works 80% of the time, but that is still bad enough to be totally usable" to "this actually works!". Maybe you don't remember how bad voice recognition was 10 years ago.
To give you an idea how important it is, on OSX you have the option to download data to improve offline voice recognition. It's something like 500 MB. And that's the result of the training.
What exactly is needed for training - audio recordings with transcripts, human validation of recognized text?
There are successful crowdsourced efforts for proofreading of OCR'ed text. Archive.org could host a CC-licensed archive of sound & transcripts.
Recognition of the human voice is almost like writing, hopefully everyone could have access.
Edit: how much disk space would be needed - TB or PB?
That is a common size for LVCSR, and you need something around that area to get good performance (maybe minimum 100h). In academic papers by Google, they usually use their own private training data set, with e.g. 1900h. (E.g.: http://arxiv.org/pdf/1402.1128.pdf)
Some crowdsourced effort to collect transcribed audio under a CC-licence would be great!
Maybe also: https://librivox.org - has audiobooks read by volunteers, plus the book text.
In all cases, the total patent life for the product with the patent extension cannot exceed 14 years from the product’s approval date, or in other words, 14 years of potential marketing time. I
f the patent life of the product after approval has 14 or more years, the product would not be eligible for patent extension.
all regulatory periods are divided into a testing phase and an agency approval phase. The regulatory review period that occurs after the patent to be extended was issued is eligible to be counted towards the following calculation:
First, each phase of the regulatory review period is reduced by any time that the applicant did act not act with due diligence during that phase. The reduction in time would only occur after an FDA finding that the company did not act with due diligence.
Second, after any such reduction, one-half of the time remaining in the testing phase would be added to the time remaining in the approval phase to comprise the total period eligible for extension.
Third, all of the eligible period can be counted unless to do so would result in a total remaining patent term from the date of approval of a marketing application of more than fourteen years. An additional limitation on the period of extension is that the extension cannot exceed five years. For example, if an approved drug product which is eligible for the maximum of five years of extension had ten years of original patent term left at the end of its regulatory review period, then only four of the five years could be counted towards extension. The Patent Trademark Office is responsible for determining the period of extension.
you can read more here - http://www.fda.gov/Drugs/DevelopmentApprovalProcess/SmallBus...
You can even download trained models: http://kaldi-asr.org/
It supports many state-of-the-art methods, like DNNs, sequence training, etc. So you can get quite good results with it. To train it yourself, of course you need some good training data from somewhere.
Translate {words} to {language}
My flights
My schedule
My packages
What time is at {city}?
How tall/old/heavy is {important person}?
{description of a photos} in my photos (like 'ocean in my photos')
Note: if you have G+ photos. description doesn't need to be typed in the photo meta info
compare {food} and {food} (nutritions)This!
My love affair with Google Now started with exactly this query. Me and my friends were debating over Hina Rabbani Khar's height[1].
[1] https://www.google.co.in/search?q=How+tall+is+Hina+Rabbani+K...
123 {currency} in {currency}
"Remind me to check my air filter when I get home"
"The number of atoms in the entire observable universe is estimated to be within the range of 1078 to 1082."
Thank you Google! http://www.google.com/search?q=how+many+atoms+are+in+the+uni...]
Actually, they all seem to me pretty limited compared to the answers you get from Wolfram Alpha.
Also c.f. the responses to one of their test questions, "how much is a quarter cup of butter?". Google makes fun of the inquiry. Wolfram Alpha gives you a thorough nutritional profile, and links to variations based on international cup sizing and different types of butter.
The result Google also gives you is wrong (since its also a just a snippet).
In the article, the question "How old is the Lincoln Tunnel" struck me as incorrectly formatted for the parser (I know, that's the point), so I asked Siri, "When was the Lincoln Tunnel built." The Wikipedia article on the Lincoln Tunnel was returned. Wolfram Alpha was listed under other sources, so I chose that. The response? "1937"
Luckily if you say "Wolfram XXX" instead of just "XXX" Siri will route your question straight to Alpha no-questions-asked.
Just a formatting issue.
Solving it isn't really easy, either. When reading a document is a caret a power, a regular expression operator, an exclusive or, or a nose on a smiley? :^)
A lot of context identification work has to be sorted out for smart agents to succeed.
It's also worth noting that Siri gives more verbose answers in "Hey Siri" mode, presumably because it assumes you're not looking at the screen.
This is somewhat frustrating if you don't know the magic incantation to make Google Now do what you are asking it to do. Siri's engineers have done a better job anticipating the various forms of the commands and handling nearly all of them.
When Siri fails for me, I sometimes ask my 7 year old to talk the same thing to Siri and she gets better results. My daughter has a more 'American' accent than me so I have concluded that Android is better at hearing through accents than iOS.
(I haven't yet read the original article but wanted to quickly comment since our observations are completely opposite)
EDIT - iOS tablet user = iPad user.
I then added some simple machine learning to filter search results for the most relevant threads which improved it quite a bit.
>!ask what is the largest prime number?
For primes of the form 2^n - 1 (known as [Mersenne primes](http://en.wikipedia.org/wiki/Mersenne_prime)), a very fast primality test known as the [Lucas-Lehmer test](http://en.wikipedia.org/wiki/Lucas%E2%80%93Lehmer_primality_...) is available. The ten largest currently known prime numbers are all Mersenne primes.
>!ask what's the airspeed velocity of an unladen swallow?
African or European?
> !ask tell me a joke
I was eating chicken tonight last night, and got a little bone in one of the fillets. So I said "Fillet? More like fill-it with bones! Haha!" Then I looked around at the empty table and quietly sobbed into my bowl.
As I said, it only works for questions that are likely to have been posted on reddit before. Not a general search engine.
http://www.amazon.com/Simple-Heuristics-That-Make-Smart/dp/0...
BTW, I worked with Knowledge Graph last year when I consulted at Google. It is an incredibly nice project. The team, who helped me when I needed help, was great - constantly improving the platform.
I am also using the IBM Watson APIs right now while helping another customer, so I feel like I am getting a broad view of what is available.
I expect that Knowledge Box, Cordova, Siri, IBM Watson, etc. are all going to get much, much better in the coming years and will change the way most people use computing devices. Exciting times!
Knowledge Graph builds on linked data and semantic web technologies to encode knowledge that is served on an efficient scalable platform.
IBM Watson is a system for ingesting large amounts of text and for then allowing natural language queries on the information in the text.
Both are valuable properties.
"What's the population of New York?"
Me: "Siri, find me a target."
<finds several targets>
Me: "Directions."
Siri: "Which target?"
Me: "Third one."
<gives directions to third one in list>
Also:
Me: "Siri, find me restaurants."
<lists restaurants>
Me: "Review for Blue Duck Tavern"
<lists review>
Me: "Other restaurants"
<lists restaurants>
Me: "Reservation for Founding Farmers."
Google Now will do some coreferences, but Siri is almost modal. You can talk to it like an assistant instead of trying to formulate everything as a search query. I had a Nexus 5 almost a year before getting my 6+, and I was always jealous of how much more practically useful Siri was on my wife's iPhone 5.
Me: "Find me a target"
<nearest target shows up>
Me: "Directions"
<changes to "directions to target">
<shows a list of targets to select>
Me: "first one"
<chooses Portland Galleria Target>
Has there been any major advances in voice recognition other than just growing your speech corpus?
I remember back in the day, most of the errors were the exact same errors that a human would make. Even you and I only really hear 95% or so of the words someone says. But we can fill in the rest from the context. It always amazes me when I dictate some sentence to my phone, and at the end I see one of the words change to another that sounds almost identical but makes much more sense in the entire context of the sentence.
edit thought I'd add this. I just had a conversation with a united agent and they probably understood less than 70% of what I was saying. it was beyond frustrating I wish that I was actually talking to my phone instead.
And usually I'm trying to decide between a one sentence typing task or asking Siri to do it. The five second wait really throws off my time "profit margin".
I work on a UX team of three. One of our guys has a MotoX, One guy has a Windows Phone, I have an iPhone 6, and based off of what happens at work it seems that Google Now and Siri are slightly more functional.
This is from merely observations, but it seems like Google Now is faster than Siri. And Siri is better at accurately hearing/understanding the words you say.
but I AM suprised, anecdotally, how good Siri is vs previous iterations on my iphone 6 and OS 8
It feels close to "good enough" for the majority of functions I actually use eg: dictate an email, get directions, lookup a contact, and dial a phone number.
My sense is that Apple will nail the base functions so that most users, self included, won't notice a difference between the two.
In terms of 80 % functionality I think all three are quite close. In terms of general features I think Android is already lightyears ahead. But I need a solid day-to-day phone with consistent usablitity concepts and I'm very tempted to give WP a try as it seems to be the cleanest here. Let's see what WP 10 brings ...
This is how I interpret most responses when people talk about Apple.
But we all have our biases as well. (Even as an Apple user I fully expect Google to always win the Knowledge Vault race. Its in their wheelhouse more so than Apple's.)
This is exactly why I'm shocked that Cortana's accuracy is so much lower than Siri's (under the assumption that this test is reasonably valid). They've got a world-class research arm and they've run a search engine for yeaaars, and somehow they dramatically underperform a company whose biggest fans would even admit has a spotty record when it comes to services. I guess the "time on the market" advantage is a lot more dramatic than I would've thought.
Heh, you not only proved kumarm's point about apple fanboi'sm but served the proof on a silver platter with a little side of dessert.
> In addition, this was a straight up knowledge box comparison, not a personal assistant comparison. For purposes of this study, a “knowledge box” or “knowledge panel” is defined as content in the search results that attempts to directly answer a question asked in a search query.
https://news.ycombinator.com/item?id=8430202
Anyway, there I was trying to find out which speech recognition software was more accurate. How do these compare to Dragon?
It seems like we're quite close to being able to actually using voice dictation without the frustration.
Do they use template databases for different topics? (like afaik WolframAlpha) Do they use (scientific) ontologies or are the templates more flat and stored in a SQL database? [and web search results as fallback]
What do you mean by a template database?
"How tall [object]?" / "[object] height"
"Becoming a [job]"
"How much is a [unit] of [object]?"
"How old is [object]"
"When is the next [event]"
My question is: Do Siri/Cortana/GoogleNow use such flat templates (per topic) or do the use an IR ontology? http://en.wikipedia.org/wiki/Ontology_(information_science)
Then we execute that symbolic representation using a variety of strategies. There are funnily enough quite a few cases where we can understand the question but don't have the curated algorithms to actually answer it.
That being said I really like Google Now but the accuracy of speech to text is what killed it for me.
I use it frequently in the car so maybe the background noise is affecting it.
edit: And #12 on the front page right now appears to be essentially just that.
"What's going to be on my ballot?"
Can Watson even answer this?