Hey Siri: An On-Device DNN-Powered Voice Trigger for Apple’s Personal Assistant
machinelearning.apple.com
machinelearning.apple.com
"We compare the score with a threshold to decide whether to activate Siri. In fact the threshold is not a fixed value. We built in some flexibility to make it easier to activate Siri in difficult conditions while not significantly increasing the number of false activations. There is a primary, or normal threshold, and a lower threshold that does not normally trigger Siri. If the score exceeds the lower threshold but not the upper threshold, then it may be that we missed a genuine “Hey Siri” event. When the score is in this range, the system enters a more sensitive state for a few seconds, so that if the user repeats the phrase, even without making more effort, then Siri triggers. This second-chance mechanism improves the usability of the system significantly, without increasing the false alarm rate too much because it is only in this extra-sensitive state for a short time."
“hey Siri... Siri”
Many, if not all, American English dialects pronounce the words in the fashion I describe. In them, the name 'Alexander' would be
/ˌæ.lɛˈ(gz)æn.dər/
while 'Alexa' would be /əˈlɛ.(ks)ə/
- in both of which, the phoneme corresponding to the letter 'x' is parenthesized.Generally in English 'x' is voiced when it precedes a stressed vowel, which it does in 'Alexander'; in 'Alexa', 'x' precedes a reduced vowel, and therefore would always take the unvoiced pronunciation. (It'd sound very odd to an anglophone ear otherwise - say /əˈlɛ.gzə/ one time out loud and see if you don't feel the same.)
That said, it wouldn't be incorrect to pronounce 'Alexander' in American English with an unvoiced 'x', as
/ˌæ.lɛˈksæn.dər/
but, while I believe some dialects of English may default to this pronunciation, certainly not all do. (Neither of the dialects I speak does so, at the very least.) This pronunciation also produces a "hitch" or break in the word between the unvoiced 'x' and its preceding vowel, which would tend to make it a little odd both to hear and to say.In “Alexander”, the normal pronunciation of the letter “x” is the voiced consonant cluster /gz/, in “Alexa”, it's usually the unvoiced cluster /ks/. (Also, the second “a” is usually different between the two, being /æ/, like the “a” in “pad”, in “Alexander” and /ɑ:/, like the “a” in “father”, in “Alexa”.)
The difference in the “x” is the normal way that the pronunciation of “x” differs when following a stressed vs. unstressed vowel in English.
I'm actually pretty impressed by Alexa's voice recognition so far.
That would trade off both security and privacy for very limited gain.
As an aside, alarms are literally the only thing I've ever used Siri for.
"Hey Siri, set an alarm for {seven,eight,nine} a.m." "Hey Siri, call {mom,dad,[friend]}"
I've created unambiguous nicknames for my friends so I don't get name collisions
She knows my next appointment, and she knows how to give directions, but she refuses to give me directions to my next appointment. If I could throttle her for her stubborn insubordination, I would. lol
It usually works for me too, but sometimes it doesn't, now I know to retry.
Using Siri or Google Assistant for more than one question at a time quickly makes me feel like I'm going insane. "Hey Siri.. Hey Siri.. Hey... Hey..."
I'm hoping Google and Apple fix these subtle annoyances. Or maybe it's just me.
But the sound of the wake word, whether it's just "Alexa" or "Hey Siri" vs "Siri, doesn't seem to deal with the main issue of your complaint, which to me is how limited "conversation" is with the assistant.
If you ask Alexa for the weather, you'll still have to say her name for any followup questions within that immediate context, i.e. "Alexa, what's the weather today? Alexa, what's the weather this weekend?".
Though there are a few functional exceptions in which Alexa will prompt you for additional information without needing to be re-awakened, e.g.
You: "Alexa, set my alarm for 6 'o clock"
Alexa: "Is that 6 'o clock in the morning, or in the evening?"
And then at some point Apple realized they had to make a longer wake word to cut down the number of false positives ("Siri" -> "Hey Siri", from 2 syllables to 3).
Google probably went through the same process ("Google" -> "Okay Google", from 2 syllables to 4).
Amazon probably deliberately chose a 3 syllable name with "Alexa" for the same reason.
I can imagine future improvements where we can have the originally imagined wake words "Siri", "Google" and "Alexa", and at that point I would be most happy with "Siri" because it would be short and not-corporate.
The initial product, before the Apple acquisition, was an iPhone app with a chat interface. I don't recall that it supported voice input.
It is still possible that the founders were thinking ahead to voice input and wake words when they named the company.
seriously?
I've seen so many people prefix 'Hey siri' when they don't really need to:
Holds home, Siri activates "Hey Siri, how's the weather?"
If could just be "How's the weather?" In that instance.
You can now reboot from Settings in iOS11 if the device is responsive and unlocked.
If it's locked or you just want to properly shutdown quickly, press and hold a volume button and the lock button to bring up the SOS mode, which includes the shut down slider.
If the device is unresponsive, it's much less obvious now: https://ios.gadgethacks.com/how-to/force-restart-iphone-x-wh... Volume up, then down, then press and hold the lock button.
Unhelpful while I am driving...
A voice assistant that requires physical controls massively reduces its utility and core purpose.
It is like Schrödinger's UI design. It is both being claimed as the solution to the problem being raised while at the same time being called entirely optional. Sounds to me like there is no actual solution, and this is a bandaid.
On the other hand, if you want to check the weather AND set an alarm, you need to activate Siri twice.
> To avoid running the main processor all day just to listen for the trigger phrase, the iPhone’s Always On Processor (AOP) (a small, low-power auxiliary processor, that is, the embedded Motion Coprocessor) has access to the microphone signal (on 6S and later). We use a small proportion of the AOP’s limited processing power to run a detector with a small version of the acoustic model (DNN). When the score exceeds a threshold the motion coprocessor wakes up the main processor, which analyzes the signal using a larger DNN.
> Apple Watch uses a single-pass “Hey Siri” detector with an acoustic model intermediate in size between those used for the first and second passes on other iOS devices. The “Hey Siri” detector runs only when the watch motion coprocessor detects a wrist raise gesture, which turns the screen on. At that point there is a lot for WatchOS to do—power up, prepare the screen, etc.—so the system allocates “Hey Siri” only a small proportion (~5%) of the rather limited compute budget. It is a challenge to start audio capture in time to catch the start of the trigger phrase, so we make allowances for possible truncation in the way that we initialize the detector.
Another interesting nugget:
> There are thousands of sound classes used by the main recognizer, but only about twenty are needed to account for the target phrase (including an initial silence), and one large class class for everything else. The training process attempts to produce DNN outputs approaching 1 for frames that are labelled with the relevant states and phones, based only on the local sound pattern. The training process adjusts the weights using standard back-propagation and stochastic gradient descent. We have used a variety of neural network training software toolkits, including Theano, Tensorflow, and Kaldi.
Apple doesn't usually do tech writeups; I imagine a pinhead at the corporate level decided it wasn't worth the risk of leaking any "secret sauce" until now.
0: https://www.cultofmac.com/390181/5-ways-hey-siri-will-change...
Great to see such corporate changes toward developer-friendliness.
At some point I expect Apple to design an audio neural network processor to put on their CPU chips which will allow them to do both phrase recognition and highly accurate speaker dependent text to speech on their devices. It will be yet another way that people who don't build silicon won't be able to compete.
[1] https://static.googleusercontent.com/media/research.google.c...
Attackers might be trying to steal your information from machinelearning.apple.com (for example, passwords, messages, or credit cards). Learn more NET::ERR_CERT_COMMON_NAME_INVALID
Access Denied
You don't have permission to access ".../machinelearning.apple.com/2017/10/01/hey-siri.html" on this server. Reference #...
https://static.googleusercontent.com/media/research.google.c...