Tone also came with the unfortunate side effect of Google software having constant access to your microphone.
0: https://chrome.google.com/webstore/detail/google-tone/nnckeh...
Tone also came with the unfortunate side effect of Google software having constant access to your microphone.
0: https://chrome.google.com/webstore/detail/google-tone/nnckeh...
The basic problem is, the wifi password needs to be shared with a device without an interface.
The traditional method at the time was the device boots with an unsecured wireless network. And an app on the phone is used to connect to the network, and share the wifi password, and then rendezvous on the wifi network to continue provisioning. I think there are some protocols for this that I don't fully remember now. There's lots to consider in this method, such as can someone sniff the transfer, how is that protected, what is the range to pick up, etc. Can someone trick the app into connecting to some other device, etc, etc. If you put mitigations in, how can those be countered.
With sound / ultrasound based provisioning (the device I'm talking about already had a microphone and speaker), the range is more limited to who can hear the device. The signal strength is much weaker, an eavesdropper might need to be in the same room. This might allow weaknesses in the overall model of key exchange to be less of a concern, as the properties change to something a lot harder to intercept.
When considering robustness of support, one nice thing about the WIFI approach is in theory you can fall back to just a web browser. So basically any device could be used for provisioning.
And the audio method, would probably support a great set of devices. We didn't actually consider ultrasonic, just an audio codec that was in use at the time and doesn't sound awful.
I feel like Apple’s ‘Find my’ tech could work to get positional data and make it possible. You’d need whitelists and the like of course.
The issue at play here is that most devices have no concept of locality or relative positioning. They may know (at best) approximate distance between itself and another device based on the type of transmission, signal level, noise levels, etc.
Phones also can't understand intent. How does it know whether it doesn't mean the person sitting behind your friend or your friend?
I think there would still be a quick confirmation button listing who owns what phone you're sending it to, and it could default to people you know in vague situations - but generally whoever is closer.
In today's world, most phones or "smart" devices are also constantly listening; I want to believe they don't listen until the trigger phrase is uttered, which could also be implemented for these ultrasonic applications, but I'm not entirely convinced and them always listening is but a silent over-the-air update or setting change away.
They can't know whether the phrase was uttered unless they constantly listen.
The microphone is active and "listening" all the time.
The firmware that detects the wake word compares the constant input stream against waveforms that are designated "wake words". Firmware can be sometimes updated for custom or trained words, but it doesn't hold a large dictionary.
If a reasonable match is found, it kicks the full recording/recognition/streaming code, squirts any buffered audio at it (to catch words that come directly after the wake word and before the full handler is ready), and then things proceed according to plan. Depending on the device and service, recognition might happen locally or in the cloud.