My understanding is that this device includes a screen. Right now I could use a 17" 16:9 iPad with a stand and maybe an optional detachable keyboard, as I find I'm using my 12" iPad Pro more and more for TV/movie viewing. I could use something like that for a bedroom nightstand.
I don't think Siri will be the focus of this product. I hope not, at least, because it definitely isn't the top assistant. Apple's strength is it's device media capabilities, and they should exploit that instead.
It's in the interest of Google, Facebook and others to get us to believe we have to give up our privacy, that it's the cost of advancement. But it's baloney. It lets them be lazy and continue to use us (the users) as the product for their real customer; advertisers. But it's just not true.
Siri's problem isn't speech to text -- try the dictation button on the iOS keyboard sometime. It's perfect.
Siri's problem is that it's very limited.
ASR quality is dependent on audio quality (plus, of course, a bunch of big-data-driven training) because it's harder to get the words right if you can't hear the sound.
Audio quality is dependent on microphone(s) quality and signal processing (higher signal-to-noise ratio).
And getting that to happen at far-field distances of 10-20 feet in noisy environments is not easy.
So... bottom line: the better the microphones, the higher the ceiling is on NLP performance.
Source: worked on Echo / Alexa for almost 3 years.
Another example is how as an Alexa skill maker, you have to provide utterance / intent mappings. Is that just used to accurately classify intent / entities? Or is it also used to identify which skill to pass a user utterance to as a part of the alexa skill service because there could be very little variation between skill names or inquiries amazon is supposed to actually fulfill when a person is talking ??
Yes, you are right, in that there are ways to blend the two (use data from one to improve the other). However, in the end of the day, the better the system can determine which words were spoken (using whatever technology), the better it can determine the meaning and the intent, and then decide what to do about it...
> Another example is how as an Alexa skill maker, you have to provide utterance / intent mappings. Is that just used to accurately classify intent / entities? Or is it also used to identify which skill to pass a user utterance to as a part of the alexa skill service
It is primarily used for the former (classify utterance / intent mappings for bootstrapping). Over time, with ML, the goal would be to help understand which "skills" apply to which intents. Unfortunately, today, that's not the case (in Alexa) and that's why the skill-specific keywords or names are still needed. From the user's point of view, it would be preferable to be able to say "Alexa, I need a ride to the airport" and have Alexa figure out whether Uber, Lyft, or the light rail service with a station two blocks from your house is the "best" option for you right now, based on price, availability, and your explicit and implicit preferences. Of course, the system would also have to allow you to specifically request a Lyft, if that's what you want.
(And, of course, it should ultimately be proactive and just offer to get you to the airport when it sees a flight in your calendar, of from having scanned the flight purchase confirmation in your email...)