Apple Is Manufacturing a Siri Speaker
bloomberg.com
bloomberg.com
It's a huge handicap when you're only allowed to train models on device using just one user's data, or only allowed to use various complicated privacy preserving training methods (I'm not talking about taking a user's data, removing the obviously identifying fields and sending it to the server, which does not really preserve privacy except for simple statistics). There's no way Apple can train the same super complicated deep learning models that Google uses, for example, when Apple can't just gobble up all their users' data. And all of the "provably" private methods for training models incur significant accuracy penalties.
So it's a little rich that users on HN will complain about Siri on one hand, and on the other hand complain about losing privacy.
In the end I think this is a fatal mistake for Apple: people love to complain about privacy, but time and time again reveal through their purchasing choices that they don't actually care, especially when there's usability tradeoffs. Apple would be better off vacuuming up all their users' data if they don't want to be crushed by Google/Amazon.
Also, I find it baffling anyone would want to put an internet connected always on mic in their house no matter who makes it. Insane.
While I would like to see a model where they all share data to train better assistants, I know that will never happen.
Ontop of all that my side project works on voice recognition and trying to figure out what users want, damn it's frustrating not being able to grok at times, but I haven't quite gotten to the point where I'll take all user input, and put it through ml. (Lack of uses haha, go figure).
My understanding is that this device includes a screen. Right now I could use a 17" 16:9 iPad with a stand and maybe an optional detachable keyboard, as I find I'm using my 12" iPad Pro more and more for TV/movie viewing. I could use something like that for a bedroom nightstand.
I don't think Siri will be the focus of this product. I hope not, at least, because it definitely isn't the top assistant. Apple's strength is it's device media capabilities, and they should exploit that instead.
Siri's problem isn't speech to text -- try the dictation button on the iOS keyboard sometime. It's perfect.
Siri's problem is that it's very limited.
ASR quality is dependent on audio quality (plus, of course, a bunch of big-data-driven training) because it's harder to get the words right if you can't hear the sound.
Audio quality is dependent on microphone(s) quality and signal processing (higher signal-to-noise ratio).
And getting that to happen at far-field distances of 10-20 feet in noisy environments is not easy.
So... bottom line: the better the microphones, the higher the ceiling is on NLP performance.
Source: worked on Echo / Alexa for almost 3 years.
Another example is how as an Alexa skill maker, you have to provide utterance / intent mappings. Is that just used to accurately classify intent / entities? Or is it also used to identify which skill to pass a user utterance to as a part of the alexa skill service because there could be very little variation between skill names or inquiries amazon is supposed to actually fulfill when a person is talking ??
Yes, you are right, in that there are ways to blend the two (use data from one to improve the other). However, in the end of the day, the better the system can determine which words were spoken (using whatever technology), the better it can determine the meaning and the intent, and then decide what to do about it...
> Another example is how as an Alexa skill maker, you have to provide utterance / intent mappings. Is that just used to accurately classify intent / entities? Or is it also used to identify which skill to pass a user utterance to as a part of the alexa skill service
It is primarily used for the former (classify utterance / intent mappings for bootstrapping). Over time, with ML, the goal would be to help understand which "skills" apply to which intents. Unfortunately, today, that's not the case (in Alexa) and that's why the skill-specific keywords or names are still needed. From the user's point of view, it would be preferable to be able to say "Alexa, I need a ride to the airport" and have Alexa figure out whether Uber, Lyft, or the light rail service with a station two blocks from your house is the "best" option for you right now, based on price, availability, and your explicit and implicit preferences. Of course, the system would also have to allow you to specifically request a Lyft, if that's what you want.
(And, of course, it should ultimately be proactive and just offer to get you to the airport when it sees a flight in your calendar, of from having scanned the flight purchase confirmation in your email...)
It's in the interest of Google, Facebook and others to get us to believe we have to give up our privacy, that it's the cost of advancement. But it's baloney. It lets them be lazy and continue to use us (the users) as the product for their real customer; advertisers. But it's just not true.
If I could just type something into Siri I would use it countless times more. Most of my frustration comes from it misunderstanding me, often resulting in me pulling out my phone and typing in my question into search engine anyway.
Additionally, I do not want my business spoken aloud.
Google seems to be doing a melded experience based on their Google I/O presentations. A lot of the third party experience will involve visual responses.
I am afraid that soon it is hard for general public to avoid the situation where when they flush the toiled some unexpected device will recognise it and report to the headquarters.
From Apple employee: "only collects the Siri voice clips in order to improve Siri itself"
Which is basically what Google says.
EDIT: both companies have promised "off-line" processing, whatever that means, but I don't know if any already does that. Note that off-line processing does NOT mean they won't store your voice clips on their servers.
Or in the car, I'll pick up the phone and she'll prepare to say "don't start tapping on that thing while you drive!", but I just say "ok google navigate to [restaurant we're going to]".
Voice recognition isn't perfect, but a couple of years ago, android crossed the threshold where it's good enough for me.
What baffles me is that opening some apps is not allowed without unlocking the phone. I should be able to say "Hey Siri, open NPR One" while washing the dishes. And I can, but it tells me I need to unlock my phone - but I used my voice because I have soap all over my hands, so no thanks.
Google Assistant is accessible by typing; it's built into Allo, and either built into Google Search on Android or the latter hooks into Allo if it is installed, I forget which.
So, I mean, you're still out of luck if you prefer interpretive dance as the interaction mode, but you aren't limited to speech.
You seem to imply this doesn't work, but I'm not seeing it.
It's why I canceled Amazon Prime (Funny, after being a customer for over a decade, with Prime since it was offered, cancelling was a single click, with zero follow-up or attempt at retention. Not even a "Oh, hey, why did you decide to cancel after so many years?")
I would suggest Roku make a digital assistant as well but it would probably come in 4 different models ranging from wholly unusable to outrageously expensive, have a purple bow tie, and need to be replaced annually with the latest model.
After all, it even has a VLC app...
Not sure why Spotify isn't using it.
It's basically a hard coded list of events that get triggered if someone says a particular phrase.
Spotify's use case isn't supported.
The iPad and to some extent the Apple TV both have top of the line hardware. The software on the other hand...improvements every 2 or 3 years don't inspire confidence.
Updates to Siri once a year would kill this product.
I just don't see a lot of utility going forward unless they liberalize the service and allow other people to develop new experiences for users. That should also give them new data to mine and allow good voice UI patterns to bubble up.
Not sure how they should go about designing a framework for developers to plug-in to, though. It's a tricky problem. But presumably they've learned some things after 'sharing' some of Siri with Uber and Facebook and that could help them move forward.
I do think that the walled-garden is concerning since you can't use it with all music services, for example, but on the other hand, if its a premium speaker system that might be enough to push me the other way and switch to Apple music. The question is is the product compelling enough to get people to buy into the platform.
In other words, a voice platform that works 10x better than another may only be slightly more useful to the average consumer while if the actual speaker is 10x better, it might be 10x more compelling.
I think part of the problem is that the things 'voice' is good for tend to have humans on the receiving end that then carry out complex, human actions. "Stu, could you schedule a dinner for me at a good local Italian place at 7:30 and invite Scott and Robin?"
Well shoot, that is a lot to unpack! It's a hell of a lot for a computer to unpack. It's a lot for anyone to unpack unless they know you and your circle of contacts. Carrying it out would require interaction with at least 3 people over a variety of mediums.
I suspect that by opening Siri up (And other digital 'assitants') it might promote the growth of infrastructure and services that would begin to make some of these more advanced queries a little more tractable.
To be really compelling, voice needs to offer the whole AI concierge experience. At a basic level, it needs to be able to deal with queries like "Find me a good place to stay in Barcelona" and "Find me a new jacket". It needs to ask questions as needed, to operate with the initiative to search tens or hundreds of sites, and to recognise the useful data in the results.
With Google, Amazon, and the rest, search is becoming more and more of a problem, not less. For non-trivial searches, finding good products and/or reliable information can be incredibly time-consuming.
So currently voice is a bandaid on top of search tools that aren't progressing much, and may even be regressing. Voice has to solve the search problem first before the recognition and context awareness problems really become important.
It looks as if the industry is to trying to do this the other way around. I'm not convinced that's going to work. It works up to a point, but the point isn't as advanced as users expect it to be, and the overall experience can be disappointing.
I never leave my phone somewhere else when I'm at home, so... what's the difference if I add an extra microphone/speaker to the system already in-place?
Even besides that, I've never managed to work out what people who cite privacy concerns are actually worried about; I've never managed to come up with any scenario that both seems worrying and seems to me like it has any plausible chance of actually happening.
If we imagine a hypothetical counterfactual world where Google records everything ever spoken in my home or near my phone, retains these records forever, and analyzes them and uses the analysis to choose which ads to show me, and promptly responds to any requests by law enforcement for recordings with no hesitation or review process, and recordings are available for people at the company to listen to... what would happen that I would be upset about? Fewer poorly-targeted ads? Law enforcement having a slightly more-trivial time than usual locking me in a cage if I happen to offend a law-enforcement officer? Someone I'll never meet listens to me having sex?
It might be the case that I just happen to have a life situation that's uniquely stable and not susceptible to being fired for political opinions or whatever the risk is... but I don't think my attitudes and risk evaluation here are really all that unusual. I speculate without evidence that most people just really don't care if they're being watched unless it's going to have practical consequences.
Ultimately, for me, it boils down to a lack of benefit for me, and a lot of benefit (current and potential) for Google and these other ad-drivers to track my decisions more than I really want or need them to. As it stands, I live a fairly ad-free life, either through payment or blockers --- so the targeted ads are less of a concern. Law enforcement concerns are also extremely limited, as are political (null).
I just don't see the benefits. If I need to know the current weather, I'll look outside. If I need to search something, it'll never be so important that it can't wait until I sit down at my computer.
I guess, overall, I'm less concerned about the privacy aspect as I am the actual personal benefit of these devices. In my mind I relate it to the Dash buttons.
I don't understand why I'd want a tiny speaker sitting in my house listening for commands, when I already have that on my phone and it can provide a display of the results, along with voice.
Obviously your use cases may vary, but with my echo I was pleasantly surprised at how useful it was. I was given a smartwatch for free and was nevertheless disappointed at how useless it was. I can pull my phone out and do stuff with one hand, whereas on a watch you need to hold one hand in place while the other hand operates the device, meaning it's harder to use and less convenient.
These aren't real selling points.
Google Home integrates very well with Google products (out of the box I thought to try asking it to 'play the latest video from [youtube channel] on my shield', and it worked, exactly as expected.
As far as speakers, Amazon started with a device that had very good speakers, and has since downgraded to a puck.
(I also own small Bluetooth speakers from Sony and Anker, and both are appreciably better for music — enough so that I go out of my way to use them while cleaning or whatnot, even though the Echo is right there.)
The ecosystem that Amazon provides is far richer than Apple can hope to provide. Google does seem to be better positioned.
Even if the hardware is amazing, I don't want Apple music (I have prime music) the voice assistant is no better and I have serious doubts about the integrations that will be available.
Whats going to make this DOA is integrations, it's not in Apple's DNA and its going to be the best sounding, most beautiful, most seemless voice assistant nobody wants. I hope I'm proven wrong.
I've ended up dropping my usage to setting timers and turning on/off lights, both of which don't work with Siri in MacOS.
I'd hope that they'd get their own ecosystem in order before adding yet another gadget, but that's not happened time and time again under Cook's leadership.
But I literally have a single contact named "Mom" that is also starred as a favorite.
Yet sometimes Siri calls my mom, and sometimes she asks me "Which Mom?" even though I always choose the same one.
Jerk.
The Google Assistant happily asks me if I want to finish researching "Care Bears" because I once searched for them on Amazon using Firefox in Incognito Mode on a friend's computer but it can't seem to figure out that I always only ever call my wife on her cellphone.
Go to your own contact card, and then set a related name for mother to your Mom’s contact card. The next time you ask Siri to “Call Mom” it will do it.
It’s unnecessarily complicated, it should be able to infer which one, but there’s the workaround.
It takes longer to ask a thing what the weather is than looking at my phone, but that's just how it is.
Right now on the other platforms devs are creating skills but there is no revenue stream for your efforts!
VR/AR is an area where everything is clunky and awkward to use and Apple could probably build something really amazing here but the market size is tiny and the use cases are outside of apples wheelhouse (industrial applications and gaming right now).
To me, this just feels like another jump on the bandwagon. I have an Echo and, while it was awesome at first, with all the skills and everything it feels less amazing than it did when I first got it. Now, I feel like it mishears us all the time, it doesn't give the answers we'd normally get unless we say things in a very specific way, and all the cool tricks ("Who is the mother of dragons?") are just gimmicks now.
Unless Apple can actually make their version useful and able to understand questions that aren't formed in "robot" as the language, I don't see the point of this.
One of the examples is playing ESPN Radio.
Hmm?
They also registered the domain apple.car
The grandparent is likening Apple jumping on the bandwagon after competitors' cars to them jumping on the bandwagon with the speaker.
Lots of others make laptops, smartphones, headphones, set-top boxes... and yet, and yet. They still turn quite a profit.
Thats a big if, though. Services aren't Apple's strongpoint, and their stance on privacy means they may have to do a lot of the processing in the device, and limit their ability to rapidly improve their offering.
They aren't? Tell that to the teams that run iMessage and the iTunes Music Store.
As for the iTunes Music Store, I could believe that it's not used as often as previously because of the popularity of streaming services, but Apple also runs one of those streaming services (Apple Music), and they run iCloud Library which gives you access to all of your purchased music on all of your devices, and of course they also run the iOS App Store and you can't argue that doesn't have massive scale.
If Apple releases the product and it's just a "me too device" with nothing new, then your statement is reasonable.
Nope, nope, nope, and nope.