On top of that the general distrust of the privacy of these systems has stoped a significant number of people (myself included) from wanting to us them at all. I don't have an in home device, and have turned off Siri on my Apple devices.
On top of that the general distrust of the privacy of these systems has stoped a significant number of people (myself included) from wanting to us them at all. I don't have an in home device, and have turned off Siri on my Apple devices.
No clear feedback, a weird timing issues where it just stalls and show the message it' about to send in case it got it wrong.
It's just a terrible UX all around.
The Google ones do support named timers, so you can say “start a pasta timer” and later ask “what’s left on my pasta timer” etc. I thought Siri added this at some point, but I wouldn’t be surprised if not.
Anyway, it does seem to have improved, but I wonder why that stuff wasn't in from day one. It seems pretty obvious to me.
If I say "set a timer for 20 minutes called A" it just ignores the "called A" part.
A HomePod can set multiple timers. A Watch, iPhone, or iPad can only set one timer. There is no obvious technical reason for it. It just seems like only the HomePod team thought it was an important feature.
This becomes annoying if you have multiple devices set to respond to “hey Siri” and the wrong one picks up the request and then refuses to comply.
I think if the accuracy was better and more content/things were available through voice it would be a pretty good input method for any scenario where you don't need visual feedback.
Browsing content, or looking it up, via voice is slow and playful, it always will be. Who what's to be saying "next", "scroll down", or have a full on conversation with an AI to try and work out what you want to play? Our fingers on our hands have evolved to be incredible at interacting with things, we are good at using them. Touch screens or physical UIs will never be superseded by voice.
So yes, there is a small use case for voice for controlling music/tv, or controlling a few things in the home (heating, blinds, lights) but thats it, I don't believe there is this massive opportunity to expand it into our everyday lives where we are constantly interacting with devices via voice.
Humans evolved language to communicate ideas, wants and desires to others for thousands of years. Obviously voice UI is not there now but maybe someday the experience won't be much different than asking the movie rental store clerk for their recommendations for a romantic comedy.
The only people who ever did that were in a romantic comedy.
I bet ya some engineer at Amazon hooked that up manually when they saw a bunch of requests failing, so that’s only gonna work for popular fuzzy naming conventions. I don’t want to have to think “is this a way lots of people are gonna request this song?” before saying it that way.
The more relevant use case is "hmm, I'd like to listen to some prog rock, let me browse what Spotify has and see what takes my fancy". Sure, I could say "VA, play prog rock", but I don't want it to choose for me: I want to browse the available content to remind myself what are my options, and choose one when I see one that looks interesting.
That's why an opensource ToS-violating assistant has chances to work better than legal ones, they can just scrape all those infos off internet. But then, once you go into that grey area, you just end up pirating content already.
my observation of people on the road has led me to conclude that Driving is an activity where people think they can do absolutely anything else while engaged in it.
1. sending messages on phone while driving, one hand on steering wheel.
2. having sex / receiving oral sex.
3. turned around, yelling at kid in back seat to not fight with other kid in back seat.
4. girlfriend having argument with boyfriend, slapping him on arm some, about how she was smart too just a different kind of smart while swerving back and forth in fast merging traffic near the Haight (I was in back seat)
5. it's getting hot in here, time to take my jacket off!
your mileage may vary of course.
Since the voice assistants are incredibly stupid I find it extremely stressful and distracting to ask them for anything while driving.
Saying "Hey Siri, text Fred <pause> I'm on my way but stuck in traffic, eta 4 o'clock" or something along those lines nearly always works fine for me and is no more distracting than having a conversation with somebody in the car with me. If Siri gets some of the message wrong I'll either send a new one using clearer speech or wait until I'm not driving to fix it if the mistake isn't important.
Sure, it would be possible to then allow myself to get distracted by focussing too much on some weird aspect of it, but equally it would be possible to get so emotional in a conversation with somebody sat next to you that you stop paying attention to the road. And we (most people at least) don't say "it's not safe to talk at all while driving", we just make sure not to go over that line of getting too distracted by the conversation.
Until you have several Freds in your contact list. Until you have friends with foreign/uncommon names. As long as you have near-perfect American pronunciation. As long as...
There are too many variables to consider and think of. Sometimes I can't get Siri to reliably understand what music I want (and my Engilsh is pretty darn good), much less anything more advanced.
There isn’t a such thing as an American accent. Ask anyone who is not a native speaker and either hasn’t been to US that long and tries to understand my natural deep southern accent. I can adjust my accent if needed and if I think about it.
- General American https://www.babbel.com/en/magazine/united-states-of-accents-...
- California English https://www.babbel.com/en/magazine/the-united-states-of-acce...
That only works well if you have an accent it recognizes, if you're speech is clear (not slurred, not lisping etc), if you don't stammer, if you don't have any verbal tics that you don't want to show up in the message, and if "Fred" is actually a simple unambigous name.
Otherwise, at best when you want to send a message to "Ioana" it may end up sending a message to "Anna" that says "I'm, ummm, oh my way! and stalking traffic ate a what was it like 4 like maybe 4 and you know what <pause>" (followed by the "4 o'clock" that will no longer be included).
While driving, I wanted to have Siri read a lengthy webpage to me. I pulled up the page, got in the car and asked Siri to "speak screen." Siri says it can't do that when I am driving! What idiot thought that was a necessary safety measure? What if I were the passenger?
Overall, I am stunned at how bad Siri is at things that don't even require AI. It's almost as if this insanely profitable company failed to invest a tiny bit of money into researching ways that people would like to use Siri.
I often go places with my sister (she drives). Her car doesn't allow pairing or swapping bluetooth connections to the car's entertainment system while it's moving. If we want to switch to my phone we have to come to a complete stop.
1. The command set is broad enough or user input is complex enough to make other UIs inefficient. 2. The voice UI is up to the task of correctly interpreting the voice input correctly most of the time
What "most of the time" means for the second item is somewhat personal and use-case specific.
For item 1, examples where voice is better right now, or could be with reasonable NLU improvements:
"Text my wife that I'll be there in five minutes."
"Get me driving directions to the nearest Indian restaurant with at least 3 stars on Yelp."
"Order six rolls of paper towels and a bottle of Windex from Walmart, delivered to my home address, for delivery by Saturday"
"Remind me tomorrow morning to review this web page"
"Create a shopping list with the items from this recipe"
"Create a basic presentation with one slide each for each entry in the table of contents for this book"
Voice can be better. As others have pointed out, as long as it's like playing Zork where half the time the response is "you can't do that" or "I don't understand", voice interfaces will continue to flounder.
Siri never triggers when I'm driving, it just doesn't hear me. I think it's because of the noise of the car or because of my music, but it doesn't work. I have to move my face closer to my phone so that it can hear me, but that's even more dangerous than using the controls.
Same when I'm in the shower and I ask it to change the music, it doesn't hear me, I have to shout and get angry every time.
For what it's worth, it doesn't work either when it's my pocket. When I come home and ask it to turn the lights on, it doesn't answer if it's not in my hand.
See here: https://support.apple.com/en-gb/guide/iphone/iphaff1d606/ios...
this is still potentially a huge domain. one could imagine a benign scenario where voice assistants enhance people's abilities to interact with each other (and digital devices) when a more potent UI is not within reach
privacy concerns (->controversial business models) and technical ability to deliver a desirable service (that people would pay for) might indeed prevent this vision from catching on in the short term
another factor that may complicate adoption might be just cultural / perceptions. It is a somewhat odd thing to be shouting at devices - especially in the presence of other people. User interfaces that interfere strongly with communication habits and behaviors established over millennia (see also wearing VR goggles) might have a harder time seeing adoption outside very specific scenarios
BTW, another use case for speech recognition is when you're carrying a baby around.
My most egregious example of this for me is that there's a grocery store near me that the Google assistant is incapable of finding because of a few people in my contacts list. Whenever I try to ask it for directions to that store, it picks (at pretty much random) one of three of my contacts instead. This is despite the only common part of said contacts' names and the grocery store is that their names all start with the same letter.
Basically, imagine asking for directions to Albertsons, and the assistant giving you directions to Andrew.