Voice assistants are not doing it for big tech
theregister.com
theregister.com
Instead, it's a guessing game about syntax and semantics, and frequently a source of frustration. There are many failure points: it can "hear" you wrong, it can miss the wake word, it can hear correctly but interpret wrong, miss context clues, or simply be unable to process whatever the request is. In my experience, most normal people either relegate voice commands to ultra-specific tasks, like timers, weather, and music, and that's that. Google and Alexa are relatively good at "trivia" questions, but Siri is a complete failure. All systems have edge cases that make them brittle.
I think there's potential here. Cortana was the most promising: an assistant that's integrated into the OS and can change any setting or perform anything on-screen would, again, be really awesome. We just don't have that. I think maybe OS-wide + GPT 4 (or later) might get closer to what we expect, but it's just not great right now. I really want to be able to say something as unstructured as "hey siri, create alarms every 5 minutes starting at 6am tomorrow" or "hey siri, when I get home every day, turn on all of the lights, change my focus to personal, and turn on the news". There /is/ power to-be-had, but nobody has really tapped it.
Maybe it's possible to learn a working vocabulary and know how to command a voice assistant. I know my way around several command lines, but I have no idea what to say to Hey Google.
Also I know this is true in other domains as well, obviously there is a common "google-ese" that people learn to narrow down their searches.
In the spirit of the title of this post, someone else also has to say something.
If your argument is that this is a "non-visual command line" there's slim hope of the layperson learning a whole secret grammar without even a goddamn man page just to do their menial tasks.
This got me thinking. Voice recognition is basically a commodity now .. there are open source AI engines that can do it offline really well. So the recognition part is solved, you can just grab it from your distro's package manager. Now there's just the language part.
Thing is, I don't want to speak to my computer using English. Aside from the enormous practical problems in natural language processing you've outlined, I just find the idea creepy[1].
What I want is to unambiguously tell it to do arbitrary things. I.e. use it as an actual computer, not a toy that can do a few tricks. I.e. actually program it. In some kind of Turing complete shell language that is optimized for being spoken aloud. You would speak words into the open source voice recognizer, it writes those to stdout, then an interpreter reads from stdin and executes the instructions.
Is there any language like this? What should it look like?
And yeah that would take effort to learn to use it right, just like any other programming language; so be it. This would be a hobbyist thing.
[1] https://i.kym-cdn.com/photos/images/original/002/054/961/748...
If you're using an averaged American voice - maybe. But it's really not solved for everyone. Google assistant can't set the right timer for me 1/10 times. And that's before we get to heavy accent Scots and others.
There are quite a few hobbyists working on local on-prem privacy focused voice assistans with conversation support.
https://www.home-assistant.io/integrations/#voice https://www.home-assistant.io/integrations/conversation/
Have fun. It is a rabbit hole.
I personally don't consider this a fully-solved problem. The best transcription system I've used is OpenAI Whisper, and it doesn't work in realtime. Maybe it's fine on small amounts but it's still not perfect. You really need error to be driven down dramatically. Zoom auto-captions are a joke in terms of how badly they work for me, and Live Text (beta) on macOS is equally dreadful. YouTube auto-captions suck. All of these use industry-leading APIs. If I'm speaking a voice command and one single word is wrong, usually the whole thing fails.
There's an entirely separate issue about things that are Proper Nouns that don't exist. For example, "Todoist" is often misunderstood by Siri. Thus, people started saying "Two doist (where doist rhymes with joist)" to fool it into understanding "Todoist". Media like anime with strange titles from other languages often flat out trolls these transcription systems. ("Hey Siri, remind me to watch Kimetsu no Yaiba tomorrow".)
You knew that you were drawing something designed for a computer to recognise as unambiguously as possible, while being efficient to draw quickly and easy to learn for you. I feel like that's the kind of notion that voice interfaces should somehow expand upon.
This is potentially far from true, depending on how exactly you draw the line between "voice recognition" and "language". I've looked at quite a few transcription services, and they fail a lot of the time for most people - those who either have a non-native accent (even if very slight!) or those who do any amount of stammering or other vocal tics.
Manual transcription:
> So no: long story short, Slum is basically the way we can have an individual [, uhhh,] instance that carries all the licenses.
(Slum is a project name in this case)
Computer transcription (MS Teams):
> So no.
> A long story shorts. Love is basically the way we can have an individual.
> OHS instance that carries all the license.
Natural language is a fundamentally wrong vehicle to convey information to a computer. It can be useful for some specific tasks, automated Q/A, simple interfaces to databases, stuff where I can't be properly f_ed to remember the syntax or the shortcut like IDE commands.
But the idea it can replace formal language is fundamentally and dangerously incorrect. I agree with Dijkstra's quip, we shouldn't regard formal language as a burden, but rather as a privilege.
What the hell? Is riding public transport or riding a bike either a burdain or a privilidge? Is Driving a car?
I am trying to control shit in my home, it should be neither.
1: https://www.cs.utexas.edu/users/EWD/transcriptions/EWD06xx/E...
I don't see why a voice assistant for the masses couldn't "train it's own users", for example suggesting the language it does expect. But even then, most times people are talking in noisy environments or talk to fast or don't have an understand of how the machine might work. Regardless, who cares. They ruin the audio environment of a home. They're good for setting timers while you're cooking, that's about it.
So maybe it's just that the subfield of natural language understanding is still too early to be really useful. Speech recognition itself has gotten really good but then understanding the context, the intent, etc, all that is natural language understanding, and that is often the problem.
Citation needed, there's a lot of disagreements and misunderstandings (some have cost lives) that could've been avoided if we didn't have 10 different ways to say the same vague thing that can be interpreted in 20 ways. You think the military uses a phonetic alphabet and specifically structured communications for fun? Or the way planes talk to ATC for example. Where precision and unambiguity is crucial, natural language always gets ditched for something more formal.
Around when Android appeared, and the first voice searches began, Google suddenly started to alias everything.
Search for 'Andy', 'Andrew' appears. Search for 'there', and 'they're' appears.
This has been taken further, now silly aliases such as debian .. ubuntu exist, and as google happily drops words in your search, to find a match, this makes precision impossible.
But, that's the only way to make voice search remotely work, so...
If you have terms you don't want interpreted broadly you can put them in quotes.
I preached the Gospel of Google when the competition was composed of web rings and Altavista, but Google in its infinite wisdom has abandoned the advanced user with changes of this nature.
https://blog.google/products/search/how-were-improving-searc...
Maybe we don't have clear intentions in the first place, maybe languages are not just ambiguous, but only meant to narrow realms of valid interpretations down to a desired precision, rather than intended to form a logically fully constrained statements. Maybe this is why intelligent entities are needed to "correctly" interpret natural language statements, because an act of interpretation itself is a decision making and an action.
Just my thoughts but I do think there are more to be said than "natural languages are ambiguous".
I only use voice assistants to set alarms. I cannot imagine voice as a primary input. Then again, many have opted out of owning desktops and laptops in favor of mobile phones. That also seems terribly inefficient.
A lot of people don't need computers in the general purpose sense. I admit my mind boggles a bit when co-workers tell me their kids don't want a computer to do their school papers because their phone is fine. But, then, I'm used to keyboards and what we think of as a "computer" and have been using one for decades--and grab one when I can for any remotely complex or input-heavy task.
https://www.theverge.com/2020/4/20/21227741/apple-ipad-pro-m...
Suggests a keyboard and large tablet is heavier than a laptop
I actually do wish there were good Mac or Chromebook choices for a travel 11" or so laptop but the market seems to have settled on a thin 13" as the floor and, admittedly, the weight/size difference isn't huge.
In response to a grandparent comment about weight for tablets: I prefer Apple’s folio old style of cases/keyboards because of weight. I have one for both my small and large iPad Pros. Whenever I travel, I usually just take one of my iPads if I don’t need a dev environment [1].
[1] but with GitHub Codespaces and Google Colab, development on an iPad is sort of OK.
Might as well go for the laptop at that point given that it can actually do far more imo, unless you ditch the phone and go for one of those half phone half tablets I guess.
For reading, I'm probably bringing my Kindle along if I don't bring my tablet.
Back to work? Sit on table, one cable and it's back to a desktop and charging up again.
Makes the whole thing make far more sense.
The killer-tech will be when we have a tablet that is as light as phone.
I grew up in the 1980s, when handwritten papers were still the norm. I do see the advantages of using a word-processor for writing papers, but don't see why it would be a necessity (at least, until University).
On the other hand, legalese exists and is the lingua franca of telling people what to do, and math exists.
And that's why all of aviation has moved to a tight phraseology, such that delegated commands are universally understood and their meaning is set in stone.
Natural language has cost many lives.
Using language to instruct humans goes wrong all the time. Just a short while ago on British Bakeoff I saw 2 of the contestants make white chocolate feathering on their biscuits by making actual feathers out of white chocolate and placing them on their biscuits. And I'm sure that will confuse quite a few people reading this too. It certainly confuses image searches. Language is a fuzzy interface. Compare to interface like clicking on a button that does the thing I want done.
Not always resulting in unambiguous instructions:
"Lord Raglan wishes the cavalry to advance rapidly to the front, follow the enemy, and try to prevent the enemy carrying away the guns." ~Lord Raglan, Balaclava
"I wish him to take Cemetery Hill if practicable." ~Robert E. Lee, Gettysburg
Every time we try to minimize errors, we formalize a language. I don't even think people use natural language to issue commands often. Commanding people is often considered rude.
I think this is really a characterization. Mostly human communication is full of errors and problems.
What is true is that when it is important enough, humans have come up with ways that minimize communication errors and frameworks to deal with ambiguity - mostly these involve training and effort though, it really doesn't come naturally.
- identify users by voice
- ask them clarifying questions
- remember the answers on a per-user basis
- understand "no, that was the wrong answer"
If you're going to provide a formal interface to the computer, you also have to provide teaching in that formal interface, which is far more of a burden to the user than the cost of the device. And we've completely moved away from that model (not necessarily a good thing, but that's what the market has chosen).
But I imagine there are significantly more who would appreciate clarifying requests by a teachable assistant capable of interacting with the entire digital world on their behalf, efficiently and intelligently.
I've said before, I would prefer a voice assistant that optimized for traversing its menu system, in response to unambiguous noises (could be high and low pitch hums or whatever) that lets me bypass the guessing game and use the menu it's hiding
That would be a fun universe.
[0] https://mw.lojban.org/index.php?title=Lojban&setlang=en-US
[1] https://en.wikipedia.org/wiki/Universal_Networking_Language
Hey Siri
Turn lights on 50 percent
For one hour
Dim over that time
Play music.
I can learn what I need to do; JUST LET ME KNOW THE MAGIC WORDS!
...I don't know how to do that
Alexa, turn lights on
...What do I turn the lights with?
Alexa, activate lights
...I don't know what you mean
...It is pitch black. You are likely to be eaten by a grue.
ALEXA TURN ON THE DAMN LIGHTS
...I don't know the word "lights"
...Oh no! You have walked into the slavering fangs of a grue!
** You have died **
Downstairs or upstairs bathroom?
Downstairs.
Sorry, I didn’t understand. Downstairs it upstairs bathroom?
Downstairs bathroom.
Sorry, I didn’t understand. Downstairs it upstairs bathroom?
Cancel.
Ok. Cancelling.
Siri turn on downstairs bathroom lights.
(Turns off all lights)
Because god help you if your device names are similar to your room names…
And I added scenes so I can say “Gondor calls for aid” and the beacons will light.
"hey siri?"
(no response, no icon),
"hey siri?"
(no response, no icon),
"hey siri?" (louder)
(no response, no icon),
"hey siri?" (louder and slower)
(no response, no icon),
reboot iphone 13 pro
"hey siri?"
works
I guess a list like language would be ideal and the pauses would be like parentheses
you great the prompt and add one or more actions to take.
Otherwise, it works great :-) We love the hands-off usage mode because we cook a lot, so adding things to shopping lists or looking stuff up doesn't require cleaning hands in the middle of prep. Also the speakers are pretty darn good for the size and work well for music.
Doing complicated things is right out though. But the simple stuff works fine.
It would be more error prone in a lot of ways, which is probably why nobody's done it yet, but it would also be a _lot_ more powerful, and fulfill the vision of conversational AI in a way the current rules-based assistants do not.
I think if powerful language models were easily accessible to normal people (in an inexpensive and completely unrestricted fashion, like with Stable Diffusion) we'd already see this happening in the open source world. Companies are going to be a lot more hesitant to try it though until they have a way to 100% prevent the models from making mistakes that could reflect poorly on the company, which is going to take _way_ longer to achieve.
It's used because companies can cheap out on buying a license for other communication applications, it is fundamentally worse than anything else in any other metric. If voice lets me respond to a message without hunting for the hidden reply because Teams shoves it below the bottom of the screen then it could be a win. Considering UX is so low for Teams I doubt it will.
But how? Even if those interfaces were actually working, it's still extremely inconvenient to talk when you can click. You have to be somewhere where talking out loud doesn't disturb the people around you. That excludes most situations: open space offices, restaurants, coffee shops, public transport, cars with passengers, and most places in the home except maybe the bathroom.
And even if you're all alone in a silent place, giving instructions out loud takes more time than configuring a screen, and will always be error prone, because the feedback will always be ambiguous and imprecise.
Except maybe if the feedback is on a screen, but then if there's already a screen, why not use it.
Then it would be called AirPods.
* Current time, and weather forecast for the day
* Upcoming meetings today
* My current commute time to work, including traffic
* NPR news podcast
So during my routine of letting the dogs out, starting the coffee, etc. in the morning, I get the daily "essential" info.
Smart home light/etc while hands are occupied like with a baby. But usecases are quite limited
I would separate out the two, actually. There's a "natural language control system for the entire OS" and then there's the actual voice part. Voice is often mostly useful for accessibility purposes -- hands full, running, driving, etc. However, the other side is that a text-based NL assistant would also be profoundly useful. On iOS, you can enable "Type-to-siri" and you can just type sentences and Siri will respond back in text.
If we make progress on NL-driven command-lines, we can actually make progress on voice-assistants, and vice versa. The catch is that the voice side still needs recognition work.
The would be trivial to use the interface on screen when appropriate, and a truly smart assistant should be able to follow the context and be aware of your preferences and mood.
This is not fundamentally impossible, we're simply not there yet.
Working from home changes that. I can see many more opportunities for a multimodal input interface. Examples:
1. My fingertips now are closer to the "reply" button below this text area than they are even to the touchpad. Touching "reply" is half a second, moving one hand to the touchpad, aiming the pointer at the button and clicking takes longer. With a mouse: much longer. Anyway, my screen is not a touchscreen. I'll click.
2. Or, with an assistant, I could have said "Click reply", provided that the assistant knows where the focus is and that it can read the form I'm typing in.
When they were a novelty I recall the excitement of trying new commands and layering in context, after many failures I've been conditioned to now only attempt and expect success with generic queries.
My biggest frustration with Alexa is getting it play the podcasts I want to listen to. Even popular podcasts with English names are hard to get just right for Alexa. The same goes for song titles and bands that are not popular, or they are in other languages.
Usually when I want to take a shower, I try to get the podcasts/music to play for 2 minutes, then sigh, give up and just say "Alexa play Britney Spears".
This is not power. This is just first-world problems.
Granted, this is for a specific user base and yes, not in coffee shops.
This kind of thing can't be built for modern mainstream operating systems because they generally prevent subjugation of the OS components and other programs, even if the user wants that, ostensibly for security reasons.
Unlike a human operator, an assistant "app" can only operate within the bounds of APIs defined by the OS vendors and third-party developers. Gone are the days of third-party software that extends the operating system in ways that the overlords couldn't (or wouldn't) dream of.
I think the problem with that is that even I, as a human, struggle to know for sure what you want.
You want to turn all the lights on in the house? Does that include the lamps in the bedroom? How about new lights that you add later? Or the ones in the garden? It's full of ambiguity. What device do you want to watch the news on? Or did you mean the radio? Do you want this to apply when you get back at 2am one night, meaning your family gets woken up when you turn on all the lights and start playing the news in their bedrooms?
I think that's probably why voice interfaces aren't likely to work well for anything beyond direct, specific, well-scoped requests: turn on the lights in the bedroom; turn off the heating at home; roll up the blinds; what's the weather like today; what's the remaining range on my car. They really struggle to deal with anything more complex – not so bad in theory, but really incredibly irritating when they make the wrong decision.
If you had some kind of 24-hour live-in assistant (a butler, maybe?), then they probably have the knowledge and intuition to make sensible decisions in response to fairly unstructured requests. But I think we're miles off getting a voice assistant to do it – not because they can't, necessarily, but because if they mess it up at all it's infuriating.
What about the shortcut for when I need to leave at 3am for some reason. then a different shortcut for when it isn't just me, but my whole family leaving at 3am. An still another for my son having to leave that early.
Jeeves can figure it out when I arrive at 2am so I don't need to program it.
Voice can do way more than we know, but we have no idea what it does or how to use it.
Standardizing the interface and providing tutorials would possibly change things dramatically.
And this goes for the back-end protocols as well.
The tech is way, way ahead of the UI and integration.
Imagine getting the power of 'git' with no tutorial and not really an understanding of what it does? Good luck with that.
90% of us would be using it in the car to do a lot of things if we really knew how to do it:
You: "Siri: Command. Open. Mail. Prompt. Recipients starting with S"
Siri: "Sarah, Sue, Sundar"
You: "Stop. Command. Message. To: Sunar. Thanks for the note. Stop. Send without Review"
Some of this already exists, but it's product specific etc. there needs to be some kind of natural universal interface - or we have to wait until the AI is really, really that good.
I don't really want them to be all that much more powerful, because natural language can be imprecise, and... there's just not much I that I want to automate in a home setting beyond some real simple timers for lights and stuff.
What if I had a bad day and didn't want to see depressing news? Or what if I came home and was talking on the phone when it turned the news on?
True automation as opposed to just telemetry and remote control can easily be annoying more than helpful.
I like the idea of automation... but I don't actually... automate anything aside from timers and reminders.
The problem is that you have many many billions of dollars have been sunk into making these devices about more than setting alarms and timers. There's actually been a lot of pretty amazing progress. But it's yet another one of those things that getting to 90% to anyone but techies who want to fiddle with their smarthome stuff or otherwise play with the technology.
Maybe they'll add an app that lets you browse possible commands so it's more discoverable.
But I'd observe that I'm going up to my brother's tomorrow and he has all manner of timers and other WiFi-connected stuff and none of it has any sort of centralized control and that's pretty normal even for people who have a lot of that sort of thing.
And, yeah, the only smart light thing I have at home is one thing that doesn't have a controlling light switch and I used X10 for it for years before I got an Alexa.
And even then, a voice assistant is essentially a user interface, not a product or service.
It could be a service if you could reliably say "Alexa, plan my trip to customer X the week of the 30th and send me my itinerary". But for now they are an alternative to a phone UI.
This seems a really high bar for voice assistants aspiring to do much more than set alarms or turn the odd light etc. on or off.
These days a large part of what people relied on secretaries for a computer can do faster, so only at the highest levels do you see them. There are still secretaries at the low levels, but not nearly as many, and they are not doing the same tasks.
And, yeah, assistants shared with a bunch of other people--as with travel agents in general--aren't really all that useful. If I'm mostly just giving fairly mechanical instructions to execute, it's probably easier for me to go online and figure out the options myself.
A secretary made a lot more sense when you dictated memos for inter-office mail and retrieving information often involved making multiple phone calls.
I know they know this about me not only because of my Gmail account but also because I use Google flights to find the flights before I book them.
Unfortunately they're not using this data to help me. Rather they're using it to target advertising to me. But they definitely have the data and the machinery to be more useful to me with more than just a few facts
It gets complicated in a hurry and for the cases where it is relatively simple (and when it gets into very complex international travel a voice interface is going to be completely useless), I can look up my options pretty quickly on a computer.
That's kind of my point. A voice assistant is just a fancy UI until they reach the level of AGI, and I don't see the point in spending billions of dollars on them to be a simple UI as Amazon seems to be doing.
Voice alone sucks, it's just too limited to be useful on a grand scale. Similarly, command lines suck too. The shell in general has the same problems that Voice assistants have, just that they have more value and had decades to mature into something actually useful. And toady we have unix-shells which reduce the problematic parts by many levels, and still receive constant improvements. This is missing for voice assistants, because unix-shells are growing and improving in an open space, where everyone can add their own things. This is not happening in big tech.
Edit: Maybe things like this happen because there are various nerds who lead these products and are good at talking the businesspeople into funding it. Maybe this was only possible at the big tech growth stage while business wasn't that good at telling the value proposition. So end result, lots more engineers get paid which is great in my book :-)
At the same time, siri seems to be getting slower and fatter every iteration so perhaps it is becoming more human ;)
“OK, I’ve created an infinite number of alarms, every five minutes, starting at 6 AM tomorrow!”
(As a native English speaker, I'm not sure what specific outcome you want to happen from that request. That's the one that makes the most sense.)
And you now have me wondering how open-ended calendar requests are actually implemented given that they can't literally have entries out to infinity. (I assume they go out some finite period and some background process periodically re-populates future entries.)
If these voice assistants were smarter about “alternative” names for every device it might be easier to use. But as it stands, it’s kind of a pain because the way you phrase each request is so unforgiving…
Oh yeah, and god help you if your device name is similar to your room name. If your room is “office” (or did I name it “the office”?) and your light is “office light” Alexa is gonna have a bad time figuring the two apart.
I have no clue how to fix this…
PS: this is why I question steering wheel free self driving cars. How will we tell these things exactly where to go when we cannot even reliably tell our voice assistants exactly what light to turn on?
I work at SoundHound where we've been worried about these issues. (I'm going to plug our recent work...) Our new approach is to do natural language understanding in real-time instead of at the utterance (turn) taking level. That way we can give the user constant feedback in real-time. In the case of a screen that means the user sees right away that they are understood, and if not, a better hint of what went wrong. For example a likely mistake is an ASR mistranscription for a word or two.
We still need to prove this is a better paradigm for VoiceAI in products that people can try for themselves, and are working towards that goal. I hope that voice interfaces that were clunky with turn-taking will finally be more naturally usable with real-time NLU.
A 'debug' area that lets you ask a command, see what was interpreted - and immediately edit or click "that's not what I wanted". But not an afterthought and not a cumbersome process like setting up an automation that is triggered by specific commands.
Imagine telling your voice assistant "You're wrong, as usual" and instead of it giving you the boiler plate "I'm sorry ", it actually offered a way to improve itself.
That sounds like a security nightmare. Someone walks by and starts changing your system settings? No thank you
In practical reality these interfaces feel, to me, as extremely inefficient. As someone who doesn't particularly like to speak, and prefers silent environments, these interfaces require more energy from me to use. Unless they are serving someone who has a physical impairment then I don't see what problems exist that these solve, but I can identify lots of problems that they introduce (not only noise but privacy / security vulnerabilities etc.)
Personal preference.
However Google's Assistant in comparison worked great, no memorization, and very useful. Sure time, weather, set timers, and alarms worked great with a very flexible set of natural language queries. Even more complex things like what will be the temperature tomorrow at 10pm, simple calculations and unit conversions. But also things like IMDB like queries about directors, actors, which movies someone was in, etc generally worked well. It seemed to really understand things, not just "A web search returned ...". Even more complex things like the wheelbase of a 2004 WRX would return an answer, not a search result.
With all that said I'm looking for a non-cloud/on site solution, even if it requires more work, most recently noticed https://github.com/rhasspy/rhasspy
There's a great little game series called Megaman Battle Network (Rockman.exe in Japan) which diverges from the mainline by showing an alternate universe where scientists focused on AI instead of robotics, resulting in a world where "Navis" are ubiquitous.
I wonder, what if our early software engineers focused on bringing natural voice control to CLIs, before perfecting GUIs first?
On top of that the general distrust of the privacy of these systems has stoped a significant number of people (myself included) from wanting to us them at all. I don't have an in home device, and have turned off Siri on my Apple devices.
I think if the accuracy was better and more content/things were available through voice it would be a pretty good input method for any scenario where you don't need visual feedback.
Browsing content, or looking it up, via voice is slow and playful, it always will be. Who what's to be saying "next", "scroll down", or have a full on conversation with an AI to try and work out what you want to play? Our fingers on our hands have evolved to be incredible at interacting with things, we are good at using them. Touch screens or physical UIs will never be superseded by voice.
So yes, there is a small use case for voice for controlling music/tv, or controlling a few things in the home (heating, blinds, lights) but thats it, I don't believe there is this massive opportunity to expand it into our everyday lives where we are constantly interacting with devices via voice.
Humans evolved language to communicate ideas, wants and desires to others for thousands of years. Obviously voice UI is not there now but maybe someday the experience won't be much different than asking the movie rental store clerk for their recommendations for a romantic comedy.
The only people who ever did that were in a romantic comedy.
I bet ya some engineer at Amazon hooked that up manually when they saw a bunch of requests failing, so that’s only gonna work for popular fuzzy naming conventions. I don’t want to have to think “is this a way lots of people are gonna request this song?” before saying it that way.
The more relevant use case is "hmm, I'd like to listen to some prog rock, let me browse what Spotify has and see what takes my fancy". Sure, I could say "VA, play prog rock", but I don't want it to choose for me: I want to browse the available content to remind myself what are my options, and choose one when I see one that looks interesting.
That's why an opensource ToS-violating assistant has chances to work better than legal ones, they can just scrape all those infos off internet. But then, once you go into that grey area, you just end up pirating content already.
my observation of people on the road has led me to conclude that Driving is an activity where people think they can do absolutely anything else while engaged in it.
1. sending messages on phone while driving, one hand on steering wheel.
2. having sex / receiving oral sex.
3. turned around, yelling at kid in back seat to not fight with other kid in back seat.
4. girlfriend having argument with boyfriend, slapping him on arm some, about how she was smart too just a different kind of smart while swerving back and forth in fast merging traffic near the Haight (I was in back seat)
5. it's getting hot in here, time to take my jacket off!
your mileage may vary of course.
BTW, another use case for speech recognition is when you're carrying a baby around.
Since the voice assistants are incredibly stupid I find it extremely stressful and distracting to ask them for anything while driving.
Saying "Hey Siri, text Fred <pause> I'm on my way but stuck in traffic, eta 4 o'clock" or something along those lines nearly always works fine for me and is no more distracting than having a conversation with somebody in the car with me. If Siri gets some of the message wrong I'll either send a new one using clearer speech or wait until I'm not driving to fix it if the mistake isn't important.
Sure, it would be possible to then allow myself to get distracted by focussing too much on some weird aspect of it, but equally it would be possible to get so emotional in a conversation with somebody sat next to you that you stop paying attention to the road. And we (most people at least) don't say "it's not safe to talk at all while driving", we just make sure not to go over that line of getting too distracted by the conversation.
Until you have several Freds in your contact list. Until you have friends with foreign/uncommon names. As long as you have near-perfect American pronunciation. As long as...
There are too many variables to consider and think of. Sometimes I can't get Siri to reliably understand what music I want (and my Engilsh is pretty darn good), much less anything more advanced.
There isn’t a such thing as an American accent. Ask anyone who is not a native speaker and either hasn’t been to US that long and tries to understand my natural deep southern accent. I can adjust my accent if needed and if I think about it.
- General American https://www.babbel.com/en/magazine/united-states-of-accents-...
- California English https://www.babbel.com/en/magazine/the-united-states-of-acce...
That only works well if you have an accent it recognizes, if you're speech is clear (not slurred, not lisping etc), if you don't stammer, if you don't have any verbal tics that you don't want to show up in the message, and if "Fred" is actually a simple unambigous name.
Otherwise, at best when you want to send a message to "Ioana" it may end up sending a message to "Anna" that says "I'm, ummm, oh my way! and stalking traffic ate a what was it like 4 like maybe 4 and you know what <pause>" (followed by the "4 o'clock" that will no longer be included).
The Google ones do support named timers, so you can say “start a pasta timer” and later ask “what’s left on my pasta timer” etc. I thought Siri added this at some point, but I wouldn’t be surprised if not.
Anyway, it does seem to have improved, but I wonder why that stuff wasn't in from day one. It seems pretty obvious to me.
If I say "set a timer for 20 minutes called A" it just ignores the "called A" part.
A HomePod can set multiple timers. A Watch, iPhone, or iPad can only set one timer. There is no obvious technical reason for it. It just seems like only the HomePod team thought it was an important feature.
This becomes annoying if you have multiple devices set to respond to “hey Siri” and the wrong one picks up the request and then refuses to comply.
No clear feedback, a weird timing issues where it just stalls and show the message it' about to send in case it got it wrong.
It's just a terrible UX all around.
Siri never triggers when I'm driving, it just doesn't hear me. I think it's because of the noise of the car or because of my music, but it doesn't work. I have to move my face closer to my phone so that it can hear me, but that's even more dangerous than using the controls.
Same when I'm in the shower and I ask it to change the music, it doesn't hear me, I have to shout and get angry every time.
For what it's worth, it doesn't work either when it's my pocket. When I come home and ask it to turn the lights on, it doesn't answer if it's not in my hand.
See here: https://support.apple.com/en-gb/guide/iphone/iphaff1d606/ios...
this is still potentially a huge domain. one could imagine a benign scenario where voice assistants enhance people's abilities to interact with each other (and digital devices) when a more potent UI is not within reach
privacy concerns (->controversial business models) and technical ability to deliver a desirable service (that people would pay for) might indeed prevent this vision from catching on in the short term
another factor that may complicate adoption might be just cultural / perceptions. It is a somewhat odd thing to be shouting at devices - especially in the presence of other people. User interfaces that interfere strongly with communication habits and behaviors established over millennia (see also wearing VR goggles) might have a harder time seeing adoption outside very specific scenarios
While driving, I wanted to have Siri read a lengthy webpage to me. I pulled up the page, got in the car and asked Siri to "speak screen." Siri says it can't do that when I am driving! What idiot thought that was a necessary safety measure? What if I were the passenger?
Overall, I am stunned at how bad Siri is at things that don't even require AI. It's almost as if this insanely profitable company failed to invest a tiny bit of money into researching ways that people would like to use Siri.
I often go places with my sister (she drives). Her car doesn't allow pairing or swapping bluetooth connections to the car's entertainment system while it's moving. If we want to switch to my phone we have to come to a complete stop.
My most egregious example of this for me is that there's a grocery store near me that the Google assistant is incapable of finding because of a few people in my contacts list. Whenever I try to ask it for directions to that store, it picks (at pretty much random) one of three of my contacts instead. This is despite the only common part of said contacts' names and the grocery store is that their names all start with the same letter.
Basically, imagine asking for directions to Albertsons, and the assistant giving you directions to Andrew.
1. The command set is broad enough or user input is complex enough to make other UIs inefficient. 2. The voice UI is up to the task of correctly interpreting the voice input correctly most of the time
What "most of the time" means for the second item is somewhat personal and use-case specific.
For item 1, examples where voice is better right now, or could be with reasonable NLU improvements:
"Text my wife that I'll be there in five minutes."
"Get me driving directions to the nearest Indian restaurant with at least 3 stars on Yelp."
"Order six rolls of paper towels and a bottle of Windex from Walmart, delivered to my home address, for delivery by Saturday"
"Remind me tomorrow morning to review this web page"
"Create a shopping list with the items from this recipe"
"Create a basic presentation with one slide each for each entry in the table of contents for this book"
Voice can be better. As others have pointed out, as long as it's like playing Zork where half the time the response is "you can't do that" or "I don't understand", voice interfaces will continue to flounder.
If you tell it you want more coffee, it should know what you like and suggest a mixture of brands you bought before and new ones you may enjoy. If you tell it you're hungry, depending on the time of day it could suggest you some takeaway you've ordered previously or something else you may like. If you say the same some other time it may suggest recipes based on what you have at home or it may suggest nearby restaurants. It should keep track of your friends and otherwise and tell you when their birthdays are coming and it would be nice if it could even suggest some presents based on things you've told the assistant before, or their wishlist on amazon or something else.
There are a lot of things assistants could do, but it needs to know you. The model where everyone has the same assistant doesn't quite work out.
Instead, you could go easy, by first suggestint the user to set up your assistant by linking it to amazon, deliveroo/uber eats/etc, facebook, and verbally sharing information as it asks. Then, over time, it could spread its inference further and further and you won't be sure or not if you already shared that information or not, but you will just assume you did, as you share everything anyway. For people who are not so open, it could stick to inferring less and being less useful.
The issue is of course, like paying for youtube, is consumers will wonder why they would want to pay for an assistant.
This is a complicated social construct even with people. If a public relations person (or whoever in a professional context) reaches out to me, I hope they've done some basic research on what my interests are. I saw you were at $CONFERENCE last month? Sure, probably. Start asking me about my vacation last month that they found photos from on Flickr? Probably getting over a line if I don't know them.
I imagine that a truly top-notch virtual assistant would always be listening and aware of your behavior and context. However, the level of trust required for such monitoring is usually reserved only for one or two people in our lives, and even then, it's quite incomplete awareness on their part.
I don't know how a for-profit company can reconcile this disconnect, though I imagine that someone will eventually try.
"Hey Siri, add more toilet paper to the shopping list" (while pooping)
"Hey Siri, shuffle my music" (while driving)
"Hey Siri, countdown 10 minutes" (while shoving a pizza in the oven)
Anything else is a shit show. Anything where trust or accuracy is involved i.e. mutating data, spending money, absolutely no way can I trust it at all and never will.
This is the main reason why I have an Echo in my bathroom! The one advantage Alexa has over everything else is that you can voice shop -- "alexa buy more toilet paper" solves the problem that much faster than a reminder for later.
The reason Alexa exists is to sell you Amazon's prices, not necessarily a good deal.
And it's not even a very frequent thing. Mostly, every few months, I look through what consumables need replenishing and I fill up the car with plus-size packages from Walmart.
"Hey Google, put dishwasher salt on the shopping list" "OK, I added 'put dishwasher salt'" (strangely, this particular bug only manifests for dishwasher salt).
Timers are useful, but sometimes they can't be shut off by voice command.
Shuffling music, turning lights on, yes fine - because confirmation that the right thing has happened is instant and effortless. Anything else, I'll use a button or a screen.
Confirmation is required when dealing with humans as well ... https://www.youtube.com/watch?v=11fCIGcCa9c (this reminds me of Alexa)
Then I timed myself setting a timer on my phone, which took 9 seconds from pocket to running.
Adding to a shopping list isn't clicking the "buy" button, no - but if it's not on the list I won't buy it and then I will have no toilet paper. I would not need a list if I could simply remember everything.
It’s not perfect though, for example when trying to add fruit and fibre cereal it will often add two items, “fruit” and “fibre”. But its close enough that when I get to the store and check the list I know what I intended to add to the list.
Are you saying this for comedic effect, or does the Alexa really do this? (I'd look it up myself, but good luck with that query...) To each their own, but I'd throw the device into the street if it pulled a stunt like that.
Then I timed myself setting a timer on my phone, which took 9 seconds from pocket to running.
To the Homepod or my Apple Watch: "hey, siri, tea timer for three minutes".
"Three minute tea timer, starting now."
I didn't think a product could screw that up. I would suppose it's a design decision between "assistant" and "servant that carries out my command without backtalk". There are times that I wish the Apple product were more "assistant" than "servant", but the Alexa product just sounds pushy.
The other day, we had to remember to book a school thing for the kid. I said "Hey Google, set a reminder for 9pm to [book the thing]".
Google replied "Here are web search results for set a reminder...."
When they fail constantly at the most basic tasks, usage is going to drop way off.
Meanwhile my success rate for the way I speak is below 20%.
just compile a list of all commands into an email and send that to me. i don't actually want to hear you talk, siri
"Hey Google, turn off the den light switch in 30 minutes"
"Sorry, for safety reasons we cannot...."
It's a light. Because it heard "switch" it thinks there might be some power tool connected to it and won't let me set delayed actions. I want to be intelligent with it like "Hey Google, turn turn the lights on when the sun comes up everday" but no one has gone to that next step.
Or how about "Hey Google, turn off the tone played when you say Hey Google". These settings aren't accessible from the voice interface itself.
Can't wait for Alexa to fail so my SmartTV will stop nagging me to use integrations I will never use. Anti-competitive but whatevs.
really makes me think they almost want to discontinue it but it's so integrated with android and the chromecast they can't really kill it
I hate how many products need to be configured from apps nowadays. You gave me the voice assistant and I can't use the voice assistant to manage the voice assistant. I shouldn't even need an iPad or phone to set it up. Give me a minimum local processing/interpretation so I can get it online and it can begin interpreting more articulate requests.
We are so far away from zen.
And, they don't really explain the syntax constraints. Which are massive.
Try ab initio without knowing how to do it, to get OK google to open an arbitrary google authored app and direct it to do something. Compared to learning how to use the OS UI keyboard shortcuts or applescript. (Which btw like Windows is basically fully documented because all the libraries are self documenting for their call structures)
The voice interfaces are universally badly designed because spoken command sentences are not well understood as a modality of command, distinct from mouse, gesture, touch or keyboard.
Until voice is baked in with a documented syntax in "man" format, i won't believe its first class.
How do I even know for any arbitrary app what voice directives it uses? How do they correlate to any other command input? How consistent is this with other commands in other apps? Does "stop now" always mean the same thing between a mapping routing app, and a tape backup app? Isn't "stop" contextually defined in a way ^C isn't?
I have a room called "Study" and the only lights there are called "Main Lights".
- Hey, Siri, study lights 100%
- Did you mean these lights: "Study: Main Lights"
:facepalm:
If it was the only "main light" across all rooms I'd hesitate to say the same mind you, nesting scopes should be respected.
What would you want it to do when you add a second light bulb? Tell you the old command is no longer unique or turn on both?
I think when the command is generic "lights" it should turn on all lights. But both are valid behaviours.
I’ve learned that I can’t say “what’s the temperature” because it will tell me the temperature my Dyson fan is set to lol. It’s not wrong, it’s just missing some context (me holding my jacket wondering if I should wear it) that I can provide going forward. Maybe I should just ask it “should I wear my jacket today”
The hard part isn't naming the stuff, its remembering what each thing is named. I could call the lamp next to the couch several different things depending on the context. If you don't remember exactly what it is named you are gonna have a rough time. I've been tempted to just put a label with the device name on everything controlled by Alexa.
And good luck getting a guest to turn anything on or off.
Or just call it dude :) https://reddit.com/r/tumblr/comments/548b3p/everything_is_du...
I'll stick to Google assistant for the extra occasionally useful stuff, but the idea of a device that won't stop working if the server goes away is pretty cool.
Do your friends have accents? Alexa frequently fucks up due to my South African/Zimbabwean accent. I joke that Alexa is racist. Google seems to handle it way better. Here's hoping that Mycroft does better (which I plan to switch to in the coming months).
Truly opened to developers they could have been really interesting and fun.
This is what happens when big companies develop technologies and think they are too valuable to share.
I can’t do this loop with modern voice assistants.
A voice assistant with contextual conversation skills, and access to an “always on” visual monitor (home projector or AR glasses) would definitely increase utility by 10x or more
Setting a timer.
The oven is in another area of the house, so when I come back to squeeze in some work, I often just said "Voice Assistant, set a timer for 10 minutes!"
And that's about it.
Apart from that I worked on some chat-bots in the past and it's the very same thing to me, just even a bit worse because of audio.
Natural language processing simply isn't there AND there is just a few very niche use cases in my eyes.
So if they go away again, I won't cry or rather my tears will dry fast.
Somebody develops something, gets a lot of buzz, doesn't deliver, buzz dies slowly, and then several years in the future somebody else actually builds the tech stack needed for it, and it takes off again.
Voice assistance, chat robots [1]... the metaverse.
[1] Several years back I attended a chat in my city where some people from IBM showed us how to implement a chatbot I think with Watson.
Beyond the trivial examples, it was nearly impossible to implement anything (Or the people giving the chat didn't know how to implement those) but the documentation was nowhere to be found beyond those trivial examples.
Pardon my getting a bit off topic, but it is interesting to see the “belt tightening” by FANGs: it seems like everyone is cutting excess staff and looking hard at which products make money and which are money losers. This may seem like a good thing except for newer product categories like AR/VR that will need a lot of experimentation to get right.
1. Hey Alexa/Google/Whatever, buy me a saddle, Selle Italia SMP Extra, white. I expect to get a proposal for the best price + shipment and that's it. Easy.
2. Buy me a replacement head for my Parktool pump. Oops, it's a PFP-3 pump. Ah, the shop selling the saddle doesn't have one... I eventually discovered that that shop has an identical part for a pump of a different brand, probably the same Chinese manufacturer. I think this would be hard for an assistant, borderline to impossible.
3. Buy me a magnetic mosquito screen of at least X x Y cm, not adhesive or velcro. Oops, I need a frame to mount it on? I spare you with the discoveries I made along the process. Either the voice assistant is equivalent to a professional installer and can see my door or I'll always do a search with a browser, and watch many videos too.
I use them to get travel time estimates, reference facts, music and podcasts, rewind / forward / pause / resume / next / previous, timers, lights, and occasionally broadcast messages on home speaker devices.
I use voice UI while cooking, running, mowing, raking, driving, changing diapers, putting on my kids' clothes, and while sitting with my wife.
My two young children (5 and almost 2) love voice UI. We ask about the animal of the day, what does this animal sound like, play that song. My older child is beginning to set timers with it. My wife recently said having a nearby voice assistant is important in a home configuration discussion.
My family and I sometimes have frustrating experiences with voice UI, like failed hotwords, shifting syntax, answering on an unexpected device, and slow responses. But we still use it frequently, and our overall sentiment with voice UI is positive.
Because it lacks VUI, my family recently didn't buy a new ~$400 product. We bought an older version that was marked-down and otherwise inferior, but has built-in voice assistant.
can you imagine hiring a production assistant that gave everything you asked them to do to a corporate competitor ?
Why would I want to do that with my life and Amazon and Apple ?
But because of the 2010s meteoric rise of FAANG (in the public imagination and stock market) they think they can quickly push into mainstream all these immature paradigms like IVA, AR and now VR.
All these technologies exist since decades, but they are not ready! billions of dollars of investment and marketing are not enough to make them so.
As small a Leap as a tactile portable device took us almost 20 years to reach a conclusive mainstream form (iPhone)
We need to accept that IVA/AR/VR are exponentially larger leaps and should remain side-shows for a very long time to come.
For example, Microsoft is finally acknowledging this, with HoloLens being now just "to help you solve real business problems".
If I could download it and run it on a private server in the basement, without any ties to the cloud, and with my own settings for privacy, then I'd be more willing to use the things.
Now it's here, it works amazing well (considering what it's doing), and . . . I'm talking to Apple, Google, and Amazon? No thanks . . .
Data, training, pattern recognition, language.
Lost yet another another use to ios.
Though, as much as these services are apparently bleeding, the in-kind payments from big surveillance had really be worth it.
1. Play me music or a podcast (this fails a shockingly high amount)
2. Call person X
3. Covert unit to unit
4. Turn devices on or off
5. Timer set
6. Reminder set
7. What time does Shabbat (my Sabbath) start.
All of the above can be done with my phone, so there's really no point in them - for me.
Siri in particular will wait as long as a full second to “ding” letting you know she’s ready to process a command … often interrupting your attempt to give the command. Alexa will very often fail to hear me say “Alexa” despite the audio conditions being quite favorable.
Both Siri and Alexa will misunderstand what you asked for and be completely indifferent to your attempts to correct them unless you interrupt them correctly. Afaict Siri can’t be interrupted by voice alone and Alexa can only be interrupted by something starting with “Alexa”. Once you are already “in” an interaction session with either of these tools you should be able to simply say “no, not that” or “no, <correction>” and the assistant should stop talking immediately, instantly, and not continue to blather on with the previous wrong interpretation.
No matter how useful these tools are in the hands of a power user, Siri and Alexa give the impression of being morons largely because of unresponsiveness, not lack of capability.
For me, the utility I get out of my echos are home automation control to control heating, lighting, audio when my hands are full or their control are out of reach. But you are never going to make a fortune out of house voice control when you sell the voice assistants at or near cost (don’t think I’ve ever paid full price of any of my Echo’s, and I tell people considering getting on to wait till they go on sale as Amazon always seem to be offering a deal/discount on them every few weeks.
But what does frustrate me if that Amazon don’t seem to uniformly “monetise” things across the Alexa platform, for example when I found out about the Samuel Jackson voice for Alexa I wanted to buy it even if it was just a novelty, but it’s not available in the U.K.
What we were promised: a personal well-trained butler.
What we got: a voice input field that you know will be incorrectly interpreting your intents far too often.
Wouldn't we all like better managers with clearer goals?
Roll this out to a doctors bedside manner, or the mistaken command given to a nurse that was a perhaps misheard and the wrong medicine given
or ...
We can surveil all our waking lives - and almost the only thing social sciences get right is large population statistics- we can find out who leads happiest best lives and be taught how to do it better.
Or we can stomp out individuality and live on a totalitarian nightmare.
But we don't get to not play the game
1) Doesn't always understand what I'm saying (voice recognition)
2) If it does understand (voice recognition), it actually changes it it something else (super dumb, no, idiotic "AI", changes someone I call almost daily into someone I never call)
3) If it asks me what I want to do with a contact, it doesn't understand "call"
4) There's no way to 'correct' what has been said.
5) I can't open "podcasts", because it will open Deezer (I don't use Deezer anymore, but it's still on my phone).
It's really just really really a dumb command line interface. And if it would just be that without the 'smart' things, it would be a lot more helpful. There are only a few things I use it for:
1) Call XXX
2) Alarm at XXX
3) Open google maps. Sometimes I say "route to xxx using google maps", and that works
Even these three common use cases fail about 10-20% of the time.
When I'm in the car and want to change route or add a poi, want to open an app, or text something, I try, but I get so frustrated which is more distracting than just typing it in your phone. So that's what I do.
Google's voice recognition is a lot better and faster. Also their search understands more context.
Voice assistants are like a skinner experiment. Solution: make it very rigid in terms of operations. People are quick to adapt this. Somehow the ai-crowd doesn't understand that a UI works best when the things you operate with are at the same location and always respond and work in the same way.
I'd compare the voice assistant experiment as a GUI where the buttons, their function, and labels always change.
"Alexa turn on my FireTV" (doesn't work, weirdly FireTV and Alexa speakers don't seem to get along)
"Alexa play The Killers" (responds with "Playing the killers on Spotify" then silence)
"Alexa play The Killers on Spotify" (works)
"Alexa play The Killers" (sometimes tells me to start scanning devices and install a skill for Sonos)
"Alexa wikipedia 'history of Bulgaria'" (works but asks me if I want to continue after every sentence, so I give up)
"Alexa stop" (doesn't stop for some apps, Amazon should have support for silencing as a minimum requirement for apps and implement worst-case scenarios if an app doesn't respond to stop in time)
"Alexa rewind 30 seconds" (doesn't rewind for some apps, Amazon should have support for rewinding played recordings as a minimum requirement for apps)
Is the monetization path hard - yes. Trying to break into specific industries takes a ton of work. Additionally - you already have staff in these cases. Would a dentist get rid of her assistant for some voice commands? not likely. A hospital still needs nurses.
It was a neat internet radio device and Bluetooth speaker but outside of making it play the radio I never really used the voice functionality much.
I remember one of the places where I buy groceries from offering an incentive if you ordered your shopping via the Alexa skill, it was so painful to use that I added everything to my basket using the website and then just added 1 unambiguous item using the skill and checked out. They gave me the discount/incentive but I never used it again.
I moved house but never plugged it back in and don't really miss it.
I also remember going to AWS re:invent in 2016 and they gave every attendee an echo dot, with a lot of fanfare to encourage Devs to make skills. I tried to make one but it was so convoluted, I gave up in the end and just gave the Dot to my parents. It broke after about a year, they never replaced it either
The amount of wrong answers or just telling me something it has found on the web, which has nothing today with the topic is just annoying. Even basic questions seem to not work reliably and looking at the whole Alexa ecosystem it doesn't even feel close to "smart" home.
Lack of context awareness, connection to basic things like parcel delivery, washing machine and similar that would make my life a little easier (because I'm to lazy to open an app to look when my parcels will arrive or when the washing machine will finish) just ruin the whole thing for me.
Besides that, I've tried setting the male voice, but for some responses it just switches back to female Alexa??
The problem is a family like us. We haven't bought a new device in 2 years and still use some of the original pucks. Amazon will need to make money somehow so I think a yearly subscription would work. Then again, if it isn't a growth segment I cannot see a tech company keeping something in today's world.
So when I interact with the screen it can show me all my possible choices at the same time at each situation. I click a menu-item and a sub-menu pops up.
With a voice-assistant I would have to wait for the assistant to list all possible choices but if they are many that takes a long time. You can only hear one word at a time. Combine that with inaccuracies of voice-recognition and it is clear that a computer screen + mouse is a much better interface.
Hey it's the same difference as with Radio vs. TV.
A computer screen-interface might be usefully augmented with voice recognition. But working with voice alone is like working blind. Gee who could've thought voice is not the killer-tech of 3rd millennium.
When you tell a mercedes navigate me to "exact road and adress" it barely gets half of it right. and usually that just the number.
Now i don't have a speech impairement and have tried in 3 languages, none of them work to my sophistication.
The worst thing though, personally, is all the shitty patents around voice activation keywords. It's disgusting.
Finally i am also an IOT guy with a lot of toys, and i would have written my own personal assistant if it would not be for the patents. After all the hardware is surprisingly cheap.
All the problems are solved too. but if you believe i would write free code for megacorp amazon without pay, you might be heavily mistaken.
To me where it's gone wrong is that like many other things the A team conceives it but the B team is responsible for its development. The A team will ask "what will users want to say to it"? But Team B says "well if they say 'blah' how do we know that they mean 'blah'"? Mediocrity creeps in.
As an example, here is a typical dialog between me and my Alexa.
"Alexa, play Brahms C Minor Piano Quartet."
"Playing: music like Coldplay."
"NO"
[music like Coldplay]
"Alexa NO"
[music like Coldplay]
"STOP"
[music like Coldplay]
"ALEXA STOP STOP"
I use my Alexa as a timer only, and when it started saying things to me like I have notifications or "did you know you can use Alexa to order such and such" when I would say "Alexa, stop the timer", I almost threw it out the window.
I'd love to repurpose my Alexas into satellite Rhasspy[1] devices if Alexa retires.
1: on HN the other day https://news.ycombinator.com/item?id=33705938
& direct link to the satellite info https://rhasspy.readthedocs.io/en/latest/tutorials/#server-w...
There are obviously still many open research questions. Like how this is then combined with actually performing the commands. But there are also solutions to this.
This is still somewhat open active research. But given such powerful models, I think in principle we could make such devices much more useful.
The issue apparently is no one knows how to make money
I see lots of conjecture on the nature of voice assistants, or why they haven’t taken off.
As someone who was at Amazon, and close to this area, let me offer a simpler explanation.
Alexa has stagnated because its leadership has little to no direction. The incentive structure driving the product teams and tech falls under two categories: 1. rest and vest 2. or write documents to build your promotion portfolio and get promoted. Then leave the org to find a job in AWS.
In fact, Alexa could be used as a text book study in empire building.
* Myriad teams that maintain or increase head count each year, with absolutely no meaningful deliverables.
* Services that could be maintained by a 2-3 person on call that have an entire team of 6-8 engineers.
* Tech directors, Sr. SDMs, and SDMs in a race to build their empire to get promoted to the next level. SDMs solemnly hiring head count for the sake of having more reports. Other SDMs hiring SDMs below them, even though there’s no need for additional management (team is idle), so they can show a larger footprint and get promoted to the next level.
Voice assistant tech may as well be hitting a brick wall for technical or human computer interaction gaps that need to be thought much more in depth. But saying this was just an insanely hard technical problem is misleading and missing the bigger picture.
If you build an organization as a pet project and then throw money at it to do whatever the hell it likes, with practically no accountability, of course you won’t get meaningful results.
In recent years, Alexa became a place to chill and wait for your RSUs to vest. I personally know great engineers in Alexa, who are now stressed out of their mind about their work and visa situation because of the layoffs. But look - you decided to join a team where you basically just hang out at work, have a series of useless meetings, or half the time, your manager is complaining for some reason that his team doesn’t all show up to daily stand ups. A ton of extremely talented folks became extremely lazy, and so here we are.
To root cause: misaligned incentives, lack of accountability and a poor work ethic, shockingly poor given at least some other parts of the company are still on the other, opposite extreme.
I use it primarily to turn on and off the lights (multiple times per day), play music, turn off the TV (but not to play things on TV, too unreliable) and to raise/lower the temperature on the Nest.
We have tried to use the Nest Cams with the Hubs as a baby monitor but the cameras feeds freeze and don’t tell you so it is actually dangerous.
There's a reason voicemail has largely been replaced by email, etc...
Then there's the spying issue.
Using voice to send commands to a machine is not practical or pleasant for everyone.
This should not be a surprise and it could have been easily predicted, but hype and FOMO pushed big companies to sink billions into this tech.
This is not the first time, and it will happen again, especially because anything touching AI leads to highly inflated/magical expectations.
There are many other seemingly simple tasks it has failed at. All I use it for now is sending texts and turning in navigation when I’m driving.
Or that the results on a smartphone can be visual, whereas these home devices don't have direct IO apart from voice? [And yes, while they can be hooked to exogenous devices like a television, that's an extra step of configuration to finagle...]
OpenAI charges 2 cents per about 750 words (for their best model), or roughly 1 cent per minute of talking. Maybe they can add LLMs as a premium feature, $3 a month for an actually smart home assistant seems like a deal.
At launch I could reliably "Hey Siri" a timer across the room, now it just doesn't work because presumably Apple have downgraded the long range microphone tech at some point to save costs.
Eventually just stopped bothering and just set them manually.
I think voice assistants will continue to have good use cases in cars, kitchens etc. where you can't use your hands so the trade-off is worthwhile.
Not everything can live on ad dollars.
Kids in India have smartphones now, no one’s impressed