A16Z AI Voice Update 2025
gamma.app
gamma.app
I don't get it -- textual support chatbots have been around for decades. Even if we accept the premise that people would rather speak to them by voice, how do voice agents represent some kind of sea change in availability?
(And I personally find customer support chatbots deeply frustrating to use for reasons that have nothing to do with the modality or the quality of the AI model. I only ever need to use one when the question I have is not answered in the documentation, which is often the extent of the chatbot's business-specific training data. Inevitably I end up being led in circles, screaming for a human.)
An utterance is something like “give me directions from $source to $destination”.
LLMs mean that you don’t have to give the system every utterance in every supported language.
I work in the cloud call center space (Amazon Connect) as projects come up. I spent 3 years working in the consulting department at AWS.
Maybe this proves GP's point that it's not a desirable development.
The platform support just isn’t there yet. Amazon Connect and Genysis are the two most popular cloud platforms for it. Besides that, enterprises are just really slow to move.
This is a quick mockup I put together for an earlier HN comment
https://chatgpt.com/share/67be86bc-4090-8010-8017-f3048fe32d...
Compare this to actually listing out every utterance.
You would need to build something like this using a custom Lambda integration with Connect and Lex. I’ve done it before as an MVP.
The FAQ page of their website already does that.
> schedule appointments
I've been scheduling all my doctors appointment through their website for years already.
> complete purchases
PG himself sold a company making websites doing this to Yahoo in 1998.
Just like all the people trying to reinvent the bus, they've reinvented the website...
> It is the most frequent (and information dense)
Second, this is false. Voice is effective when the sensory context is available to both people, e.g. in the dinner table where "pass the salt" makes immediate sense. Otherwise it is an erratic form of communication, prone to misunderstanding, often repetitive and redundant.
It is not more information dense, but it is the most immediate. The latency of AI applications makes its immediacy less useful.
Voice is simply natural to humans. Downloading an app to learn about the departure of the next bus is not.
I used voice bots to let my 5-year-old play role-playing games (e.g., checking into a hotel) or let my parents (60+) call a fake car dealership.
It's amazing to observe. They behave as if they're talking to a human, especially when doing it via a phone. That is exactly the UX a computer system should have—simply a phone number and voice.
As soon as people have to learn something new (a new webpage, a new app, etc.), something is wrong.
- noise: I expected that this will be solved soon. Eg. LiveKit just announced a VAD model that works on human speech behavior and not voice detection - privacy: this seems to be a cultural thing. And can quickly change. People moved quickly from everyone on their Bluetooth headset (mid-2000) to calls at all 202x
You still get the best results by talking like a robot.
That’s… quite the claim. I guess we’re picking the worst people, the best voice-based AI, the easiest of scenarios, and a total desire for humanity to remove other human from interaction.
Pretty dark and sinister if you ask me.
This is one those claims that's like....yea I guess you can go on the internet and just say things.
What a stupid slide deck. Jesus Christ.
What if you were able to get helpful support, 24/7/365, with no time waiting in a queue, in your own language (regardless of the service provider's location and 'native' language support)? And the company was able to provide the product and support for it cheaper, resulting in less cost to you?
We're far from there, but I expect it'll happen.
I do like your idea though -- it reminds me of William Shatner donning boxing gloves and "fighting" to get you the best deal on priceline.com (gosh, I just checked and that's from 2016!!)
My beef with AI voice is it's so fucking slow. As someone use to podcasts at 3-4x speed, I can't wait to ditch human interaction if as voice agents adopt variable speech rate.
I'm sorry but I have a young daughter and I don't see that at all. A few years ago we moved to a different country and she had to learn a new language; initially she would use voice-based translation because that new language had a different alphabet so writing would have been harder. Before using the voice-activated app she would move to an empty room where no one could hear her, place the device just next to her mouth and whisper into the microphone... Did not seem to be natural at all - and she is a digital native (the iPhone existed when she was born).
If anything it seems that young people are less likely to trust tech, compared to us. She and many of her friends have permanently placed a piece of tape on their macbook's webcams. When I tell her that back in the day people didn't care much about security and we would re-use the same password everywhere, she looks at me like I'm an alien. Kids don't have the same optimistic outlook about tech that we had back then. We grew up being told that tech would change the world, they grew up with primary school courses about the dangers of cyber-bullying.
I will note that the model has been successively nerfed, massively, from launch, you can watch some demo pre-launch videos, or just try out some basic engagement, for instance, try asking it to talk to you in various accents and see which ones Open AI deems “inappropriate” to ask for and which are fine. This kind of enshittification I think is pretty likely when you are the only one in town with a product.
That said, even moderately enshittified, there’s something magic about an end to end trained multimodal model — it can change tone of voice on request. In fact, my standard prompt asks it to mirror my tone of voice and cadence. This is really unique. It’s not achievable through a whisper -> LLM -> Synthesizer/TTS approach. It can give you a Boston accent, speculate that a Marseille accent is the equivalent in French, and then (at least try) to give you a Marseille accent. This is pretty strong medicine, and I love it.
There’s been so much LLM commoditization this year, and of course the chains keep moving forward on intelligence. But, I hope Ms. Moore is correct that we’ll see better and more voice models soon, and that someone can crack the architecture.
I'll take a professional actor over TTS any day - incomparably better quality even with the best TTS.
(This is still a crazy impressive amount of work, they clearly labored over matching things to facial expressions)
In any case, I think the biggest win is that tons of books which have never received audiobooks now have the option of getting a way better alternative than legacy TTS tools. Even if current TTS tools are a bit limited, they still feel like a massive leap in quality from what was available a few years back. Making it trivial to generate better audiobooks will help make tons of information more accessible to people.
The choice of audiobook is rarely going to be between a professional actor and TTS, but between no audiobook at all or a TTS version.
Anthropomorphism is to AI what skeuomorphism is to UIs. I can’t wait for us to move into the “flat design” era of AI, where instead of being patronized with phrases like “Hi! I’m Bobby! Your intelligent AI assistant, how can I help you?” we just get something cold and straight to the point like “Ready for Instructions”, in some crunchy byte encoding. Sorry for the rambling, I’m a little drunk.
That the potential for scams and emotional manipulation seem much higher than any "positive" use cases
But the whole point of this medium is that you want the humanity and personality. Otherwise just use text.
I don't believe that. For input, maybe (you do draw things probably to explain stuff, or send reference documents). For output, not at all; it really sucks. Not only is reading faster/more economical (if you can read of course, but that's another story); adding visuals (images, charts, but tables, animations, videos, calendars, kanban, mindmaps etc etc) aka GUI really helps in communicating. That's all GUI.