I have a Mini that consistently misunderstands broadcast requests and says "sorry I'm not playing anything right now". When it occasionally speech-to-text converts a broadcast word, it consistently cuts off the first word or letter, even when it's "I'll be right down" users will get "L L be right down".
It used to support simple offline requests like SMS and Navigate when data was unavailable. No more.
It used to integrate with Google Keep. No more.
No longer recognizes the word "torch" as a synonym for flashlight. Why would I be asking to turn on my phone's "porch"?
Painfully slow replies, even under ideal network environment... Just spinning forever, often until a timeout that it doesn't even have the decency to respond with a proper error message.
It's just amazing how they launched a product with a clear "this is where we are, this is our vision of where we're going" and they still sell it but instead they're going in the opposite direction.
Google Maps: I swear, about half the time I try to activate voice search, it sits and spins before even accepting any voice input at all. Why can’t it just start reading the microphone right when I activate it, and then submit the saved audio whenever it’s done getting set up? It’s so abysmally poor that it’s usually faster to scroll through recent destinations or literally grab the phone, unlock, and put in a destination.
This is just a market begging to be disrupted. I want to see a startup combine Whisper, GPT, and a competent TTS model into a killer voice UI!
I must disagree. ChatGPT-style LLM functionality with ElevenLabs-quality realtime voice synthesis will absolutely supercharge these products. The ability to e.g. answer kids' questions in simplified English according to parental prompt guidelines, or drill down on complex educational topics, or maintain context over many back-and-forth conversational interactions will be huge.