OpenAI rolls out Advanced Voice Mode with more voices and a new look
techcrunch.com
techcrunch.com
1. The low latency responses do make a difference. It feels miles better than any other voice chat out there.
2. Its pronunciation is excellent and very human like but it is not quite there. Somehow I can tell instantly that it’s a chatbot, it feels firmly in the uncanny valley.
3. On the same note if I was on call and there was a chatbot on the other side of the call I can instantly tell. It’s a mix of the voice with the way it responds, it just does not sounds like a human talking to you. I tried a bit to make it sound more human like, asking it to stop trying so hard in conversation being briefer etc but I wouldn’t say it made things better
And so my final review is, it is a big achievement over anything out there, nothing else comes close but it is like video game console graphics. You can instantly tell it’s not the real thing and because of that I find it harder to use than just typing to it.
That to me is precisely the reason to still use it without hesitation, because once it starts getting very-much-human, I don't know if I want to use it unless I really have to.
I think there's a lot of merit in keeping it sounding just a little artificial so that it is easier to have some psychological distance from what is already an overly anthropomorphic experience.
In religion/religious studies, there is the occasional debate of whether or not deities are/ought to be anthropomorphic, and atheism of course finds the whole notion ridiculous. Considering that our hopes and dreams with AGI can often feel religious -- maybe it's time to take that same lens towards AI.
Systems like this have existed since the 90s e.g. Dragon albeit far more rudimentary.
And the issues are exactly the same: (a) discoverability, (b) efficiency and (c) recoverability.
It is so much easier to have a screen with fixed options that you interact with, can easily see your journey and can go back for any mistakes. Versus with our voice which is the clunkiest, slowest and least precise input method we have.
Most people talk faster than they can type.
Which at least for those of us speaking non-US English is never the case.
And you have only to ring your bank and try and transfer money between accounts with it reading out every account number and asking for confirmation every step of the way. Versus a few clicks with a mouse to see that for almost all operational tasks voice is cumbersome and inefficient.
Because voice interfaces don't have the equivalent of a delete key or allowing the user to quickly select a different option.
I can't keep a conversation going with AI as easily as a person because of my poor skills, no fault of the AI.
I will improve over time, and there is no reason I won't be able to become as natural as jean luc Picard telling his starship what to do.
Some related issues I have:
- my thoughts always seem to be jumbled when talking to AI
- I rush to talk quickly because any pause seems to trigger a response
- I worry words or DSL I use won’t be interpreted properly
This all leads to a pretty poor voice experience for me, and I usually forget half of what I want to talk about.
My perfect voice assistant would sound like Auto from Wall-E, which is supposedly a blend of MacOS' Ralph and Zarvox voices. Along the lines of (bear in mind I just wrote this directly into the terminal and didn't spend any time actually blending them lol)
say -v ralph -r 180 "I'm sorry Dave. I’m afraid I can’t do that" & ; say -v zarvox -r 180 "I'm sorry Dave. I’m afraid I can’t do that"
And yeah I'm almost convinced that the whole voice interaction thing came about because they interact with the computer in Star Trek using voice commands.. which is probably just because watching someone type everything into a keyboard would be some boring telly.I assume there are folks that do use it and do like it, but do they like it more than just pressing buttons to do things? No worries of being misinterpreted or having to speak like a robot at Alexa because it's failed to turn the lights off 3 voice commands in a row now. It's awesome for accessibility, don't get me wrong, I'm talking in the sense of the primary and most commonly used interface.
[1] Not a criticism, fellow soulless machines.
In my tests so far it has worked as promised. It can distinguish and produce different accents and tones of voice. I am able to speak with it in both Japanese and English, going back and forth between the languages, without any problem. When I interrupt it, it stops talking and correctly hears what I said. I played it a recording of a one-minute news report in Japanese and asked it to summarize it in English, and it did so perfectly. When I asked it to summarize a continuous live audio stream, though, it refused.
I played the role of a learner of either English or Japanese and asked it for conversation practice, to explain the meanings of words and sentences, etc. It seemed to work quite well for that, too, though the results might be different for genuine language learners. (I am already fluent in both languages.) Because of tokenization issues, it might have difficulty explaining granular details of language—spellings, conjugations, written characters, etc.—and confuse learners as a result.
Among the many other things I want to know is how well it can be used for interpreting conversations between people who don’t share a common language. Previous interpreting apps I tested failed pretty quickly in real-life situations. This seems to have the potential, at least, to be much more useful.
(reposted from earlier item that sank quickly)
Surprising that there isn't a 'hey siri' for chatgpt yet. Obviously, that would make this sort of feature infinitely more useful. This is what monopoly gatekeeping looks like.
The limitations in this feature show the problems with both EU proactive regulation and US underregulation.
Bad regulation has become the biggest issue standing in the way of useful software for humans.
Sort of a middleman approach and certainly not perfect, but you can invoke ChatGPT with Siri using Shortcuts.
https://help.openai.com/en/articles/7993358-chatgpt-ios-app-...
Nevermind, I deleted and re-installed the app on iOS while on VPN and now it works!
I was particularly impressed that it corrected the pitch accent of some words I said in Japanese. I speak Japanese fluently but, because I began learning as an adult, I have a foreign accent that I am unable to lose. One major component of my nonnative sound is my inadequate acquisition of the pitch accent system. Nobody ever corrects me in conversation and it would be annoying if they did. If, when I started learning Japanese forty years ago, I had had a bot that could hear and correct my pronunciation, I would have less of a foreign accent now.
Some prompt engineering is needed, though, to get rid of that excessive praise. In my next tests, I will just tell it not to praise me at all. That should work.
1. It's a bit too agreeable, example: "thats an excellent point" etc every single time.
2. It understands surprisingly well. example: from experience, when I explain something vaguely, my expectation is that it would not understand, but it does most of the time. It removes the frustration of needing to spell out in much more detail.
3. It feels like talking to a real person, but the way the AI talks in a sort of monotonic ways. Example: it would respond with similar tones/excitement every time.
4. Very useful if you need to chat but doesn't want to chat with humans about some subjects like ideas, and explainations.
____
[1] Which, btw, I think deserve better sentiment. On benchmarks, the new Gemini Pro seems to be better than GPT-4o. It's just not so hyped...
That's disappointing. I wonder if it's related to legal issues, technical issues, or just doing a phased rollout?
[0] https://www.macworld.com/article/2374452/apple-intelligence-...
If they were using your voice recording for training purposes, then of course absolutely. Which then for a good reason.
Or it could also be about the way the voice recording is transported or where it's stored, for how long it's stored etc. Because speech to text could be achieved even offline and not in cloud and be there for shorter period of time. You can delete the recording and just use the text.
If the voice recording is tokenized in an ongoing conversation then it would be stored indefinitely.
Their AI product wasn't ready for other languages, not even British English where the DMA definitely doesn't apply.
A person would have to be very naive to believe Apple directly here. They just want to generate bad press for the DMA.
My suspicion is that Apple is stomping its foot as loudly as possible about the evil European Commission while getting language/localization support in order.
I wouldn't be surprised at all if they suddenly and miraculously discovered a way to somehow still be compliant in the last minute and launch the feature as planned to not endanger their sales.
edit: Sorry if privacy hurts your feelings, people who downvote me.
You're refering to the EU here, right?
Not that everything they do is good. But something being illegal in the EU is a bad indicator for sure.
AI it's harder to say, especially with the attitudes of the people making it being a split personality mix of "wow this is incredible it might kill everyone" — that's enough of a surprise to legislators that I don't know what to predict nor who has the right approach from USA (e/acc), EU (conservative), China (not sure, suspect 'harmonious' but acknowledge I'm projecting a national stereotype).
E.g. one valid use-case would be about storing your voice recording on the cloud in a tokenized format.
Because the model is now directly taking your voice, which I assume as tokens, it can't be immediately deleted as opposed to speech to text, which can be quickly used to convert to text and then deleted.
Specifically, the report called out GDPR costing small businesses 'more than' 15% of their profits.
This is indeed quite a hurdle. Privacy isn't really the issue - it's regulation that understands that complexity has a cost. We, as developers, should understand the deep, deep cost of complexity.
As a business owner that respects privacy a great deal, GDPR and regulation like it are still an immense hurdle - the cost in understanding and in doing things to the letter of the law is, I think, hard to grasp from the outside. Regulations occupy a huge amount of space in my head that was previously filled with making a better product.
PDF available here:
https://commission.europa.eu/topics/strengthening-european-c...
>except for jurisdictions that require additional external review
It's just speculation on my part that that's the actual reason for some of those markets.
It might have to do with (old) EU specific regulations. I know the UK adopted many EU laws to expedite brexit.
1. This is just speculation
2. It could just be enough reason to cause review
Unfortunately the EU strongly encourages this sort of confusion. EU Commission staff and MEPs constantly say "Europe" when what they mean is the EU. And because most of Europe is in the EU, it's especially confusing.
(Also language reasons as well - though it seems it works with other languages)
How well does it work in the UK (and understanding its regional accents?)
During this time, I dont believe I actually had access to it, as it wouldn't hum, laugh, pick up on my voice tones etc
Got my hopes up!
ChatGPT describes this as "A rich, deep, and smooth tone that is pleasant to listen to for extended periods. This often comes from good control over pitch and timbre, creating a voice that resonates well."
If you watch youtube, voices in this theme are the Pirate Software guy, and the voice of The Infographics show.
There are similar voices for every gender, race, and nationality. As an American, Morgan Freeman comes to mind as a comfy black, masculine narrator voice.
All this is to lead up to my point that companies engage in a meticulous science when deciding who should voice roles, and especially when the product itself is literally just a synthetic voice and they near limitless capacity to shape it.
With that in mind, here are the voices that OpenAI wants us to hear:
Breeze: ambiguous gender, white, feminine
Juniper: female, black Maple: female, white Spruce: male, black, masculine Arbor: male, Australian, masculine Sol: female, white Ember: male, black, less masculine Cove: male, Sal Khan, less masculine Vale: female, British
The only one that could be considered a narrator/radio voice is unambiguously black (great if that's your preference). It just seems weird that they would intentionally exclude a masculine white male, and that sucks because those are always my preferred voices when I'm looking for audiobooks or choosing a computer voice. It sucks in particular because OpenAI is not staffed by dumb people—this exclusion was intentional, and that's obnoxious.
My last note on the Advanced Voice feature is that it makes my phone HOT within a few seconds, which will limit it's usefulness on sunny days when I need hands-free use the most while the phone is mounted to my dash. This is when the device is already liable to overheat (display forced to dim, lagging due to shutting down CPU cores, and in the worst case the phone shutting off and refusing to work until it gets cooler).
PS: You can modulate the standard voices with a custom system prompt, including asking them to speak at a lower tone or with an American accent.