AI Voice Generator: Text to Speech Software
murf.ai
murf.ai
In my opinion this tech is bad - and the more we spend time listening to artificial voices I would bet it can have a disregulating effect on the listener's nervous system.
There is also a unhealthy trend on YouTube where creators actually voice their content, but they speak really fast and they cut all the pauses. It's really stressful to listen to in my experience, and I believe also unhealthy for listeners on the long run.
It's no wonder that some creators who are just chill in their videos, sometime attract a wide audience, become a father-like figure almost - they could talk about anything - because younger people nowadays are just starving for this co-regulation effect.
Like I'm watching a certain "Dwayne" and I don't need to agree to everything he says.. but the delivery is so calm and grounded , and there's none of that speeding up / cutting pauses non-sense, that it genuinely helps me as I am recovering from trauma. It calms me down.
It's kinda unfortunate that at same time modern trauma models are gaining ground on YouTube, all about vagus nerve, fight/flight/freeze etc, the concept of capacity in the nervous system... at the same time you have an increasing assault from this really disregulating content...
I guess all I can say s more than ever you have to be really aware of what you consume.
I grew up in a family of fast talkers, at least 2 generations - predating YouTube by decades. Nature or nurture? Who knows?! Family events are lively, and we've traded stories about when people occasionally ask us to slow down.
I find listening to slow speakers a little annoying because of the lower information density per unit time. What's more important: knowing how the story ends, or the subtle inflections and dramatic pauses?
The inflections and pauses are the story. People are often disappointed when a good story ends.
If the goal of your story is just to densely transmit information then maybe you should just print bullet points on cards and mail them instead? It would save you the time wasted traveling to events.
This is exactly the reason why I prefer emails to meetings! Half[1] the meetings I attend can be replaced by emails, preferably with bullet points as you said.
Edit: as child comment has pointed out, we may be talking across each other: for stories those are important but not crucial. For professional communication, I want as little subtlety as possible
1. Perhaps more. Meetings are a huge time sinks. Few people are effective at presiding over them: unactionable ramblings, repetition, demanding that people who don't need to be there to attend. Interestingly, the higher ups who invariably demand this "face-time" use similar arguments to yours
Email could eliminate most meetings if people bothered to invest time in anticipating questions and preempting them. Providing clarity and insight rather than vagary. But they don't, thus we have meetings.
99% of the time the person asking for the meeting can even be bothered to write an agenda or give you any opportunity to prepare. At best you get vague subjects like "discuss stuff".
I'm even starting to see this laziness appear in search results for answers to obscure questions being buried in 15 min long YouTube videos which end up being screen captures with some middle school AV club quality title sequences because that's easier than writing.
But my original remarks were in regards to story telling.
Is there any evidence of this?
I feel like I myself have 'disregulated' in maintaining rapport in face to face conversations over the years. As in I now feel that I don't know what to do with my eye-gazing during a face to face conversation specially with first acquaintances, I don't feel particularly introverted just feel ackward/unsure what to do, when it was fairly effortless a few years back.
Then perhaps this is exactly what text2speech packages will optimize for ...
It also makes a great case for me to use Audm more which has real people reading news (usually longform) aloud. Often it's even the journalist who wrote the piece.
There are some things where AI voices absolutely ruin it, but it's not always a requirement for "emotions" to be felt in the speech we're listening to.
Outside of the generated glitching in the sound here, my main complaint is that sentence umpteen sounds the same as sentence one. When we speak regularly, our intonation and cadence moves over time and the subject matter. A sentence here sounds okayish, but all the sentences in a row sounds like they're generated discretely (which I assume they technically are), and all the cohesion is gone.
I'm okay with robots always looking robotic and synthesized voices always sounding synthetic.
I have no problem with robots becoming more human-like in their dexterity and locomotion, prefer that artificial voices be intelligible. But apart from "look what we can do" I see no need for either to ever try to pass as human.
I agree, the jump cuts in a lot of videos can be exhausting.
I would expect that anyone working on scripts with voice-overs professionally would want to use their favorite movie/audio editor. That means from a user perspective, a "AI Voice VST/AAX Plugin" is strictly superior to whatever cloud GUI anybody builds. (EDIT: Also, running AI as a SaaS means murf.ai needs to pay for pricey datacenter GPUs. Any user-downloadable software will have much lower operating costs.)
And the big elephant in the room with speech AI is that it's so easy to copy the tech. Just like Stable Diffusion did with images, TTS developers just train on public audio from the internet, so there is no dataset moat. And arXiv is full with papers that produce pretty good results, if implemented correctly. And NVIDIA has a collection of freely downloadable TTS models with good/usable quality. To me, it seems like it's only a matter of time until someone builds a high-quality open source TTS VST plugin and then all those SaaS offerings are basically worthless.
In effect, what I'm asking is: What is the competitive moat here? How can murf.ai defend against a motivated high school kid with $100k in EC2 credits?
My friend who runs a Shopify store asked for this. They are not going to fiddle with VST plugins or local/cloud GPUs.
The ideal TTS product for such a person would be something like: sign up and pay > choose voice > paste text > download audio
Like: invideo.io (no ad, just a very impressed user)
This assumes that we eventually get over the uncanny valley that we are all sensitive to when it comes to voices.
I think, yes, but there is a chicken-and-egg problem. I think ones that are starting to see a little youtube/tiktok income would be the most likely.
Think I read somewhere you can retrain tacotron II on a new voice for something like $6 on google colab, been wanting to try it with the ScotRail voice recording dump they did a while back (just because) but haven’t gotten around to it yet.
It seems to me like speech-to-speech would be much better: start with your best attempt to produce the audio yourself, with the emotion, rhythm and timing you want. Then let the AI do the "last mile" transformation, taking your voice and making it sound like someone else, like how neural style transfer can change a picture to another style.
But yeah, fully agreed, for individual projects speech-to-speech appears to be a better idea, much more data to work with in there. Otherwise it will be a Vocaloid-like experience, where you have to tinker with the intonation of individual words.
There is significant work in this area, too, e.g. Zero-Shot Voice Style Transfer: https://auspicious3000.github.io/autovc-demo/
See vid for some discussion around it - https://www.youtube.com/watch?v=_5uCvcyD0Eo
Anyone know why/how this company appears to be growing quickly?
You have to use the API, but if that's fine with you, it's definitely worth it.
There is some rather annoying process to setup an account with resource groups etc required but it does give you 500k free characters, or you can just abuse the free demo applet on the website without signing in (may need to clear cookies and reload once in a while), and just tape the audio that comes out.
I had very good luck with some of the Azure voices to create a YouTube video. My favourite right now is Sara (US English), because in testing she sounds the most emotionally natural.
Interestingly, if you choose a voice from another language, and ask it to speak English, sometimes it will replicate a non-native accent, which I found somewhat amusing
I own a creator platform (with 500k or so Voice Actors) and have been very interested in AI Voices so I've been watching this develop for a bit.
IMO, Murf's marketing page has better results than their product.
I think the VAs on my platform are in trouble, but they still have a little ways to go.
I'm an indie entrepreneur, so it's just me on this project, but it's been great fun.
Is there an error on the page or just it just not play anything?
Trying to reproduce.
Edit: figured it out. Will fix!
The same way some people like to put up a marble statue of their heroic deeds, others like to record themselves for the internet. In my opinion, both types of people want to avoid being forgotten and surely if you become a famous TTS voice, you'll have a Wikipedia entry...
You can try some really really really interesting things with half a million users
That's what the AI model needs to do to get a similar level of performance.
It’s possible that doesn’t matter to most people, and the art world will have to realise that mass-produced schlock is all the public really wants. We’ll see.
It's like saying: for hammering, a hammer is better than my pinky finger.
That's kind of the point of tools.
Do you assume it will stay the way it is?
I close those videos within seconds of recognising that the voice is synthetic.
I'm not sure why my reaction is so strongly negative (I don't have this for GPT or SD). My first thought was "Infinite free generation means infinite A/B testing, and I don't want to be part of that", but that should exclude those other AI also.
Unless the pricing is aggressively cheaper, can't say I am that impressed with the product.
I selected several different voices, but it only generated between 2 and 11 seconds. Only got up to the first sentence...
As for its core functionality, sounded good enough for my modest needs.
Like other voice synthesis software, though, it does not seem able to adjust the pauses and intonation to indicate emphasis and contrast the way a skilled human narrator does. I wonder if that will be coming as the AI becomes more meaning-aware.
"Our system has detected content that might be inappropriate...we request you to remove such content."
I was sent this moments after signing up and entering one single word starting with F.
I was looking for better TTS for my AI video presentation generator side project. Which one has the best voices out of those offering an API?
I gave up.