VALL-E: Microsoft’s new zero-shot text-to-speech model
mpost.io
mpost.io
VALL-E: Neural codec language models are zero-shot text to speech synthesizers - https://news.ycombinator.com/item?id=34270311 - Jan 2023 (136 comments)
Personally, I find that their samples aren't near anything I'd call "dangerous". I cross-compared the baseline examples to the VALL-E ones when the paper dropped, and found several that were garbled in the usual robotics sounding TTS failures.
Probably a good thing that people are getting alarmed before a true indistinguishable voice cloner exists, though.
> "Dear Fellow Scholars, this is Two Minute Papers.... "
The way. He talks. Is like. He puts. A period. Every. Word. Or two.
So... Like Shatner? Sorry, couldn't resist.
It’s like when Lex Fridman goes on a tangent about love. I just can’t do it lol.
I don't think you'd be able to pretend it's a speech given on TV or whatever, but I think they're probably good enough for phishing e.g. the usual scam of pretending to be a stranded family member, but this time rather than texting/WhatsApping them, you can do a live audio call.
Just (like the newspapers used to do to celebrities) try out lots of PIN numbers on random voicemails and lift a message from a family member if you get in. I think given that mobile phone reception isn't always stellar anyway, this would be be very effective.
Layer in a background track of random street noise and prompt the target with deflections "Sorry, the line is really bad; it's breaking up a lot. I can't hear you properly but if you can still here me, can you send me $X".
a real life $5 wrench type solution
At the rate things are progressing that's about two years away.
The impact on his quality of life - imagine not being able to communicate at all - would have been massive were it better.
I expect it will be a while until we can fully utilize that data, but I have to imagine that something could be done today to preserve my voice (while I am still in my prime). Effectively, this would be a sort of vocal cryogenics, betting that we can do something today that will allow us to take advantage of future technology.
I wonder what sort of recordings and other data you'd need to get this right assuming what TTS might look like in a few years' time
Hawking is very attached to his voice: in 1988, when Speech Plus gave him the new synthesizer, the voice was different so he asked them to replace it with the original. His voice had been created in the early '80s by MIT engineer Dennis Klatt, a pioneer of text-to-speech algorithms. He invented the DECtalk, one of the first devices to translate text into speech. He initially made three voices, from recordings of his wife, daughter and himself. The female's voice was called "Beautiful Betty", the child's "Kit the Kid", and the male voice, based on his own, "Perfect Paul." "Perfect Paul" is Hawking's voice.
[0] https://www.wired.com/2015/01/intel-gave-stephen-hawking-voi...Something like this could let someone keep their "original" voice. Hawking may have preferred to have something that sounded like his voice before he lost it than a completely new voice.
I agree though, Tortoise TTS did a lot of similar work IIRC by a single person on their multi-GPU setup. Really impressive effort. Did they get a citation? They deserve one.
edit: reading other comments it seems there is a misconception that the model takes 3 seconds to run? That isn't the case - it requires "just" 3 seconds of example audio to successfully clone a voice (for some definition of success).
rtx5000 maybe but not sure how much of a value improvement there is
You were merely an unfortunate casualty of Nvidia's product marketing scheme (and a commenter's slightly imprecise reference to it) here.
Do other people read these types of headings and just think immediately about how freaking scary this stuff is going to get...if it's actually accurate ?
The only possible solution I have seen being talked about is digitally signing any and all "real" content, similar to GPG/SSH.
I don’t think it’s elitist to think many people will be duped. I think it’s arrogant to think you won’t.
There's also the unravelling biography of https://en.wikipedia.org/wiki/George_Santos#Scandals (you know someone's doing well when their "scandals" section has ten subheadings! And that's all old-fashioned manual fraud)
(I'm now wondering how far you could get with an elected official who's entirely deepfaked and only appears in recordings. Probably all the way to being sworn in, at which point someone has to appear in person. People have occasionally been elected despite being dead, so it's not impossible)
If electoral boards around the country or at the national level were all sounding the alarm for irregularity in voting returns or procedures then an election might have been stolen. As it stands, the election results match the earlier polls which is a strong indicator that, at best, the election became a tossup at the time it was held, and the military also didn't find any reason to suspect fraud or error.
Unless the people doing the swearing in are also fake.
This matters a loooot more than the smart enough people realize too, because it's so difficult to imagine what it's like being south of that threshold, the potential consequences are difficult to imagine.
> Or do you think you yourself will fall for this as well?
Yes.
If the news confirms what they think is the truth, it doesn't need manipulation, just a fake heading
https://www.snopes.com/fact-check/climate-protest-ukraine-bo...
for example (and in that case it even had the true caption on the video, just not in English)
Maybe, just maybe, human society will move back to personal interaction, since apparently you cannot trust call, voice and video, anymore in the future.
Which implies a much narrower, less diverse, more elitist system of both politics and economy, as it gets harder to trust people from outside your existing circle.
> No video, audio, or text can be presumed real unless signed by the human subject.
Where is the controversy in that?
Follow the thou g ht through to how it would work and what it would mean in terms of privacy for the average person.
Your niceties seem dystopian to me.
Companies can't even wrap their heads around sim swapping scams and social engineering of secret questions.
Journalists would have a hard time showing that an interviewee actually said what claim they said.
Wire-tapping by the police also would become useless in court.
⇒ I would adjust it to “unless it has a reliably chain of trust”. That would mean that people could judge differently about whether to accept some things as true, but I think that’s unavoidable.
But hey, at least this will solve the problem of journalists quoting people out of context: they won't be able to claim you've said something without you personally signing off on it with your cryptographic key.
I don't quite understand what you mean. We already trust phones and computers not to be compromised for things like online banking, so why not allow these devices to sign recordings using keys lent to them by the main trust device? You could allow revocation of these keys.
Sorta, kinda. "We already trust phones and computers" because we have no better alternatives yet.
There's a reason online banking is at the front of the push for device attestation, and thanks for Google happily obliging, rooting your Android phone is pretty much pointless now. This is coupled with a parallel push to do your banking through a mobile app, or at least to use it as a second authentication factor, which makes the app necessary just as much. This means we're already undergoing the transition I'm talking about: in the nearby future, you'll have to use a pristine, unrooted, unmodded phone from a blessed corporate vendor, to perform basic functions like paying for things and managing your account. And yes, such device will be appropriate for signing everything else too - and the vendor owning the key will be your Trust Authority.
All digitally approved contracts I‘ve ever encountered were e-"signed" (i.e. me emulating a paper signature at a computer), not digitally signed, due to lack of a mutually trusted public key infrastructure.
In that sense, I would be sad but not surprised if we instead moved to a system of "trusted recording devices", where scanners, cameras and microphones by some vendors are defined to be trusted, signing their outputs and leaving the underlying workflows (signatures, recordings of verbal affirmations etc.) largely untouched.
I think that's a pretty apt summary of the state of digital ID in Germany: Commendable technical basis (with some privacy-preserving aspects too, including selective assertions, e.g. "older than 18" or "EU citizen").
When I called about actually getting such an ID card (several times) as a non-citizen, which is explicitly mentioned as a feature of the scheme, I only received the acoustic equivalent of blank stares.
For that reason, Germany introduced a "non-identity" (i.e. not doubling as a photo ID) e-signature card in 2021, but apparently almost nobody is requesting it, so most administrative offices don't know how to issue one...
My primary worry is for criminal misuse. There's a class of scams where the scammers find elderly victims, call them, and pretend to be their children or grandchildren on vacation in a plausible city and suddenly in need of money to help get their passport back or a ticket home.
There are any number of cases where I have accepted an order worth large sums of money over the phone from someone I have previously dealt with and feel familiar with their voice.
https://www.fidelity.com/security/fidelity-myvoice/overview
https://www.schwab.com/voice-id
https://blogs.perficient.com/2016/04/18/3-benefits-of-vangua...
When it's a text message it's likely going to ring a bell somewhere (hopefully an alarm bell), if it is a voice message and it sounds like the original many more people are going to fall for it.
The scammer starts a video call with the person they want to impersonate and record it. When that person takes the call, the scammer doesn't say anything. This creates a 5-10 seconds video of the person looking at the camera waiting for the scammer to say something, until they get fed up and hung up.
The scammer then calls the victim. They offer a video call as verification straight away, and they play that short video.
E.g. in the EU we have ID cards and passports with biometric information and NFC, and there was a beta of a phone app by the French government where it would read your photo from the ID card/passport, and compare with a video selfie. That way you get fully local and secure identification allowing you to do stuff online that would otherwise require you to show up in person to a government office.
1. Ignore voice messages
2. Ask pointed questions to any live caller, esp. one calling from a new number and asking for money...
"In case we get cut off, what phone number can I call you back on?"
If they actually give you a number instead of hanging up (which has NEVER happened in the many times I've used this line), just hang up immediately and call them back, to see if they answer.
EDIT: Also looking forward to AIs scamming each other, fully-automated.
As soon as AI gets out of hand and starts randomly hacking things we're probably out of luck anyways.
“If the rise of an all-powerful artificial intelligence is inevitable, well it stands to reason that when they take power, our digital overlords will punish those of us who did not help them get there. [...]” - Bertram Gilfoyle
In this bleak scrape the bottom of the barrel world, I can't help but think it's only a matter of time before customer service departments start selling voice data to the highest bidder. And to counter that there will be services where everyone ends up using voice changers, except close personal contacts. I want this fiction to stay fiction.
As opposed to what? Writing is even easier to forge, no?
As a millennial, I'm way ahead of you on that one. The government and telecom companies just need to crack down on number spoofing. Any day now..
With AI we have lost all privacy and identity. This makes nonsense like novelty theory and timewave zero seem less nonsensical. Lets hope we make it through the great filter.
My thoughts on this are that oftentimes things at the extreme ends of the distribution are usually in some way, shape or form problematic. Compared to all other living creatures we know of, humans are right at the extreme tail end of the intelligence distribution, and the result is somewhat pathological when compared to other creatures. I think nothing that's north of a certain threshold required to make advanced technology can make it through the great filter without destroying itself in the process.
Of all my concerns for humanity, excessive intelligence is not one of them...
Our earliest (and easiest to detect) transmissions should already be far enough away to be far below ambient noise I think/hope. So maybe we made it through unless someone starts shouting. Own-goals are also certainly a thing, but it would take a serious mistake to end all of us, or even set us back that far technologically on a geologic time scale.
I must be a pessimist in this regard. I don't necessarily think it would take a serious mistake, rather just the nature of complexity itself will likely become a factor at some point and I can imagine certain scenarios playing out based off that.
When you compare the length of time of humanity spent pre-neolithic-revolution with the short ~14,000 odd years since, I can't help but feel our current trajectory won't see us reach hundreds of thousands more years into the future without a serious backslide in that time.
Humanity is pretty darn resilient, but the last few years have highlighted to me that modern day life is some weird fantasy land we've created for ourselves that has an expiry date at some point in the future.
The government just should give you a card sized mini computer with 2 large primes stored in it, and everyone is perfectly fine. In fact, more secure, than the current situation?
It'd be better to ask something personal. "What did we do to Christmas/Birthday/[other event] last year? Who was with us?"
https://en.wikipedia.org/wiki/Prime_Computer
https://en.wikipedia.org/wiki/Prime_Computer#/media/File:Pri...
“ Here’s my new account number. It’s really me. Slippers in the toaster.”
[1] https://www.theverge.com/2020/2/18/21142782/india-politician...
The unendorsed version is far more dangerous.
That said, I don't understand why the parent and the article paint this as a bad thing.
It's a new use of technology to repeat an old lie. I think it's a bad thing, but it's hardly the most dangerous and novel application of this technology for evil.
Do you understand what is first-past-the-post voting system?
My employer is known for tracking a decade behind the curve and we already have financial processes in place to prevent this vector from being exploited. That leads me to believe the majority must already have as well.
That is, you have a "filter bubble" which tells you that there isn't going to be a Nigerian prince with some money waiting for you (unknown sender, unlikely message). Reduced cost fakes force people to make their filter bubble smaller by raising the threshold for faking. More broadly, the question of "high trust vs. low trust" society.
Unfortunately the move towards it is nothing new, we gradually lost text and images and now the remaining media is closing.
This will probably decimate online public forums completely and off-person communications will en up being P2P through trusted devices. Closed communities in WhatsApp, Telegram, Slack and Discord already replaced forums.
Maybe if infestation of closed communities with machines becomes a thing, maybe the real life gatherings will make a comeback?
I don't know but I used to mock Sci-Fi movies about portraying aliens as uncivilised that don't seem to have the means to have developed all that tech but I'm no longer sure. Maybe what will happen is, we will create a symbiotic life where machine will no longer be tools but partners and they will take care of us to reduce us to our basic instinct.
What cases do you think it will be used wrongly in? It can't do 1:1 realistic phone calls, the processing it'd need for that is beyond the reach of any average user.
I just would like this so I can copy-paste books/text and listen to it in the voices of my favorite narrators, but can't have that, can we? No, it'll be super dangerous, I'll listen myself to nuking a country or some such.
Spreading propaganda is the most obvious use case, assuming the resource requirement remains high. I don’t think it’s far-fetched either. If these models become less resource-hungry, then we’re in for much worse.
If you really like them and their work, why would you want a forgery?
If you can't see the danger in this, perhaps it's more that you're ok with it.
It'll be interesting when someone releases the first MMO which combines GPT models and VO generation to fill out the in-universe world with dynamic characters who can react to surrounding events.
Like the first one "just his scribbles that charmed me" sounds so weird compared to how the sample sounds.
Second one, sounds very close but again "just his scribbles that charmed me" sounds off and wrong. His "scrouples? that charmed me".
Third one as well, it's very up and down (not sure the correct science words for that type of speaking), "Dynamo and lamp.... Edison realized". The second one is completely flat and seems computer generated.
Overall these don't seem very good to me at least. It's clearly not to the level of some other AI coming out where it is very difficult to determine the human part, I could easily pick up the human and generated version from these samples.
https://www.ibiblio.org/ebooks/James/Turn_Screw.pdf (see p4)
The uncanny valley makes it more spooky to me. Like a reanimation sort of thing.
Would not not be time to measure a model's success on the actual job? Like feeding a simulator with actual data from real-world traffic scenarios and running Tesla's, or any other company's, autopilot in it?
Seems like we're very often interested in something that looks like the job rather than the job itself - maybe because it's more easily measurable? We then took this skewed ambition and developed AI after it.
Is there any other sensible goal function for a text to speech model?
> Similarly, ChatGPT is not trained to output sensible and meaningful statements but rather statements that appear to be to a human reader.
Almost certainly not true. If they could they would make it output sensible and meaningful statements all the time.
> Like feeding a simulator with actual data from real-world traffic scenarios and running Tesla's, or any other company's, autopilot in it?
Do you seriously think that self driving car companies are not doing this already?
Yes: to transfer information aurally. And the models are there, and have been for quite some time. They can simply stop with the trying-to-fool part.
But no matter what, it's unnecessary to make a tool that can copy anyone's voice. That just lowers the threshold for abuse, while not adding almost nothing of value.
And with the way the models work, once you have a model that can sound human, it is unfortunately very easy for it to sound like any individual human as well.
The correct inflections, pauses, annunciations, etc. are all important to humans, especially so for audio books and similar things that need to immerse.
Otherwise you have to strain to listen, similar to listening to someone with a heavy accent.
Perhaps real human readers can help?
> The correct inflections, pauses
Because a model that can imitate a voice is still not capable of that. There's no need to have model that can do that. A robotic accent is best. Or perhaps you like to see your politicians make all kind of bizarre statements on youtube.
Certainly, and they do, but computers can annotate much faster, without attrition, and cost much less.
Generally you'll have human recordings for what you can, and TTS for anything missing.
There are a lot of books, live streams, podcasts, articles, etc. in the world.
> There's no need to have model that can do that.
I wouldn't say there's no need, or otherwise we wouldn't talk that way. It's a human element and the generated speech is for human ears.
> A robotic accent is best.
Best is subjective because that's a human preference. That being said, I'd say your preference is a far outlier.
Most people's "best" would be what they are used to hearing; the speech of native speakers.
- Wolfie is fine dear. Wolfie's just fine. Where are you?
Jenette Goldstein is a true chameleon. She is so effective as an actress, it is nearly impossible to recognize her from role to role.
Of course, it happens over a pay phone so perhaps with the full vocal range in person it would have been different.
The market economics were also vastly different at the time, streaming really has changed the industry, and the quality of tv shows has improved dramatically.
Recently-ish: Coco, Soul, and Inside Out are among my favorite Pixar movies and all compete with the golden-era Pixar classics in terms of being memorable and story-first.
The Disney acquisition seems to have crushed some Pixar magic. You can no longer assume every Pixar movie to be gold, but they can still turn out top-quality content.
It's about the pacing, pauses and slower character and world building.
Newer movies are built for second screening and maximising action sequences to not bore the audience which kills all atmosphere and sense of time and place - the worst example of this is the new Avatar movie, absolutely horrible in every sense of the word, while the first one was alright as i remember it.
You simply need to "set the stage", explain why this story is important and why the characters deserve sympathy or hate before you start your 3 hour action sequence - this step has been removed for some reason.
Verhovens old movies were the same - there was a nerve, a seriousness, a reflection beneath the action, now it could just have well been created by an alien algorithm without a sense of the actual human experience.
I wonder if it's a ridiculously extrapolated but misunderstood tic-tokification of cinema to please marketing? Because i've seen 20 second tik-toks with more emotion and character introduction than a lot of newer movies.
Sure but this assumes movies are about characters and emotion. Michael Bay gets a lot of shit, but sometimes you just want to see things explode. Sometimes you just want to see robots fighting for two hours. Sometimes you want to see California swallowed by a tidal wave and the emotional character plot lines get in the way.
I would pay to see a movie called “Two Hours Of Giant Rocks Hitting The Earth: No Characters Edition”, and the fact Michael Bay is rich means a lot of people agree.
That's the truth right here. I once ran across kung-fu movie called Chocolate (iirc). The premise was an autistic girl who was good at fighting. The entire movie was her walking into a room, kicking major but, then walking into a different room to kick more but.
It was great.
But would you go to see Giant Rocks 2?
> the fact Michael Bay is rich means a lot of people agree
This is the argument often made for Avatar against the "no cultural impact" observation. It still doesn't have any quotable lines or memorable characters.
If the graphics were better and the explosions were bigger, yes!
Theres a difference from both Avatar 1 and quite a few of Michael Bay's older movies - they still have have a story arch.
Just a few minutes of character building and a few intermezzos and these movies would be much, much better in my opinion.
But i also hate too much CGI. I don't know what happened to "well dosed", it makes what's dosed so much more valuable.
Jokes aside, it's funny how that sounds the most boring stuff to me and you'd have to pay me to sit in a chair for 2 hours watching that.
With the explosion of synthesis software and digital audio workstations its easy for a small team to score a sound track on the cheap.
As with everything tough, there's always a few outliers, like (both) of the Tron: Legacy soundtracks.
Other interesting soundtracks done by musicians not known for soundtrack work you might find interesting: Fight Club (Dust Brothers) Event Horizon (Orbital with London Symphony Orchestra) Chaos Theory: Splinter Cell 3 (game) (Amon Tobin)
In reality there's as good or better movies now and a ton of crap from back then, too.
I have to specify blockbuster because otherwise it doesn’t make much sense to compare Terminator to a small indie movie
Also: Top Gun: Maverick
I’d also compare _Everything Everywhere All at Once_ favorably against a lot of old blockbusters.
Modern technology aside, Top Gun: Maverick feels like it could have come out thirty years ago.
> Top Gun: Maverick feels like it could have come out thirty years ago.
I've not seen it, but to what extent is that because it's a remake of a film that came out 30 years ago?
Comparing computing power is a bit handwavey, but former Pixar employee Chris Good estimated that the SPARCstation 20 render farm cluster that rendered Toy Story had only half the power of the 2014 Apple iPhone 6. https://www.quora.com/How-much-faster-would-it-be-to-render-...
I thought Terminator was a relatively low budget movie. Not an indie though.
People like to criticize back-references as a cheap way to entertain the audience knowing the referenced works, but I think it's a legitimate feature. A movie or a show doesn't exist in a vacuum; watching a sequel or a work in the same universe, I expect to see both subtle and direct call-backs.
Watch Jurassic park and Jurassic world back to back, or The mummy vs the remake with Tom Cruise, too many cameras, too many view points, too many cuts
My only real complaint was that the time between them landing and the Harkonnen's invading was too short and didn't cover all the intrigue/politics occuring on Arrakis before the invasion.
In the movie it came off as: Okay Leto you get Arrakis now LOL JK WE'RE INVADING IMMEDIATELY. But that being said, I understand there are length constraints and I think the movie was already about 2.5 hours so I can forgive them
I'm fine with how they did it in the movie I just never really jive with shows where everyone sucks.
“You sense that Arrakis could be a paradise,” Kynes said. “Yet, as you see, the Imperium sends here only its trained hatchetmen, its seekers after the spice!”
Paul held up his thumb with its ducal signet. “Do you see this ring?”
“Yes.”
“Do you know its significance?”
Jessica turned sharply to stare at her son.
“Your father lies dead in the ruins of Arrakeen,” Kynes said. “You are technically the Duke.”
“I’m a soldier of the Imperium,” Paul said, “technically a hatchetman.”
Kynes’ face darkened. “Even with the Emperor’s Sardaukar standing over your father’s body?”
“The Sardaukar are one thing, the legal source of my authority is another,” Paul said.
And yep, it’s adapted from a a book!
Dodgy animatronics were first scoffed at, and then forgotten pretty quickly as everyone just got into the pure enjoyment of that movie.
This is a nigh on 40 year old film, with a lot of target references, yet it still hits the mark..
That said, rewatched transformers too, purely for the joy of the initial transformation (childhood toys coming to life..), and thoroughly enjoyed it (though it's already feeling a bit dated).
Aside from the voice acting, there's nothing redeeming about those movies and felt like someone just pissed on my childhood. Any aspect of the cartoons or toys that inspired joy and wonder was lost in translation.
The initial transformations were spoiled by trailers and also just disappointing in general. The transformations are largely incomprehensible and they might as well have used Star Trek transporter FX or a flash of light to transform them.
That movie is so well-paced even by modern standards.
- Quentin Tarantino
With that said, I was surprised to see how decisive, resourceful, and adaptable the mother was. Hears a noise in the attic, grabs a knife. Gets attacked, puts creature in nearest blender/microwave and hits switch. Hears another noise, grabs two knives because one wasn’t enough last time.
I believe Paul Thomas Anderson partially left film school because his professor was shitting on the movie. And he was damn right.
I think Cameron’s secret weapon is the camerawork and editing. He really lets the camera breathe and has an incredible sense of “beats” for every cut.
For whatever reason, he through out story and pushing people to their limits for technology and spectacle. It makes me sad to think that the last movies we’ll see from Jim Cameron are likely all in the Avatar universe, written by committee and acted in front of green screens.
The director's cut makes it even more clear - Ripley lost her baby while she was in cryosleep, which gives her character's desperation to save Newt even more urgency.
Top Gun Maverick would be a data point that a lot of people agree. No matter how good the CGI, actors con a sound stage in front of a green screen just behave differently than actors or stunt people actually doing action stuff.
I think it was always supposed to sound a little dumb. The T-800 is acting as a father figure, and bridging the cultural gap with his surrogate ‘son.’
James Cameron would have been 36 when he was making Terminator 2, and he's a polymath with an eye for detail, so I'm sure he deliberately went for that layered meaning.
A fascinating practical effects scene, cut from the theatrical release, is when Sarah is removing a chip from the T-800's head in front of a mirror. Instead of using CGI, there is no mirror but a hole in the wall, where you can see actual Arnie and Linda Hamilton's twin sister acting as the "reflection"; the people closer to the camera are actually a dummy (or a double with heavy prosthetics for the hole in the head) and Linda Hamilton. This means they are sync'ing their movements to simulate a mirror!
https://youtu.be/wrDo7wVXrBQ?t=122
(I understand this was done both to avoid showing a camera reflection on the mirror, and also to avoid using a dummy for Arnie's face, which worked in the original Terminator but would have been too noticeable for T-2's era).
Just yesterday I learned that in Cliffhanger (also aged amazingly well), they paid one million dollars to a stuntman to actually travel the harness between the two planes while airborne. Not bad for a day's work.
(Edit: with artistic license obviously. It was plausible-sounding instead of cringe technobabble).
EDIT: I just noticed that the ones who replied to me interpreted it the other way around.
I think there might have been other giveaways in person where the T-800 is concerned.
What made it even more amazing was that the T-800 ran on a 6502 processor.
I felt obligated to find and post the link to the scene. What a great movie.
That deadpan delivery always makes me laugh.
Or for games where you can get the language for the NPCs all sorted out and change dialogue right up until it goes gold.
Part of this was to voiceprint me during the conversation.
The agent assured me that this was more secure than asking me the customary identification questions, and that their system would just use the voice identification in the future.
While I agree with you that the net effect of such technology will be negative, there's no way regulation can keep up.
What about this: You have a voice actor whose voice is part of a company's brand (maybe for an animated mascot). You can now also use that voice for dynamic text, for example audio books or for a voice assistant.
I don't want to live in a world where people are doing jobs that provide no value but only exist because "everyone needs a job". I would literally prefer we give them the money and they do no work than forcing people to do a useless job just for the sake of them working.
Do people really want their jobs to be the equivalent Sisyphus pushing a boulder up a hill aimlessly everyday? There's a reason why this is seen as a divine punishment in the myth.
I agree that I would prefer to move to a society where we don't have to work, but do you honestly think we're moving in that direction? We'll get the "there's no work" part, but not the "and here's some money" part.
I wouldn't worry about that work not existing. At the very least, with the population declining we'll need people to take care of the elderly and I don't think we're anywhere close to automating health care providers. Humans always seem to find more work to do.
A high school friend of mine did that and became the voice of the "pronounce this name" feature on Facebook.
Maybe pick a couple of different ones for context-aware messages. Something familiar/reassuring for most messages, something less comfortable for "terrain!"-type alerts. If the model supports it, it could even morph between the two as the criticality of the message escalates.
https://www.youtube.com/watch?v=NAv1o9C9qow
https://www.youtube.com/watch?v=4NOmalcIZZw
https://www.youtube.com/watch?v=eCvqSo0C5ns
But in this case, couldn't your argument have been applied to photoshop or video manipulation software? Those are meant to deceive, right?
1. Speech to text
2.fix up/edit text with GPT-3
3.text to speech in the original speaker's voice(s), preserving prosody and inflection with Vall-e.
If done with every participant's consent I don't see how it's not legitimate.
Guns are regulated, but we still have metal detectors.
But, obviously, the technology is getting to the point where a decade or so from now she'll be able to have a GPT-like chat with me with my own voice. The first company to offer that to the loved ones of a deceased person will make a fortune, not for any mode of deception but just to soothe the hurt.
I mean this kindly: your lack of imagination is not an adequate replacement for actual facts
I don't lol. OP is a bundle of sticks. Every time there's a tiny bit of tech progress with this stuff they just go "MAH REGULASHIONS PLEASE PAPA GOVERNMENT SAVE US!"
People just don’t realize that technology is what is behind the ability for people to wreak havoc on an unprecedented scale. One specific organism can’t do that much, even with a sword. But today, every person will have more and more power, and you can’t possibly stop them all!
oh please. You never wanted to listen to an article/book in the voice of a narrator you liked? I do. Maybe you should read a bit more.
Just because you lack in thinking of the ways it can improve our lives doesn't mean everyone else doesn't. That's a you problem. Are you so deprived of free thought, and so insecure of your capabilities, that the first thing you do when seeing new technology is turn to the government to curb it? if you don't like it, no one should use it?
>regulation
ah there it is. the 'answer' to every technological progress you have is "regulashions!"
>guns are regulated
thankfully they are not in some places. So it should be regulated the same as they are - Not at all.
I've wanted to record readings of some of their favorite books to pass on to them, but if I don't get the chance this seems like a way to get some analog of the experience.
May you elaborate on that? Do you mean that you need a large training set of your voice and you need $$$ in order to train the models on an expensive GPU?
Instead, usually companies train such models in the cloud. GPT-3 for example used 800 GB of training data and cost about 5 million USD to train. Extrapolating from that, I guess that the same setup would cost 375,000 USD to train (although I assume that this model has waaaaay fewer parameters making it a lot cheaper -- but I can't seem to find how many parameters it has).
If someone else already spent that money to train the model, then you could just take the "training weights", feed them into the model, and it would be as if you had already trained it -- at which point you'd only need to provide your own voice and retrain it on your own GPU for a short time.
I'm by no means an ML expert though, so I could be totally wrong on this.
I think a remake of the movie "Gas Light" using these kinds of AI technologies needs to happen as a sort of social commentary on where this is going.
- ChatGPT
- DALL-E
- VALL-E
- Stable Diffusion
maybe it's just a reflection of my interest being peaked with AI junk and google ads -- but I feel like I'm seeing it more and more even in HN.And the biggest drop isn’t because crypto is worth less, it’s because Ethereum doesn’t use proof of work since November (Bitcoin hasn’t used GPUs for years). The explosion in novel AI models we saw last year predates this change.
Google does have their TPUs(Tensor Processing Units), however it is not cost efficient budget wise, so unless you have some kind of deal with Google or compute credits it doesn't make sense. They have pods upon pods of TPU clusters though so the main selling point of TPU training is that you can get your training done really fast with just the ease of scaling your workload to more TPUs.
So if you needed a big model like GPT-3 trained in a single day, you could spend an ungodly amounts of money and get it done with Google TPUs. Otherwise if you can wait weeks or months you can go with the standard Nvidia data centre solution and it'd be cheaper at the end by a significant margin.
But I'd be interested to see what its reliability is like on various types of Scottish accent.
Might give it another go on an easier task.
The only thing I'd use Whisper for is transcription part of it. Then use ChatGPT for the translation.
Surprisingly, I tested it with my mother, she has a very broken english accent/dialect and it worked fine for her. Works amazing actually and I'm busy building some tools around it for her to test.
Hmmm, now I'm wondering if you're actually an advanced AI that has already simulated what's going to happen.
Underneath all the commercial interest some specialists inevitably must keep adding to the knowledge base / capabilities but good luck finding any objective account of that process.
At the moment, ChatGPT and Stable Diffusion look like the same order of magnitude of advance.
It turns out it comes from a book called "The Foolish Dictionary: An exhausting work of reference to un-certain english words, their origin, meaning, legitimate and illegitimate use, confused by a few pictures" by Wurdz and Goldsmith. (Talk about nominative determinism -- a dictionary by Wurds?)
Full entry:
HAMMOCK From the Lat. hamus, hook, and Grk. makar, happy. Happiness on hooks. Also, a popular contrivance whereby love-making may be suspended but not stopped during the picnic season.
My education is in biometrics and our undergraduate seminar group had a student and his supervisor with essentially identical voices.
But that's just anecdata - the human voice doesn't carry a lot of unique information and on top of that it changes over time.
In biometric identification the iris is king. 3D face scans are a distant second, fingerprints on the border of usefulness and the rest like gait, voice, keystrokes etc. get rediscovered every decade or so just to, again, not yield results.
> Scammers get hold of someone's voice which is synthesised in ~real-time whilst they're on the phone with that person's mother, asking them to transfer them some emergency funds to a new bank account.
We won't be able to trust digital sound, digital image, digital text, and soon digital video. We need proof of authenticity / proof of authorship / proof of humanity!
This resonated with me when I read it. Seems clear where AI is going, so it seems clear that we need some sort of "authentication" that can reliably differentiate a real human from an AI one.
For example I noticed it couldn't quite pick out the accent on some samples because they were so short. But if the model had more example words to hear then I'd think it would accurately understand the accent..
In some sense I feel like we already have good solutions in place when dealing with this problem in other contexts.
We can for example get some guarantees about “who wrote this code” with signed git commits for example.
The problem is a lot of our commonly used communication protocols were never designed with them in mind.
So you would need microfone with TPM chip and you are good to go. You just need that video player which can verify the data.
Making things much easier for scammers. Voice phishing and social engineering is now easier than ever.
this will enhance accessibility to a vast number of people with disabilities. Also imagine reading your books/articles in voice of your favorite narrator.
A normal person will have to work really hard to do a 1:1 real time voice obfuscation with the processing it'd take.
----
Why are you lot always doom and gloom and small-minded, is beyond me.
I like all of the above use cases you suggested, I like science advancements. I’m sure this will help plenty.
But, at the same time, we should be really careful and aware of what we’re creating.
The SBB (Swiss Rail) is now using a quite sophisticated TTS system [1] that sounds very clean and has a Swiss German accent in the high German its talking (Zürich accent based which is what most of the complaints are about). The old system was built using over 10k recording which was a enormous task
The system now is able to announce pretty much anything including reasons for delay etc. and it does not sound like a bunch of recordings attached to each other.
I am hoping it will be expanded eventually to switch accents depending on the region just how it already switches to French or Italian depending where you are.
[1] https://www.tagesanzeiger.ch/so-klingt-die-neue-stimme-der-s...
I've given up on this a long time ago. Ads are almost always in a Zurich accent, except if the ad is for extremely regional stuff like secret cheese recipes or tourist ads for Graubünden.
Zurich is the tech hub of Switzerland and the accent seems to be the default. Then again, who in their right mind would want an AI to talk with an accent from Thurgau (/s)
Have you watched it/did you like it?
I'm regularly astonished at how bad international calls in particular have become, and you're regularly subjected to these even domestically, since so many call centers are in the Philippines or India these days. And this despite bandwidth being cheaper than ever.
I've followed some of the research on prosody transfer, etc., but it still seems bad in the TTS systems I've heard.
After being ridiculed for my horrendous English pronunciation I watched hours of English movies and repeated every single sentence I heard until it sounded right.
They'll sing never written verses, non-existent sounds in a velvet voice.
"Text to speech? Anyone voice? What?"