Anger from voice actors as NSFW mods use AI deepfakes to replicate their voices
pcgamer.com
pcgamer.com
It's the same with other game assets - textures, models, animation etc. In a mod, you just have to match the original game style closely, there's no other way. The common view in the game industry is that it's fine to do because modders usually don't make money. (at least they didn't before Bethesda and alikes let them...)
In that example, someone else was paid instead of the original actor while using their voice, doing the job that neither the original voice actor nor copyright holder approved (in the context of NSFW mods the article is talking about). What makes it more ethical all of a sudden? What makes it less of a plagiarism?
It seems to me that the actors in the article don't want their names to be associated with NSFW mods in particular (which is understandable), and the AI part is a red herring.
Anyways, it is the custom there that Nexusmods moderation reacts to suspected mod asset stealing and such an issue can be raised by any website user (as mod authors often leave and are no more around). However, for the voice cloning/impersonating, it requires a notification from the original VA, or some other credible confirmation that the particular VA is against this.
Someone or some group on a crusade against NSFW material seems to be trying to use those actors in their witch hunt.
This is like hating libraries because they buy a book once and let everyone read it.
For the analogy to work, libraries now provide the service of generating new books/stories based on works under the library’s care. As a library patron, you’ll give the librarian a prompt, and they’ll generate a book for you.
This book will be intrinsically dependent on the authors of all of the books in the library, but those authors will not be credited or paid for the new work.
“Librarian, please generate a fantasy novel in the style of George R. R. Martin that continues where A Dance with Dragons left off”.
If this is what libraries did, authors would not want their books there.
This is not what libraries do, and the analogy does not hold up.
It plagiarizes then effectively puts them out of business.
Isn't that a concern for any artist? We've had this discussion with digital art and photography, and decades ago with electronic music, remixes and sampling.
AI is just enabling this on a larger scale, which will disrupt many fields, but copyright law will broaden, and artists will find ways to adapt or change careers.
We had very different discussions about all of those things.
There is a certain structural similarity between AI and these past advanced in the form of: new thing disrupts old thing.
But I think it’s deeply problematic to take that analogy much further. Take digital art. I don’t think it’s fair to compare the impact of the advent of digital painting tools with the advent of tools that systematically ingest all paintings and the remove the need for the original artist entirely.
If removing the artist entirely was part of that discussion, I suspect the tooling and legal landscape would look rather different today.
> AI is just enabling this on a larger scale, which will disrupt many fields
“This” and “larger scale” are doing a lot of heavy lifting here.
Nuclear weapons just enable this (blowing things up) at a larger scale. But these weapons also show us that scale introduces risks and factors not present in any prior iteration of the profession of blowing things up.
My point is not that AI art tools are as dangerous as nuclear weapons, obviously, but that “it’s just x at larger scale” breaks down when the shift in scale is large enough.
The result is something entirely new, for which the past rules of engagement no longer apply.
That said, we've had similar challenges before, and society has adapted. I'm pessimistic about the long-term existential risks of AI, but the short-term disruptions to jobs and the legal changes that will be required seem manageable, and are not the doomsday scenario that the media makes them up to be.
> But I think it’s deeply problematic to take that analogy much further. Take digital art. I don’t think it’s fair to compare the impact of the advent of digital painting tools with the advent of tools that systematically ingest all paintings and the remove the need for the original artist entirely.
The invention of photography in the 19th century certainly had the same, if not greater, impact for painters. Yet artists adapted, and paintings were able to coexist with the new technology. Photography opened up new avenues for art, but it didn't eliminate the demand for the traditional art form.
So will happen with AI-produced art as well, I think. The markets and our media feeds will be flooded with it, but the demand for human-produced digital art will still exist. It will be challenging to filter and curate human art, especially as the line will be blurred, and many human artists will take advantage of AI. But I don't think any of it will entirely make human artists obsolete.
Anyway, this is all speculation from my side, so I concede that I may be wrong, but it's interesting to think about, and time will tell.
The ability to monetize the original voice (in the case of voice actors) now that AI can re-create it with convincing effect.
Take the humans out of the equation and it's no wonder the VAs get snarky.
Sure, inbreeding is a problem but once it patterns enough diversity, there simply won't be need of human voices paid for anything.
Modders going out of their way to find voice actors who can accurately reproduce an out of budget voice actor sounds like a labor of love that would not scale without serious money.
Having a gaming pc with a modern-ish nvidia GPU (not hating on AMD, rocm is just painful right now) scales very nicely when you apply ML to the problem.
Scales to the budget of bored tinkering teenagers with a lot of compute from chasing gaming framerates, gaming put mini ML workstations in a lot of kid's bedrooms if you think about it.
As of right now it is pretty much turnkey as it is.
This isn’t just about who has enough money, but about the ethics of using an actor’s likeness at scale without their permission and without paying them.
I think the fact that other VAs have been hired by mod teams in the past is only relevant if you believe that using AI is somehow equivalent.
Regarding copyright, it was previously not practical to copy a voice, so while it may be true that it is not copyright-able today, this is not a sufficient argument for the ethics of such copies, which may or may not be in sync with current laws. If copying voices was possible when copyright laws were established, the legal landscape would probably look different than it does today, which hints at the necessity of changes going forward.
AI is breaking new ground, and at this stage of the conversation, precedent is interesting in terms of how it reveals the gaps in current legal frameworks, but this should not be mistaken for such gaps being acceptable.
Many of these new conversations will reveal two things:
1) The things happening right now are legal
2) There are legitimate questions about whether or not they should be
Tom Waits successfully sued Frito-Lay for using a Tom Waits soundalike in a commercial.
Typically ‘W’ and Clinton —though the ones for Lincoln and Washington is obviously imagined and presumably any rights would have expired.
2) Obvious mimicry and impressions are okay. No one thinks it really is W or Clinton.
3) They are local ads, and thus no one notices
Here, modders are impersonating a character played by a voice actor and the copyright on the character belongs to the game company, not the actor.
Assuming the game company declines to enforce their copyright, it falls back to the actor to enforce their publicity rights. Which I'm not sure they can, since it's not them but rather a character they play.
When modders hired soundalikes to extend a character, it was equivalent to finding a similar looking actor. Now, using AI, it's equivalent to using deepfake video. Legally, I assume there's a difference which might be important if the game company decided to sue the modders. But from the perspective of the voice actor, I don't think it matters. Their copyright isn't being infringed.
On the other hand, you could just have different voices for the different characters, which a book narrator can't do. I thought about doing that for my books, but it'd just be way too expensive.
It's more than just following an annotation about sarcasm or sighing or sounding depressed.
But yeah, that probably IS going to have to suffice for the vast majority of popular works.
Elevenlabs announced a new professional audio model with a more involved training process, I'm curious to see how good that one is, and how much control it gives you.
One morning in May 1979, Dan and Janet were going over a problem with the latest build when Grant knocked, “Hi, I can come back later if you're busy with something.” Dan looked at Janet to see if he should leave. “No, come in. Dan was just being a jerk as usual.” Dan tried to look hurt. Grant said, “I'm all settled in now. You should come up.” Dan looked confused, “Wait… I thought you lived up North.” “No, I transferred down here to the Tor project.” “Wow, I thought people only moved from South to North.” “I'm starting a new trend.” “Are you up on the second floor now?” “ Yep. Come up and visit sometime.” “What a thought. Is it like here?” Dan had never been on the second floor or anywhere else in the building. Grant went to the door and made a show of looking in both directions, “Looks very similar. Anyhow, what are you doing for no-driving day tomorrow?” He looked at Janet, who had a quizzical expression. “The company just said we're not supposed to drive alone tomorrow because of the gas rationing. They're going to station guards at the parking lots to check.” Dan already knew about this, but Janet didn't. “Don't know. Maybe Dan can come by and pick me up. Then, we'd be a carpool.” Grant flashed a look of displeasure for an instant but quickly erased it. He was hoping Janet would come to his house, and they could walk to work. “Do you live near Janet?” “Hermosa. Janet's right on the way.” Grant recovered, “I was going to suggest we meet at my house and walk to work. It's only about a half-hour walk.” Dan asked, “Where do you live? Hawthorne?” Grant looked vaguely insulted at the suggestion that he might live in Hawthorne, “Manhattan Beach, near Marine and Aviation.” “We could pick you up, too! You're right on the way.” Grant wasn't expecting this resolution, but there it was. He wrote his address on the whiteboard, which Dan copied down into his notebook. They agreed on a time tomorrow, and Dan and Grant both left.
This is just a single generation pass though, and the voice is a bit unstable and is maybe not perfectly consistent with the direct speech, ideally you would tweak it a bit, maybe break it into chunks, isolate the direct speech and then cut it together, etc.
Still, it's pretty good I'd say, and as I mentioned previously, being able to give specific directions is around the corner, along with a "proper"/professional training of the voice model on a specific voice.
the most obvious thing to me is that Janet is female. It does seem to have different voices for Dan and Grant, at least.
You can see him and me here:
https://www.youtube.com/watch?v=LGchbES0DhU
the commenters are all his followers, I guess. I don't know any of them.
Voices aren't that unique, at least from the perspective of a human with average hearing. I believe that by analyzing every single frequency of a perfect recording, you can determine the speaker fairly precisely, but for a normal human ear, plenty of people sound "almost alike".
Here’s an audio AI example by ElevenLabs with celebrity voices.
It's been pretty wild the past few years playing games and realizing the difference between today and tomorrow AI will have on gaming.
Contract negotiation and licensing just needs to catch up.
I'm not surprised the gaming community is going hard and fast on AI voices, especially voice engines for the huge amount of text in lots of older games which would cost a fortune and a half to voice.
It will become so diluted that using human actors will be seen as special and not the other way around. Kinda like synthetic diamonds vs blood diamonds.