Audiobox: Meta's new foundation research model for audio generation
ai.meta.com
ai.meta.com
The more rational voices in my mind, though, become more and more afraid of a world where the only thing you can trust is people sitting right in front of you. That makes the world of information pretty small again.
It's basically Dungeons and Dragons with an AI dungeon master who can generate video in realtime. Which would be awesome, but like Dungeons and Dragons it wouldn't be easy to keep the player on track.
Edit: just one more! Imagine actually having to complete quests in a given amount of time, because the rest of the world continues to revolve. People being mad at you because you left their children to die in the dungeon after arriving a day too late, because you were busy running an errand for someone else.
I think you might be able to do it with a lot of prompting, and having a database that functions like a wiki for the current and past states of the game world.
If you got really fancy, you could also make it a pseudo MMO where the content of the story you create with a character could be used as the basis for a NPC plotline in other people's worlds, possibly reducing the amount of content needed to be written.
If it got popular you could also use it as a research tool, where you could force some subset of the player population have some interaction, and be able to test and get a dataset for counterfactual reasoning.
The next 5 years will be wild.
I’m hoping that like digital instruments, I’ll be able to splice in digital voices instead of finding singers.
Definitely recommend looking up some Synthesizer V covers to see how realistic singing voice synthesis has become in recent years. It’s also free to try lower quality versions of the voices.
RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.
"The RVC model is a Retrieval-based Voice Conversion system using AI for high-quality voice cloning. It utilizes artificial intelligence to modify or clone voices in real-time." Source: https://speechify.com/blog/rvc-vocal-models/
Would make a nice vignette in a film about a dystopian future where video can be be generated cheaply and of sufficient quality.
AI is still too expensive and performance intensive to run in games cost effectively, and truly powerful AI is probably another 10-100x cost increase.
On the other hand, novels will be rapidly replaced by visual novels. The cost of having a novel fully illustrated and voiced will go down 1000x. A high quality illustration used to cost $500-$1000 (A day's work from a high-tier commercial artist), soon it will be about $0.5. I'm not counting in the author's time to prompt the images, because it would have costed them way more time to communicate with the illustrator anyways.
The entire boundary between novels, comics, cartoons etc will blur. Like if a newly written Harry Potter can have thousands of illustrations set in Hogwarts and be fully voiced, the standards for a movie adaptation will be astronomically high, which will in turn drive AI use in movie production just to keep up.
Maybe. But I think you make the mistake of considering games that combine existung AAA features + AI as where it will first impact games, where I think it will first make its mark in games that don’t use hardware heavily for 3d rendering by opening up new modes of gaming.
> novels will be rapidly replaced by visual novels. The cost of having a novel fully illustrated and voiced will go down 1000x.
The cost if having art made isn't the only reason novels aren't fully illustrated now, and AI doesn't impact any of the others.
“I’m sorry, but as an ethically trained AI I cannot engage in this sword fight. Violence is never the answer.”
Yeah. It’s gonna be very immersive :P
I could see Microsoft making the first move next generation since they're knee deep in it.
Maybe soon game developers will realize black holes and the uncertainty principle are great for efficiency.
Gives me cosmological analogies: in the far enough future, the only things you can see in the night sky are the members of the local supercluster.
I'd imagine a mixture of general prompting/training the LLM on techniques like those used for plot progressiom by human game masters in TTRPGs and guidance via systems tracking progress and injecting contextual prompts based on mechanisms like those used for tracking and guiding plot progression in GM-less/GM-replacement systems (e.g., the Mythic Game Master Emulator) for TTRPGs.
The AI-Voice revolution leads us instead to the problem of authority, and of imitation of authority. Too many of us take it as given that an authority has the truth.
The AI-Voice / AI-Content revolution allows low quality actors to imitate relevant authorities.
So, to address your question, we need to study philosophy in schools, so that kids learn to think for themselves.
There's definitely some gamers who would like to never talk to a character in a video game, even if it means hours of bumbling around for the blue key that some NPC just told me about while I ignored 100% of the text in the game.
"a world where the only thing you can trust is people sitting right in front of you"
Oh and I think real people can be a source of missinformation as well. Also holograms(or brain implants) might have a breakthrough soon, as well. So all in all I think we are living in interesting times.
Yeah, definitely. But at least you have a chance to detect their intentions; for some forged digital information, you don't even have that. "Interesting times" is one way to put it...
You could care less if the content you already consume on forums like HN were generated by AGI. Just sayin.
Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understand is that the model weights have to be swapped or altered per se to be able to commercially reuse this. Correct me if I'm wrong.
and important, if you have more than 100m active users:) 4. Restrictions If you are commercially using the Materials, and your product or service has more than 100 million monthly active users, You shall request a license from Us
So, looks like it's absolutely fine to use, except for IT behemoth.
As for how to use, API, I think. Interesting applications are possible. Like interactive mobile robots. Assistants for people with disabilities, both software and wearable.
Interesting times... this will be called AI revolution probably. It's already not a joke, after several ups and downs.
Openly distributed perhaps, but definitely not open source. The license appears as closed as meta (research-only with some leeway for other uses.)
Do you have any truely open-source general audio generation models yet?
I know about StyleTTS2, which is open source (MIT) and uncensored, but that model focuses on speech generation only. Having an non proprietary model like audiobox or Qwen-Audio would be really nice.
This technology will present serious challenges for the verification of covertly recorded audio. It will of course ultimately become widespread but I’m not inherently bothered by the idea of slowing down its release. Giving researchers extra time to examine possible detection techniques seems helpful to me.
There will be people that will spend almost every waking hour with one of those things attached to their face if they can also make this device lightweight and comfortable
It's not clear as to what the expected outcome of this 'responsibility and safety research' effort is. Is the idea to nerf the tech such that it can't be used for morally/ethically nefarious purposes? If so, then is the "speech research" community the group best fit to do that work?
Edit: nvm, seems not from this line "In the coming weeks, we will be opening up the application here, along with an interactive demo that will showcase Audiobox’s capabilities."
Does Meta usually provide a web interface for them or do you have to download and run locally?
A long time ago, there was a great story in a Shadowrun supplement of all things about a hacker that got trained to teach an "ai" how to break into computers. It was basically at a child's level, emotionally, and the only world it's ever known was "the matrix" (yes, really -- and written almost a decade before The Matrix came out). Eventually it turns out that it's not an ai, but a corporation was stealing kids and sticking them in Virtual Reality at birth to train a team of super hackers.
1. No TTS audio output is tamper-proof. Their "safeguards" will be busted, and quickly. Whether via a small adversarial NN, some basic DSP, or just...holding a cheap recorder near your speakers, maintaining audio file provenance has no chance.
2. Impersonations have vexed humanity since the invention of vocal cords. Insofar as it's soluble, it's been solved -- authenticity is determined by a fluid mixture of context, trustworthiness, and the authority of involved parties & institutions. Always has been. Always will be. If I could drill one idea deep into every tech evangelist's head, it'd be: The solution to every problem isn't automatically "more technology." But hammers see only nails, so the vicious cycle continues, and society deals with the consequences (e.g. cryptobros decentralizing money...by slowly reinventing banks, but with more fraud).
3. This secret audio ID "feature" is probably harmful. It adds needless complexity. At best it exacerbates a false sense of safety because impersonation is trivial. Bad guys can emulate it on authentic recordings to discredit them as "fake." Nobody who'd actually benefit from such safeguards will respect them. News says this audio that affirms my confirmation bias is fake? Nah, the news is fake.
Meta knows all of this. Optimistically, I hope it's just lip service to concern fetishists; plausible deniability for the knife manufacturer when a bad guy uses one. Pessimistically, it might be pretext for an about-face on their OSS commitment. "Oops, researchers trivially broke our safeguards. Shucks. That's scary. Guess we'll build a moat instead of an OSS community. Think of the children or terminator or whatever works these days"
I suppose we'll see.
Support open source models by celebrating their release and pressuring companies to release them, and oppose closed source AI or face a very bleak future for you and your descendants.
You may be having fun with “Open” AI’s API today, but you’re supporting and celebrating the collapse of society into megacap AI elites and a majority paying for metered access to old technology.
"I just hold on to all the money, 'cause bitches can't be trusted with it. We pool all the kissing money together, see? But if you wanna buy anything, you just talk to the bottom bitch, and then the bottom bitch talks to me.
Do you know what I am saying?"