Generating audio for video
deepmind.google
deepmind.google
But I literally can't keep track anymore of which AI generative combinations of modalities have been released.
Crazy how two years ago this would have blown my mind. Now it's just, OK sure add it to the pile...
Subscribed to GTP-4o (or whatever the paying one is called) for translating / finding typos / summarizing / etc.
Zero brand love and I'll switch to something else (maybe some future Claude model?) the second something better/faster comes out.
Tons -- and I mean tons -- of people have spent money on it. Because it's worth it, it's generating actual economic value for them.
It's also exceptional at making IEPs/Learning Plans for certain things I'd like to learn for the week etc which I am already somewhat familiar with. I use it as a rough guide and it has worked well so far.
https://docs.google.com/spreadsheets/d/1O5KVQW1Hx5ZAkcg8AIRj...
And here's some that I personally recommend and are "free" to use:
TXT2VID / IMG2VID: https://lumalabs.ai/dream-machine
TXT2MUSIC: https://suno.com/
AI TXT2SPEECH: https://murf.ai/
PDF Summarize (You can just use 4o or Claude though: https://askyourpdf.com/
AI ChatBot: https://janitorai.com/ https://www.chub.ai/
TXT2IMG / IMG2IMG: https://playground.com/
Obviously SD 1.5/SDXL/Pony
and so much more.
Using a recommendation algorithm similar to TikTok’s, learn what each specific user are into, and instead of showing content produced by other users, produce custom-tailored content on the fly, perfectly matching the type, tone, style, length, and rhythm each user likes.
Ideally without making anything up.
How is TTs recommendation system different from YT? Other than suggesting lower quality content that's irresistable?
TikTok seems to manage to more quickly identify users’ interests and surface content based on more signals, aggregated over a longer period of time, without relying as much on conscious users’ actions (ie "follow / subscribe"), producing a wider diversity of recommendations.
There’s also the odd suggestion every now and then, probably used to gauge a user’s interest in a different category.
I have no idea anymore if this is sarcasm or a straight up belief.
What serious professional would gamble on hallucinations?
The only point of these kinds of platforms, for worse and for worse, is to give users what they want. So hallucinations wouldn’t matter, as long as the end result matches users’ preferences.
I am not implying this is a good thing. Or a bad one. It’s just a step further down the same path we’re already on, while taking an unreliable and costly middle-man (content producing users) out of the picture.
I’m just certain it is an obvious next step given, on one side, these platforms’ goals and incentives, and on the other how generative AI capabilities have progressed in the past couple of years.
I’m pretty sure the smart play, if they want diverse, engaging, and surprising content will involve leaving some room for people to create things and somehow reward them for it.
But whatever they make won’t only be used as content to show to others, but also as new training data to feed the machine.
Try imagining this concept applied to newscasts.
That said, storage is far cheaper than GPUs at the moment.
If the sound is already being generated at a specific time, surely you can make it generate an output that can be consumed by existing audio mixing tools for further refinement.
The problem with doing these all-in-one integrated solutions is that you're kinda giving people an all-or-nothing option, which doesn't seem that useful. Maybe I'll end up being proven wrong.
This one is particularly annoying as I worked for years as a sound engineer and have recorded or produced the soundtrack for 10 feature films and some large number of shorts. What's going to happen with this is directors or producers are gonna do this at home for every scene in a burst of over-enthusiasm, realize the totality is Not Great, and then demand someone like me fix it, but for 1/4 of what the job used to pay, arguing 'but most of the work is already done'. It's all so tiresome.
Unless they can have enough raw, unmixed sample, this depends on how well they "unmix" them.
Totally and that is 100% what is coming. For a great many pictures too: why generate a picture full of lightning issues / approximation when you'll soon be able to generate and entire 3D scene and render it properly.
We've mastered 3D rendering and audio engineering.
I want the 3D models and the 3D scenes. I want the individual tracks (and combine them in Dobly Atmos or whatever shall be cool).
And that is coming, no question about it.
That'd interest me (a musical hobbyist) more than the "whole track" generators, for sure.
I imagine it's a harder task tho'. Presumably, if you give the same source material (video, prompt) to the AI multiple times, it will generate different pieces of music. So if you do a series of prompts, each one specifying a different instrument or group/bus, then you (or the AI) need to arrange for the parts to blend correctly, follow the same cues and assemble to a coherent arrangement. Is that one pass with multiple outputs, or multiple passes/prompts with one output each?
I have got the impression (from casual reading) that the music generators don't inherently "know" about different parts of a piece of music. They just know about the final output.
https://www.youtube.com/playlist?list=PLQvwVDViTLXu4usHto8PH...