Text2LIVE: Text-Driven Layered Image and Video Editing
text2live.github.io
text2live.github.io
Almost anyone will be able to easily remix any type of content however they want.
I'm just waiting for the GIF keyboard that creates GIF's based on your prompt instead of searching through an existing database of them.
That will be truly next-level.
Although, I have access to DALL-E 2 and I have to say it's not as great as it looks. It is not productizable as is and couldn't replace an artist. It can't do text, the art styles change wildly and aren't controllable, image quality degrades with longer prompts, and most importantly their training data has everything vaguely interesting censored out of it.
Which is surprising since GPT-3 is very complete and a little too uncensored. Not sure if they think visuals are worse or they just have a lot of extra ethicists messing with it now.
That could be of people you’ve made deepfakes of, but it could be OpenAI embarrassed that their outputs look bad, or their users embarrassed that they got an NSFW output for an SFW prompt.
I think the censorship is also about copyright though; it barely knows any characters while dalle mini users are generating stuff like “Shrek on trial for murder” all day.
From earlier prompts I tried, “Minecraft” makes things look conceptually blockier rather than visually blockier, if that makes sense. So a “Minecraft mansion” will turn into a midcentury-modernist house with flat walls and roofs rather than being made of voxels.
That’s a pretty neat effect, but if you use names of video games in dalle or other image AIs they tend to start generating UI elements on top of your image and it looks bad.
Finally a way for the Asylum* to cut VFX techs and editors out of the loop altogether and allow producers to keep the good cocaine for themselves.
* the people who brought you the Sharknado franchise
Most AI is currently done by big, "serious" companies that both care about liability and bad press of being associated with that sort of thing, and also have a lot of AI-ethicist-type folks on board who care a lot about what they consider "misuse". Some AI designs try to limit the NSFW training data used in the model (e.g. DALL-E 2), while others try to fine-tune or censor results after the fact (e.g. GPT-3).
Right now the adult-oriented AI applications lag slightly behind the most cutting-edge ones, but actually make up a shockingly big percentage of the consumer base, both current and potential -- adult content is probably one of the biggest actual potential applications for AI, and there are some really fascinating ethical questions around it (e.g. ethics of AI-generated porn vs real life porn, considerations around real people, minors, other illegal content, etc.).
Generally, adult-oriented models are either hobbyist clones/finetunes of existing models, or just existing models that people have figured out ways to get to work with adult inputs. There are plenty of AI model hosting services out there that have no qualms about being used for shady or even illegal purposes, so it's difficult if not impossible to stop it from the server provider side.
We need to be thinking more about what how we want to handle that sort of thing socially/culturally/legally, because it's gonna happen whether we want it to or not.
That's cool and all, but also really stupid and a pointless distraction compared to how novel the underlying mathematics and science are. This will quickly become a commodity and humans will acclimate to seeing such tricks. The content produced won't even be considered particularly impressive.
Damn. I was hoping the singularity would be better.
I'd be all for it! Just not clear on a plausible path for how this better future comes to pass.
Though, a question is whether the good storytellers can be found easily? It seems like the situation is similar in fan fiction.
1. Their priorities are wrong so they are not the best minds
2. If (1) is false because the best minds can have stupid priorities, then The Best Minds is not the be-all-end-all of everything
Just think: when I ask you “walk in to the workshop, grab a hammer and a box of nails, and meet me on the roof to help me secure some loose shingles” your mind is already imagining the path you will take to get there, what it will look like when you locate and grab the hammer and nails, and you’ve filled in that to get on the roof you have to meet me in the back yard to climb the ladder, which I never mentioned.
All these tiny details your mind can do effortlessly take huge efforts like CLIP to sort out how to make it work. And even CLIP is only text and images. There is a lot more to go from there.
A lot of people focus on DALL-E and the artifacts that come out along the way, but these are not the destination, just little stops showing the progress we are making on a much larger journey.