Make-A-Video: AI system that generates videos from text
makeavideo.studio
makeavideo.studio
A fun related experiment, I thought it was fun to see what kind of movies AI would generate, so I created a "This Movie Does Not Exist" website[1] that auto generates fake movies (movie posters + synopsis). It basically uses GPT-3 to generate some story plots, and then uses that as a prompt (with in-between steps) for Stable Diffusion. Results may vary, but it definitely surprises sometimes with movies that look and sound amazing!
[1] This Movie Does Not Exist: https://thismoviedoesnotexist.org/
The best chess seems to be when AI is used along with humans. I think image and video AI will best be exploited when human input is also taken into account.
There is still something special about human creativity, I think AI will just be another tool to expand that. At least, in the short term I would say 10 years perhaps. AI will probably one day take over all aspects of creativity and humans won’t be able to contribute.
Poker on the other hand I think human players still win vs GTO solvers, but again I may be mistaken here too.
Also an outsider, but I think this has changed in the last year and that AI now is consistently better than top-tier humans at even no-limits poker.
https://www.sciencedaily.com/releases/2019/07/190711141343.h...
https://www.cs.cmu.edu/~noamb/papers/19-Science-Superhuman.p...
I don't think this is true anymore. I don't think I've heard about successful centaur chess games in years. I would love to be wrong there though (in particular if anyone knows about how correspondence chess games have been played in the last 2 or 3 years with the availability of Leela Zero and Stockfish NNUE).
Based on that thread, it looks like centaur chess is close to dead.
> Human input in top-level ICCF [correspondence games with chess engine support] games is now 99% eliminated, other than personal preference in selection of openings.
(and then regrettably irrelevant thereafter).
I think it is a legitimate worry, as the pace of progress is considerable. These tools are impressive, and are only going to get more impressive: more people should be talking about where this is headed.
Even still, if the EU chooses to regress, I can imagine a lot of non-EU-based companies—especially smaller ones—just choosing not to deal with them anymore
EU would ban any technology that disrupts the labor market to an extent it leads to widespread protests among a segment of society. For how a somehow similar thing has played out in the last 40 years (instead of banning technologies, it has mostly been instituting quotas, paying subsidies and banning/taxing imports), you can look at its agricultural policy.
Incidentally, Uber has been banned in the entire country where I live for 8+ years now.
But what a government can do, and what its politicians will do are very different things. The political backlash for going nuclear on corporations would be massive and sustained. Not even good acting corporations want that kind of thing happening.
Regarding, Uber and your country: your point makes my point stronger.
A whole country is now off limits to Uber, but on the whole, across all the countries, their aggressive behaviors has been a win for them anyway. Just another cost of business.
If one trained using e.g. a tiktok like dataset showing viewer response measurements for each video, and do conditional generation on those response values ("prank video watchers are highly likely to watch the full video"), are we really that far from a system that learns to generate content that attracts and hold eyeballs? Not so long ago there were a lot of concerning trend pieces about how youtube had a network of creators making bizarre, disturbing or transfixing videos being watched entirely by young children. Before that, it was clickbait listicles. "Bad" content that can get eyeballs can still wildly steer what humans create and consume. I'm wondering if in 2 years we'll have an enormous number of short videos that we all agree are "bad" but which are nevertheless constantly watched.
AI is the winningest in chess, but the real life purpose of chess is to produce interesting gameplay for people to watch, and so AI is less good than Magnus at that. You’d need the AI to throw games and write press releases.
"Adam Sandler is like, in love with some girl, but then it turns out that the girl is actually a Golden Retriever. Or something.""
A South Park classic imo
I agree, and I think when that happens, it will tend to increase the value of curation. High quality curation that is, probably done mostly by hand, as opposed to the at-best-mediocre automated curation that is commonly used.
It could be bad for things like YouTube, for example. I think there will be an arms race between generated video content one one side and automated curation on the other. I mean, you can still leverage viewer choices for curation (looking at what people are watching a lot of), but that is just shifting the burden of curation to users. Few people will be willing to sift through dozens of cheaply generated crap videos to find something they actually want to watch.
I am sure I will watch MY 15 year old's attempts, and maybe a few from my extended circle but most of my consumed content will still come from what makes the cut to Netflix or HBO etc. Technology like this will empower the truly creatives once it has matured. I would expect closer to 20 years than 5 however.
Ah, I can see what's wrong there, just turn down shit randomness and increase the true gem sampling steps to proportionally increase the input weight of true gems and fix your output quality.
Basically - if it empowers creative people, their output will be fed back in and parameterised.
I don't know what artists, truck drivers, Uber/Lyft/taxi drivers, delivery drivers, programmers, doctors, judges, fast food workers, etc. are going to do.
Surely if we get to the point where programmers are no longer needed, humans have essentially been replaced by AI? Since the AI programmers could just program better and better AI?
Under the current system, the rich can do just that while everyone else literally starves to death :)
https://www.theatlantic.com/magazine/archive/2018/10/yuval-n...
The idea that rich people will all leave and start a different rich-people-only economy that somehow takes all the economic activity with it isn’t how it really works, it’s the plot of Atlas Shrugged.
https://thismoviedoesnotexist.org/movie/the-terminator
Brings up the age-old question of how much the learning in these models is just memorization. Though in cases like these it’s hard to tell.
It's crazy that it just made up those names...
Why send a soldier from the present to the past when you can also have a soldier from the future send to the past!
I know what you mean, but I also laughed at that.
"I bet one legend that keeps recurring throughout history, in every culture, is the story of Popeye." - Jack Handey
https://thismoviedoesnotexist.org/movie/in-the-land-of-oz-th...
Which reads like a bad translation of a bad translation. Like the the old joke about the AI program which was supposed to translate "The spirit is willing but the flesh is weak" from English to Russian to English, and after the roundtrip came up with "The vodka is good but the meat is rotten."
Unfortunately, that's already happening.
https://www.youtube.com/watch?v=w7oiHtYCo0w
From what I can see, YouTube has done quite a bit of work to cleanup YouTube Kids, but it's kind of an arms race.
There's this worrying issue in AI ethics discussions where most people seem to assume the problems and dangers of AI are still off in the future, that as long as we don't have the malicious AGI of sci-fi stories, then AI and "lesser" algorithmically generated content isn't harming society.
I think that's not true at all. I think we've seen massive damage to social structures thanks to algorithmic feeds and generated content, already, for years now. I don't think, just because they aren't necessarily neural-network-based, doesn't make them something to not worry about.
So I don't see AI as a particularly, different, worrisome problem. It's an extension of an already existing, worrisome problem that most people have ignored beyond occasionally complaining about election results.
It's scary to think about it but seems plausible—like if someone can make an app with Tiktok-like ubiquity of only AI content. Although to your point I imagine there will be so much nonsensical noise that curating will become a useful skill, it is today but even more so.
I really don't think it applies to us in this context though because I think that a decent number of humans don't care whether some content is AI generated. Furry porn is all hand-drawn and people still like it despite it not being real.
This just gave me a disturbingly vivid vision of exactly what you described. A seemingly likely future where everyone's in their homes scrolling endlessly on a TikTok-like app where there's literally infinite content being generated by the AI all the time, and as people like and dislike certain types of content, the AI just gets better and better at generating new videos...this is honestly kind of terrifying. I have no doubt this will exist one day, and that it'll print money for one company while billions of people are spending all their free time consuming it.
https://twitter.com/dvorahfr/status/1575508907593711618?s=46...
The description text does not convert sentences with carriage returns (or probably, newlines) into separate div's or whatever html element you'd prefer, FYI! Otherwise, very cool!
Edit:
Worst (best?) tone clash: https://thismoviedoesnotexist.org/movie/stalked-by-a-friend
Great work with this as is!
https://github.com/snap-research/NeROIC
https://github.com/threedle/text2mesh
From what I understand with NeROIC, it's not particularly meant to be able to generate an 3D model that can be imported into Blender (or other software). It requires more work to take the meshes it generates to do something with it. See https://github.com/snap-research/NeROIC/issues/10
I too was looking into it to generate 3D models for some software I've been working on.
Let the AI generate random blender files.
Render them.
Train on source => render mapping (which is 1:many)
Repeat.
You should be able to get an incredibly high-fidelity decoder to go from image to blender source.
Then you're in luck today as this was just submitted: DreamFusion: Text-to-3D using 2D Diffusion
The paper talks about "pseudo 3D attention layers" that are used in place of temporal attention layers for each dimension due to memory consumption. It seems like AI research is vastly outpacing GPU development.
Even then, these videos are only like 50 frames long - and a real movie you would want to be hundreds of thousands of frames long.
We can’t do it. AIs can sort of do it.
Latent diffusion models already demonstrated that operating on a compressed representation gives far better results, faster, but I don’t think we’re anywhere near the limit for what’s possible there. It’s no coincidence that this is how humans work.
They put an eye tracker on someone and captured their motion when walking in some rough terrain. You can sort of see that the person is focusing on the most likely place their foot will go next.
[1] https://www.youtube.com/watch?v=ph6uUHq3a-g
I think that we will discover that there is a more efficient way to encode temporal relationships, which appears to be "just throw transformers at it." My guess is that it will be in a more conceptual latent space that this attention will be applied.
Yes, but consider that most films are made up of many different shots, each of which are often just seconds long.
Obviously you could do 'human assisted' movie making where humans decide the storyboard and make directions for each shot, and then that isn't necessary.
It's a good thing to be fair, forcing research teams to optimize their projects is beneficial and creates a competition for limited resources. This gets a bit skewed when we consider a university research team vs. a MANGA type company, but the team behind Stable diffusion proved that innovation can come from unexpected places.
What's interesting to me is how this is so similar to human imagination. Give me a description and I will fabricate the visuals in my mind. Some aspects will be detailed, others will be vague, or unimportant. Crazy to see how fast AI is progressing. Machines are approaching the ability to interpret and visualize text in the same way humans can.
This also fascinates me as a form of compression. You can transmit concepts and descriptions without transmitting pixel data, and the visuals can be generated onsite. Wonder if there is some practical application for this.
>https://makeavideo.studio/assets/A_knight_riding_on_a_horse_...
The horse's face is all wrong
The gait is wrong
The interface with the ground & hooves is wrong
The knight's upper body doesn't match with the lower and they're not moving correctly
I think ultimately the right path is something like AI automated Blender. AI creates the models & actions while Blender renders it according to a rules based physics engine.
It doesn't seem that the fundamental inability to understand what is going on in the scene is stopping models of this kind to eventually lead to realistic results.
Same applies to DALL-E and GPT.
It's superficially close but when you look at details they're all slightly off. To wit:
> the knight's body absorbs the force of stomping on the ground…
But the Knight doesn't have any way to see through the helmet.
And I was surprised about how small the wiggle room was the first time I interacted with GPT in text or saw the first images from DALL-E, since I too expected them to be (severly) limited by not understanding what's actually respresented in the input/output.
With new versions the wiggle room shrinks further. So I guess the question is, whether it will be able to shrink enough to be satisfactory. We will see…
The downsides: there will be less money out there for creators, because that becomes a commodity. You will be able to make money if you are known for quality content polishing, editing and generally bringing that last 5-10% of generated content to look perfect (all the way until automated tools are trained to do that as well). AI will automate and improve most white-collar jobs. Instead of generators, everyone will become curators, as taste will be more important (until that's trained into the system as well).
For deeper levels of consequence we have to look at history: how did the world change when people finally got paper after most writing was done on animal skins (more got to write, the richest or most powerful didn't have the only say), or water piped to your house after you had to carry buckets and dig wells (It freed up time for everyone for more interesting tasks). Now GPUs are going to be the new paper and the new PVC. Yes, software has been eating the world for a while, but you won't be able to brainstorm without AI generating the first pass.
The idea of a blank slate creator has been dead long before ML tools were introduced :)
Examples: https://make-a-video.github.io/
Demo site: https://makeavideo.studio/
I am told live demo and open model are on the way.
Though I must admit that if I didn't have friends holding my hand through the minefield of modern cinema, I would also just stick to books.
Why spend money on a film with new IP and ideas that you're not sure will be popular when the data science team has already worked with marketing to figure out exactly what movie will sell well?
Good luck finding your movie with compelling and thought provoking writing in the big pile of movies produced by comittee to show up above yours in discovery algorithms.
You want a thought-provoking Bourne-style action thriller with hints of Jane Austen and a bollywood dance sequence? How about a Matrix sequel that lives up to the first one, but ends just the way you like it? Just ask.
Or how about the movie Clue with the three endings except an infinite sequence of "or maybe it happened this way" sequences. I mean how else are we going to get a sequence where Darth Vader and Tim Curry reenact the "No I am your father" scene while Martin Mull dies in the background due to a heart attack.
You could conceivably write a script and feed it into a machine and have a decent 1080p rendition of the movie with consistent characters and voice acting which you could use to better pitch your movie idea to people, or get to watch a movie you created between you and your friends even if no one else ever gets to see it.
The current trend of remaking movies as 10 hour miniseries (and then making more and more seasons) is Not Great. Whereas I could be fascinated by a quirky-but-compelling original movie, I'm less attracted to 10+ hours of hyper-polished content. Sometimes I've watched a series and thought 'that was good, but it could have been a better movie.'
In the future, kids will be making their own Star Wars movies from home. All kinds of people from all kinds of backgrounds will make novel films that would ordinarily never be made, such as "Steampunk Vampires of Venus", starring John Wayne, a young Betty White, and Samuel L. Jackson. This is absolutely the future.
I'm working on building this. I'm sure lots of others are too.
We're in a situation where the very best algorithms (like the one used by Netflix), doing exactly what they're designed to do, create inequality and the vacuous economy of influencers we have today. Look at Steam or any other marketplace: they're all the same, with 1% of the players getting 99% of the prizes. In a very real way, the only winning move is not to play.
I would suggest that this tendency of capitalism (economic evolution) is unstoppable, and that it must be attacked from a different angle. If we don't want to inevitably end up in late-stage capitalism that looks like neofeudalism, then there has to be some form of redistribution or people spend the entirety of their lives running the rat race to make rent. Traditionally that was high taxes on the winners, but UBI would probably work better. Unfortunately, the very same people who win are the ones most resistant to any notion of a level playing field or social safety net.
So I feel like there may be no solution coming. We're probably looking at long slow decline for the next 15 years or so until AI reaches a level where economics don't really make sense anymore, since economic systems by definition control the distribution of resources under scarcity. Without scarcity, they're pointless. And we moved into the age of artificial scarcity sometime after WWII, probably in the late 1960s, but certainly no later than 1990 with the fall of the USSR and the rise of straw man enemies like terrorism, using divisive politics as the primary means of controlling the population. Noam Chomsky saw this coming before most of us were born.
In other words, when anyone can wish for anything by turning sentences into 3D-printed manifestations of their dreams, then artificial scarcity quickly loses its luster. Because the systems of control around dependency no longer work. Then a new fear-based enemy comes along to fill the void, probably aliens. I wish I was joking.
I built a quick search engine over that data:
https://webvid.datasette.io/webvid/videos
Wrote more about that here: https://simonwillison.net/2022/Sep/29/webvid/
It's pretty wild to see how quickly the space of Generative AI Media is coming along.
I started a newsletter on the topic, called The Art of Intelligence (GPT-3 came up with the name) with the first post going out last Friday on the topic of how far are we from AI generated videos, and simulated worlds like the Holodeck given the rapid progress of these visual A.I. Thought y'all might find it interesting: https://artofintelligence.substack.com/p/dall-e-stable-diffu...
This type of progress also reminds me of a really lovely publication from 2017 in Distill.pub, on the topic of these A.I. enabled creation tools - I think y'all would enjoy seeing what folks were thinking even then: https://distill.pub/2017/aia/
Are GPU vendors (well, gpu vendor, as far as I can tell) focusing heavily on increasing VRAM? My understanding is that transformers are pretty quick to train, but have significant memory costs.
When they say that video is infeasible with memory... does that mean that if we had enough memory (128? 256? gb) we would be able to realistically train such networks with temporal attention?
This is insanely exciting. It looks like we are limited, at this point, only by compute.
Eg. https://www.youtube.com/watch?v=r_0JjYUe5jo --> https://vocaroo.com/1hgjjnVNqWjk
We're also working on film generation.
For some reason, unacceptable uncanny sounds is a much wider valley than unacceptable videos/pictures. The hand holding is uncanny in the family video, but I'm fine watching it for a second - it doesn't cause pain the way that same error would in music.
However, a creative system curated by a human could end up creating useful outputs, could it not?
Something that can be further refined by humans is more interesting. There's people looking into AI-based sample generation which is a lot more promising than full song generation, IMO.
Exactly, that is what I was trying to say. The way I look at it is that most people who have Ableton installed cannot create an amazing song. Now let's say they are able to prompt a Stable Diffusion Audio system with a prompt like kanye type beat with flute melody in the key of E.
The system might output 90% hot garbage, but it's easy to skip that within seconds of hearing it. So they clip and loop the good part, add whatever personal skills they do have, and upload that.
And wow, I just found out that OpenAI's Jukebox[0] was creating this stuff two years ago. This seems like the lowest hanging fruit to me, compared to visuals. Also could be extremely lucrative. I wonder if we are already listen to ML generated music and it's just not advertised?
[0] https://openai.com/blog/jukebox/ related post: https://news.ycombinator.com/item?id=23032243
Artists don't always become famous or popular because they're the best, instead it happens because they're pliable in a business sense and fit into a bigger picture of what the product is supposed to be.
I'll give it two years, at most, until we get pretty good audio and video generation from AI.
Have they? If anything, the past few months has shown that after the initial hype dies down, people find the limitations very quickly. It was only weeks ago that people here and on Reddit were proclaiming that graphic designers and artists were no longer needed. But while Stable Diffusion is great at making somewhat surreal images of Jeffrey Epstein eating an apple in the style of Picasso, it's not very good at, say, making a sleek, modern user interface mockup for a bank's login portal. And it turns out graphic designers are currently only paid for one of those things.
You could finetune the Stable Diffusion model to generate UI mockups and would receive much better results.
Imagine a future of Prompt Wizards, who are able to coax the AI to generate things in a very specific way.
Although we would probably need a much greater level of human curation. The way algorithms curate on youtube and spotify just doesn't really hit the spot.
Perhaps stability and Dall-e already kind of showed that the value is not so much in the physical act of creating, but more-so in the ability to express something that the AI can represent and which can connect with you.
It's AI written blogspam, but for images and video too. The signal to noise ratio is getting worse and worse
Every creative platform is going to be flooded with the equivalent of a Reddit comment.
(yes I am aware of the irony)
Maybe, if memes are the peak of creativity. (And who am I to judge?)
If we look at static images, the bar for distribution is zero and bar for creation is near zero since the arrival of these new AI tools – though it was circling zero before that.
And what seems to be most widely shared is memes, in my feeds anyway. When Stable Diffusion landed, people giggled about the president of my country rendered in the style of Grand Theft Auto. After a week of that my feeds went right back to memes.
Every human with an internet connection can draw like Picasso now, but it doesn't seem to matter. Because what we mostly seem to want is to take part in a conversation and get some validation, it seems to me.
After Neuralink and a few other companies torture enough poor monkeys, they eventually figure out how to create high bandwidth brain computer interfaces.
We go through a few more paradigm shifts in computing and get 10000-1000000 X or more performance increase. Metaverse protocols have advanced to allow for seamless integration of simulated persons and environments across multiple clusters.
The software continues to improve.
What you could get is a simulated realm with simulated AI characters living their own lives. But groups of real people are plugged directly in to the simulation and can influence it with their thoughts. There may be some sort of rules to ensure a certain level of stability. But basically you just think "there should be a storm today" and maybe visualize some strong winds. And then it happens in their world.
So at that point we become Gods.
I am probably getting carried away because I am tired.
if you read “France declares war on Canada”, you’re not gonna believe it unless it’s coming from an extremely reputable source. so why would you trust a random unsourced video?
the absolute worst thing that’s gonna happen is that video-based social media is gonna be flooded with low (or even high) quality AI videos. and I ask you: who the fuck cares? are these places doing wonders for society as it is? what’s a bit more rubbish amongst all the rest?
I can think of many, many more upsides than down
The fear is real and only seems fantastical because life is often stranger than fiction.
Instead of training on data these AIs will soon train on "creativity" and these layers of containerized thought will merge.
To do this it would need to want to live and know what fear is. It's just a piece of code, without being conscious there is no need to be online/alive for it. What people call AI today is just a few/hundreds GBs of data that reacts to some text and push it through a heuristic like system to get an answer doing fancy pattern matching. It doesn't do anything when there is no input, there are no processes, no thinking, no anything, it's "dead". Fear and instinct is not taught in any species, it's inherited prior and biological process, you would need to explicitly code feeling into AI for it to be able to "feel" anything.
A lie that is repeated a thousand times becomes truth. We are not talking about one out of place, weird news that would appear once on someone's newsfeed. We are talking about mass flooding.
> unless it’s coming from an extremely reputable source
It's 2022, no one is verifying sources
>verifying sources
I’m talking about if your news came from Reuters or the BBC compared to coming from realnews247.io or a facebook post. the same applies. if you’ll believe some text on the BBC, you’ll believe a video from there, and if you’ll believe the words of AnonTruther on facebook, then you’ll believe their videos too. this makes no difference
I mean, there are some theories that the mental illness epidemic is partially caused by internet use. I did strict 2 months Internet fast and I can attest it has a healing effect. Does it have entertainment-to-death effect? I guess there is a bigger than expected number of suicides that wouldn't happen have the people not been addicted to Internet. But I agree with every passing year we are moving further and further into entertainment-to-death with our civilisations.
Makes me kinda sad to work in tech industry.
> A golden retriever eating ice cream on a beautiful tropical beach at sunset, high resolution
example is terrifying.
But actually, this technology is super exciting. Imagine a future where movies and games are choose your own adventure.
Besides just textural content, it's intriguing to consider the possibilities of full-3d roguelikes.
https://makeavideo.studio/assets/a_golden_retriever_eating_i... (webp)
That grasp though.
These things still feel a bit like e.g. Google/GCP services to me: Super appealing at first glance, quite close to what you want, but somehow never quite there. Maybe they'll asymptotically get there, eventually? Perhaps that statistical model can't really make it to the level we want it to?
Never say never, we've come a long way since GPT-2! All this was unthinkable back then
Image-to-image and tuning already addresses many of these issues; just as inpainting works really well, it won't be long before we have select-and-repair, where you add an additional prompt like 'improve this part - the ice cream is fine, just work on the dog's muzzle.'
[1] https://makeavideo.studio/assets/A_knight_riding_on_a_horse_...
"Our goal is to eventually make this technology available to the public, but for now we will continue to analyze, test, and trial Make-A-Video to ensure that each step of release is safe and intentional."
Are they really going to do a replay of OpenAI and Stable Diffusion? Deja vu coming soon.
My daughter was born in 1998. The Internet made a bit of a phase change during her early childhood and there were many instances where there was no precedent to guide a decision on what to do, so we just had to wing it as parents based on our own intuition and values. I certainly got some things wrong but we made it in one piece. I just worry a bit about what kind of new challenges my daughter will face in raising her son in this new environment, where powerful organizations wield unbelievable resources to create deep, life long co-dependencies on their revenue streams and intellectual empires.
I can't imagine how poorly people will speak and mentally process emotions in 20 years.
If you took an average working-class, "blue collar" person from hundreds of years ago, they would speak very differently from how those books are written.
Surely you can see there is a sharp drop off in vocabulary and ability (or desire) to convey complex ideas.
How many of these books are representative of how average people spoke then?
This will absolutely snowball easy videos all over social media such as TikTok, Insta and YouTube.
Some of these look like existing videos that have just be garbled - check the clownfish one and the litter of puppies - maybe that is because the prompts aren't that detailed or they just need to up the "creative/randomness" factor.
The sloth one though has a more specific prompt and came out looking more better? and more original.
With the huge amount of data being used to back these, how do we know the "uniqueness" of the generated content? is it original or is it just mangling of existing content?
Even if it is just garbling existing content, it is pretty amazing. Frame by frame the mutant unicorns float across the beach hah.
The shortcoming is that, since they don't use any video/text annotation (just image/text pairs), complex temporal prompts would probably not work, something which require multiple steps like "a bear scoring a goal, then doing the Ronaldo celebration"
Fascinating times, either way this is allegedly being open–sourced, so I'm hopeful others will be able to build on top.
I'll just say it now: this is a mistake, quite possibly a huge mistake. The average human is not intelligent enough to deal with computer-generated video that they can mistake for reality, and so this can and will become a tool for despots.
People will end up watching news about events that never happened or product reviews for things that don't exist.
It's like self-driving cars. They use almost very effective statistical models, certainly better than our previous models, but they never seem to shake off that "almost" and become truly effective.
Any more specific reason why you feel this way? Curious
That's a bold prediction. Why do you think that?
The first thing I thought was the exact opposite. This isn't very good, but it's only version 1. Motion pictures are less than 150 years old. In another 150 years I bet virtual filmmaking will progress a lot.
For the record, I'm actually rather bullish on self-driving cars. There's nothing physically impossible about solving the problem, but I'm not surprised it's harder than it sounds. But I don't see humans being fundamentally prevented from solving the problem in the same way that humans are fundamentally prevented from ever engaging in everyday space travel.
It's not hard to imagine that this kind of thing could end up doing a lot of the heavy lifting for things like background scenes in the future, opening up the kind of stuff we saw in The Mandalorian, Game of Thrones, and the LOTR film trilogy to increasingly lower and lower budget productions.
I think in a lot of modern stuff they go too far. They now use CGI for things that could easily be practical effects, but they go with CGI because it's simply cheaper or because they want to A/B test different colors of wall paneling behind a character in post production, or who knows what. The end result is apparently good enough to make a billion dollars, but I can't stand it. Movies don't feel authentic anymore. It's hard to describe rationally, but the word 'soulless' sums up how movies made the new way make me feel. Even the scenes which are wholly practical/real get degraded; the excessive CGI and compositing used in the rest of the movie cast a miasma of unrealness across the entire movie.
https://64.media.tumblr.com/248d25e2185a58bf827d329490480fb9...
This video does a good job of explaining why, I think: https://www.youtube.com/watch?v=DY-zg8Oo8p4
For real though, I think the best CGI has always been when it's a light touch, providing only slight enhancements to an otherwise mostly-practical scene. And that's also when it's most invisible— so while superhero movies are obvious CGI-fests and can be clearly said to have been "ruined" by it, I think the most interesting modern CGI use is in lower-budget productions like TV shows, enabling the insertion of fantasy- or period-themed backdrops that would never be possible if you had to actually come up with all those props and extras in real life.
And I think to the extent that this is already happening, it's a lot harder to track because it's so much more subtle than the bombastic in-your-face effects of a Marvel movie showdown.
It would be interesting to see a revival of a show like Drive [0] as that was pretty ambitious and expensive for its time, but might be a lot more possible to do now.
I remember listening to the DVD commentary track for X2 (from 2003), and loving that scene near the beginning where Nightcrawler is hiding out in a cathedral but losing control of his teleportation powers because he's being remotely controlled/possessed by the villain of the story. Bryan Singer talks on the track about how great it was to fling the character (can't remember if it it was actually Alan Cumming or a stuntman) from the rafters, and how much better and more weighty it looked to have a "real" shot and just have to remove wires/harnessing vs filming it as a blank canvas and having to insert the character digitally.
Laypeople tend to think everything is linear, technologists tend to think everything is exponential; more often than not, reality tends to be sigmoidal. Technologies have exponential takeoffs followed by logarithmic plateaus. We're clearly well into the exponential phase of deep-learning ML, but it's only a matter of time before this approach hits its logarithmic phase.
Of course, the hard part of sigmoidal prediction is determining where in the curve we are. Does the current paradigm have an even steeper part ahead of it? Maybe. And yet, we could just as easily be right in the middle of the function, with a leveling off coming as the state of the art gives way to incremental improvements.
The much more likely scenario IMO is that people get used to the artifacts and notice them less.
A good path forward is to fuse these image-element compositing tools with some of the 3d scene inference ones. So you start out with 'giant fish riding in a golf cart, using its tail to steer', then give that as ground truth to a modeling tool that figures out a fish and a wheeled vehicle well enough to reference some canonical examples with detailed shape and structure, the idea of weight etc. Then you build a new model with those and do some physics simulation (while maintaining a localized constraint of unreality that allows the fish to somehow stay in the seat of the golf cart).
The groups supporting such absurd claims largely boil down to:
- Money-driven researchers in the applied AI field. Those are the people that spam popular ML conferences with barely novel contributions other than some minor tweaks on code from their previous papers.
- People unable to critically think and evaluate the significant limitations of SOTA methods occasionally marketed as AGI breakthroughs.
Last is virtually the entire HN userbase and the one that needs to be taken the least serious.The first category is much more troubling however, since they can significantly influence research directions due to the broken citation system in academia (more citations --> higher quality contributions).
I agree with your deeper criticism, though preferential attachment/ranking is very much How Humans Do Things. You could do a much improved citation system by expanding the time dimension and looking at papers that were unpopular at first and then attracted wide interest later.
Of course, academics also have a tendency to over-cite (because they don't want to be rejected for inadequate literature review), so there are incentives to cite a bunch of research whose premises or conclusions you hope to overturn.
Ok? It's a nascent technology. Look at the original DALL-E blog post from last year [1]. Now compare it to DALL-E 2 and Stable Diffusion.
I don't think AI will be able to create a movie anytime soon, but I think it will become "good enough" to serve as inspiration for creatives, or to replace simple stock footage (Much like SD and DALLE-2 is now).
It's not over for traditional moving-making. It would be decades before the software and hardware could surpass. But it will improve tremendously, just like computers do for nearly everything.
It's just the first version and seeing Stable Diffusion come out while openai's tools were coming out were something to remember and think about.
That Luddites have been successfully maligned as irrationally anti-technology crazy people is a propaganda victory by industrialist factory owners and their friends, the newspaper men.
IMHO the best argument for unfettered innovation is the impossibility of slowing innovation globally; it can be slowed in one country, but that country can't force all the rest to get with that program and will eventually be overtaken by technologically superior foes.
Reading the paper, it seems to be the "right" approach (separating temporal / spatial for both convolution and attention). Thus, I am optimistic what remains is to scale it up.
What you’re going to see is a race to the bottom, the same as with claymation films and 3d.
It will suddenly require a lot less (expensive, highly trained) people to make the same films.
You’ll still need expensive highly trained people to do it, but with different skills and a lot less of them can do a lot more a lot more quickly.
…and that means that some studios will make bad, low grade films… and some studios will make amazing films using hybrid techniques (like 3d printed faces for claymation).
…but overall, the people funding movies will expect to get more for less, and that will mean a downsizing of the number of people employed currently in certain roles.
Traditional film making over? Hm… it’s complicated. Is it over if the entire industry changes, but people are still making films? Or is that just the “traditional” part of it which is over?
It’s definitely going to change the industry.
People will still definitely film things.
…but, I wager, less people will be doing highly payed skilled manual work, which will replaced by a few people doing a different type of AI assisted work.
…and we’ll see some really amazing indy films, of small highly technical teams producing content with very little physical filming.
While these early versions are primitive, as a traditional filmmaker I think in a couple years these technologies will creatively empower visual storytellers in exciting new ways. The key will be developing interfaces which allow us to engage, direct and constrain the AI to help us achieve specific goals.
Exactly. Every step forward in creative technology is additional leverage for the artist with a vision.
I thought the same way you did about speech-to-text and image search, back in the day. boy was I wrong.
who is saying that?
You'll see an excellent mostly- or all-AI feature within 5-10 years. There will be terrible ones before that, maybe 2-3 years. The first really good one will be enjoyed on its own merits, ad the artificiality of it will come to light after it has gained popularity.
*produced, assisted, or discussed on set one time during lunchbreak”
Well, not quite obsolete, but you can see a significant qualitative difference between the two approaches.
I didn't know you could be drifting with a horse.
Something like: there are 10 peers, A sends to B, B waits for C, yadayada.
Imagine, gives a initial frame and a reference video on the side, that would be pretty dope.
Right now the pool of content you might be interested in is constrained to all the content that has been made. But there may be better content that does not exist yet that you would be even more interested in. The future is going to be very weird, but also very entertaining
Note the "AI by Meta" watermark in the example. So Zuck gets a cut as well.
Meta is falling fast imo
Like in 10 years you could plug this tech into high end VR and get a prompted reality dynamically generated that would be indistinguishable from our own.
Conventional physics giving birth to our universe is currently the model with the fewest assumptions.
What would have to take place to give rise to a universe the size of a data center, running an AI model of a human? It feels like we have to bake in assumptions of stable physics, a rise of a stable system for that data center, and some path towards creating it and modelling a human.
That said, if we believe we're capable of running billions of believable simulations, then we're more likely to be in such a simulation than ground reality. But a datacenter pocket dimension bakes in a lot of assumptions that make it less likely than our own universe.
Remember in this case the simulation is not running a model of physical reality - it’s only running a neural net that is fed into your senses.
In that case there are no ‘atoms’ it’s just a concept fed into your mind. Just like in dream it’s hard to question reality, you just go along with it.
It's equally important to understand and study the rules of this waking dream you are having while reading, not because there are actually atoms out there but because it behaves like there are. You can do physics without metaphysics, altough the two usually get entangled somewhat.
And if it is simulating our universe's atoms, then that part of it basically is our atoms. But is a neural net doing that really simpler than the atoms just running with a more direct mathematical model?
> Like in 10 years you could plug this tech into high end VR and get a prompted reality dynamically generated that would be indistinguishable from our own.
They said that 10 years ago about VR and it still is dog shit
The original Oculus was a step change, and current VR is a step change from that.
8k VR is coming and will be close to indistinguishable from reality visually at least.
At the end of the day you're still sitting/standing out there with two tv screens 2cm from your eye balls
If the screen is wide enough with eye tracking and divested rendering, how could you tell the difference?
The usual presumption of the simulation being a deliberate construction of a conscious being makes the whole thing seem like nu-religion for people who reject supernatural things. With the presumption of a being deliberately creating the situation, you pull in these notions: We're special, we exist on purpose, we are probably being examined and judged. This reeks of religion.
The neural net itself is built on a much simpler substrate in an external universe.
That's assuming there even really is one - senses can be hijacked, just like in dreams, so we may only think we have a brain - so strange.
Well, from the outside reality's perspective, it's helpful for people to spend the first few decades of their lives in an early 21st century simulation, just so they can gradually acclimate themselves to all this technology.
/folly
But even so, this era feels like it could be a singular phase shift. Maybe.
I know it's not the most loved Stephenson book, but bear with me (spoilers warning). The book features a couple of major plot points that feel increasingly prophetic:
- A global hoax, carefully coordinated, convinces a good chunk of the world that Moab has been destroyed by an atomic bomb. This is managed via a flurry of confusion, thoughtfully deployed pre-recorded video footage, social media posts, and paid actors who don't realize the scope of the hoax at the time. Naturally, a massive chunk of the population refuses to believe that Moab still exists even after the hoax is exposed.
- A group of hackers deploy and open source a massive suite of spambots that inundate social media and news sources with nonstop AI-generated misinformation to drown out real-world conversations about a topic. This is used specifically to drown out real conversation about a single individual as a proof of concept, but soon after has repercussions for the entire net...
- Thanks to exactly this kind of spambot, the "raw" unfiltered internet becomes totally unusable. Those with means end up paying for filtering services to keep unwanted misinformation out of their perspective of the internet. Those without means... either don't use the internet, or work in factories where human eyes filter out content that can't be filtered by AI.
I worry that exactly these kinds of developments are speeding us faster and faster down the road to a dystopian world where the "raw" internet is totally unusable. Right now, stuff like captcha and simple filters can keep out a lot of low-effort bot content on sites like Hacker News and niche forums (I think of home barista, bike forums, atlas, etc). Sites like Reddit are losing the war against bots and corporate propaganda; comment sections across the rest of the internet lost that war long ago, and just didn't realize it.
But those filters and moderators can only keep up with the onslaught of content for so long. What happens when GPT is used to spew millions of comments and posts at a forum from millions of ephemeral cloud IPs? And when DALL-E creates memes and photographic content? And when Make-a-Video enables those same spambots to inundate YouTube and Vimeo? It's clear that captchas are not long for this world either.
Will we see websites force more and more users to authenticate as a "real human" using their passport and government-issued ID? Maybe a Turing or Voight-Kampff test? And what does it mean when there's no longer a way to participate on the internet anonymously? As far as I can tell, limiting a site to only real human users doesn't guarantee quality -- all you have to do is look at Facebook to understand that. And somehow, despite being an incredibly easy target for bots and spam, niches of 4chan retain (racist, insane, and confusing) traces of genuine thought and conversation.
The internet has been such a valuable tool in my life, and I still love browsing blogs, forums, etc. to learn about unique people doing unique things. What happens when human-generated content is hard to come by and near impossible to distinguish from promotional AI garbage? I fear for my ability to discover new content in that world.