Runway – Create impossible video
runwayml.com
runwayml.com
If these demos fairly represent the user experience then it's slick as hell and further blurs the line between editor, compositor, DI specialist etc. Much will depend on whether the ML marketplace can be competitive with many mature commercial offerings from software studios that have no intention of letting their lunch be eaten. Video production involves a lot of bleeding edge technology but the client base is also suspicious of new providers or do-it-all solutions at first, and very loyal to products and tech support offerings that have got them through difficult projects in the past, so there will be a big hill to climb between people acknowledging it's cool as hell and their willingness to sell it to a producer who is making a 5, 6, or 7 figure bet on what some editor/VFX geek is telling them.
The web browser/cloud storage is an issue. It could be a plus in many circumstances, but it's also a big barrier to any production that's working in a remote location without reliable internet, especially given the massive data volumes involved. That limits a lot of the use cases to post production, and many producers and directors are going to be wary of starting on one platform and finishing on something else; nobody likes workflow changes unless they can be shown something seamless, like having your Avid/Final Cut/Premiere/* project leave your machine and show up frame-perfect in Runway, along with a definite answer about export time like being able to take in 12 hours of video and have it online within 24h. There will also have to be a lot of questions about security, downtime, and being able to get your project back out of Runway if money runs out, the editor gets fired, creative differences rear their head etc.
Looks like an instant win for short-form projects like personal shorts, music videos, commercials, corporate, demo reels, spec pieces. Potentially very good for reality TV indie and low-mid budget films once the above questions are answered. Toughest nut to crack is large budget or episodic TV where there's very stiff competition and contractual or professional commitments already in place, but doable within 5 years.
The main drawback I observed was computation speed, most edits required >20 seconds of loading/buffering (with a 1Gbps internet connection). I presume the computation occurs on their own servers, so with beefier hardware the performance could be increased, however my intuition is that the app would perform faster if running natively (rather than a shared resource quota on a remote server).
In this regard, the lack of performance could be a major productivity killer, my hypothesis is that it is the video processing/manipulation which is taking a long time (and not the model classification). Many companies have tackled this problem space with completely video transcoding chips (such as YouTube), however this generally still incurs long periods of waiting.
My best guess is that the web can be a path of least resistance and that’s why this launched on the web, but as soon as they have available resources, they could package and ship a native app (hopefully one that can take advantage of modern chip ML tech)
I'm currently working on a fully-automatic version of this based off some excellent research from the University of Washington. It's still in its infancy, but if you'd like to follow my progress, I post occasional updates over on https://nomoregreenscreen.com
This is just the beginning of what's possible in the intersection between DL and creative filmmaking, and I'm really excited to get to be in this field at a time when compute is cheaper than ever and all the information I could want is available for free on the internet!
AI Masking: check
AI Depth estimation: check
AI Flow estimation: check
Flicker artifacts just like the public AI models: check
EDIT: Also, this isn't actually new anymore? A quick check found two very similar startups: unscreen.com vfx.comixify.ai And AI rotoscoping has been part of DaVinci Resolve 17 since Februrary: https://www.blackmagicdesign.com/products/davinciresolve/wha...
Us nerds often forget that possibility and usability are not nearly the same.
Usage happens when non-technical people like editors are able to get their hands onto the technology.
The ML people don't make it easy to work with the output their models produce. I was playing with TensorFlow/MediaPipe last month, to see if I could get them to play nicely with my canvas library. The results were quite promising[1][2][3]. Still, I think making it easier for devs to use these ML models in various ways needs to be prioritized.
CodePen links (all request access to the device camera):
[1] - TensorFlow body-pix model - hide the background in various ways - https://codepen.io/kaliedarik/pen/ZEeoZaP
[2] - MediaPipe Selfie Segmentation model - hide the background in various ways - https://codepen.io/kaliedarik/pen/PopBxBM
[3] - MediaPipe Facemesh - draw on the face in real time - https://codepen.io/kaliedarik/pen/VwpGrVG
People who do a lot of video editing usually already have decent PCs on which to do it.
There’s also the data play of making the user keep data there and pay rent forever to keep it accessible. You can download your videos but you lose edit history etc.
Likely there's something I'm missing, here.
Those videos or GIFs or whatever they are work fine in Chrome and Safari, so it turned out that I was missing something. :)
Perhaps you have autoplay blocked? They are set to autoplay by the looks of things.
Relatedly, I've never understood why some books have 4 pages of testimonials in the front. Once you get past 5 of these are you really more likely to buy the book?
I can see this could be appealing for people who don’t want to learn or install professional software, I think there is value in that. I’d like to see a client side version of this app, many people have beefy gaming GPUs that could run the models client side
You're not going to beat Fusion, Nuke, or even consumer tools like Hitfilm at the VFX game. A better use of this tech would be to take the area you improved in and turn it into a plugin for all or one of these.
Guess the actual target demographic gets it?
This is different though. No cutting or masking, the actual A.I. itself is doing it automatically. No filming the same takes with the same camera position 3 times in a row then doing tedious editing that takes hours.
I have my doubts as to how well it would work... But assuming it did then it's a real game-changer in the editing world because it could not only save a ton of time, but actually produce new abilities, like cloning with object crossover, and while the camera is moving around erratically.
Also, the demo tells me the target audience is not pro video folks (film/TV) rather individuals. They might find it harder to keep spending so much on shuttling data. The power of ML in video has been clear to many but we need these tools to work offline.
Are there great ML specific chips for PC/laptops? I mean like the ones that Apple keeps talking about? I am not in the domain so I don't know, but I guess GPUs are the best chips for this, is there any reason this software would not work on a beefy RTX 30X0 based device?
Update
1. I mention "even in India" because I keep seeing bandwidth pricing here is still relatively cheap, globally speaking. OK, not cheap by typical Indian household measures but if you are doing pro video then yeah, cheap infrastructure.
I haven’t actually read the article, so perhaps they say as much. But if so, that’s surprising.
EDIT: thinking a bit more carefully, the limitation is upload speed, not download speed. In that area, the US has been lagging behind in tech.
Still, I wouldn’t bet against it being viable. I remember the first few months YouTube launched. The video quality was atrocious, the worst on the internet by far — this isn’t revisionism; other platforms tried to compete on quality.
Didn’t matter; youtube won. And now we enjoy 4K steaming.
It can do progressive sharpening too, so that e.g. if you pause the video you can see the full quality.
I agree that it's not ideal, but a lot of people would use it if the price was right. I would.
As far as the video assets being on the server, sure, it may suck to upload those. But people do (sometimes) upload HQ assets to youtube, so it's not unheard of.
At $35/month Davinci works out cheaper in less than 10 months...
Machine Learning is also genuinely making new things possible. Two Minute papers has a video here: https://www.youtube.com/watch?v=22Sojtv4gbg demonstrating a model that trains on driving footage, then adjusts video game footage to look almost indistinguishable from a real world shot (and in real time!). This is technically a video game application, but it's not hard to imagine how this could be used in a video or movie context.
The extent to which this site/software actually captures these broader trends is probably minimal, but they're definitely out there. I think we're on the edge of another huge shift, like the one where CG became more practical and down-to-earth so it could be used anywhere. A lot of the same stuff will be done, but by a couple of people rather than an army of editors.
Edit: I just spent the last 25 minutes persisting with it, I'd say their auto tools are actually worse than AE. I guess it's a consumer product, but even then... seems like you'd be spending a lot of time for little reward, for example, it's not clear to me how I might refine this mask: https://share.getcloudapp.com/rRujBABR
Edit 2: I'm gonna be a little more generous and say I think the mask itself is pretty good, it can deal with some pretty complicated situations well, but the tooling is very remedial.
Not sure about the cloud based aspect. I am teaching video online right now. Bandwidth issues is making many students drop the course.
Any modern mirrorless camera is doing the same analysis on the scene with object detection and tracking. So, they're just doing it on a video stream without phase/depth data.
Sony, Nikon, Canon, Panasonic and Fuji have similar technologies built in their cameras, but having this on your desktop for using with your videos is a nice step forward.