Film: Frame Interpolation for Large Motion
film-net.github.io
film-net.github.io
The work described in the linked article is also extremely impressive and feels almost unreal, in any case.
I don't know; I feel the real-world applications are still missing and what we are now seeing are tech demos (impressive ones!) and gimmicks. I'm still waiting to see all this ML stuff to be used in a productive context.
Also Retrobatch has a ML based image classifier, but I didn't try that yet.
Other than these, yet all these are impressive tech demos. We need them in production and preferably open source (model + training + process) versions.
I think Dwarf Fortress has the story generation part. The aesthetics/graphics part not yet..
And I think its procedurally generated, but with complex and strange results.
https://www.reddit.com/r/dwarffortress/comments/2ztnkw/i_thi...
I suspect Ultima Ratio Regum is a step up from DF even in terms of procgen world lore. I mean the creator (author?) of that game wrote their PhD thesis on the subject
News and sports, in particular.
It got some weird pushback, but Peter Jackson's film "They Shall Not Grow Old" really helped make it's subjects so much more real by cutting through the limitations of the old footage from the 1st World War. Being able to apply similar techniques more cheaply and quickly will bring a lot of old footage to life and make the past much more real for the viewer.
As someone in the field of computer graphics , where there’s been considerable ML research over the past few years that are more reliably applicable to people’s lives , most of the exciting stuff doesn’t make it to the front page of HN even if it’s posted here.
There’s been lots of research in the past few years. The initial shiny stuff makes it on here, but it’s the follow up iterations that are highly catalyzing of change that don’t because public interest in those topics has waned in the interim.
I should add that my point is more that there’s not been a slow down in ML research, and the ebb and flow of interest on HN isn’t indicative of accelerated progress in the field. Simply of what things have captured the public mind.
Anyway on to the links…
Disney had a slew of papers out this year, the facial motion retargeting ones in particular are very interesting for use in production of films and “metaverse” characters
https://studios.disneyresearch.com/machine-learning/
Luma have had a lot of progress in their Nerf capture and on device representation which will likely have huge effects for e-commerce use cases among other things
https://captures.lumalabs.ai/unbounded
Nvidia released an ML based version of OpenVDB that will potentially improve effects in films, but could be huge for games
There were also a ton of neural rendering papers at siggraph that I still need to separate in my head mentally since I saw them presented back to back, so I apologize for just sharing a dump
https://twitter.com/neural_fields/status/1555947856271446018...
Apple released some neural rendering content too that has a lot of implications for spatial product training
I wouldn't call moving pixels on a screen "real-world". Are these technologies going one day to have a physical effect on our lives, like, in the real real-world? I very much doubt it.
It's week two, give it two decades.
Or come and build some!
Amazon in 2000 didnt knew it would become infra for the world with AWS
Probably because of falling into the uncanny valley [0].
Don't get me wrong, it's an incredible feat, and seems to handily beat the other automagic interpolators (eg, 3:49 in the video at the bottom of TFA) in terms of minimizing "pop-in", but it's still clearly present in dentition.
I was going to download this thing and generate a bunch of samples to send to my family tomorrow, possibly dumping them right into the uncanny valley and being too unobservant to notice I was doing it
https://film-net.github.io/static/images/000204/interpolated...
Would be interesting to see if this can be made more context sensitive - i.e. algorithm recognizes this as a person's head and fills in details more intelligently.
[0] https://replicate.com/google-research/frame-interpolation/ex...
Animated frames are supposed to convey intention. They’re fantastic at doing this since you can manipulate every detail of every frame. The idea that you’ll just run an AI through it, that might work for dialogue scenes of a typical Japanese TV anime where intention is low and mostly it’s indeed grunt work. But I would imagine it would be a bit lifeless - unless someone trains an ML specifically for anime using good animation as a reference.
Basically just moving between two frames is an example of extremely poor animation.
Source: am animator, sort of.
In an animation those two photos would be drawn and created as keyframes which would then get interpolated many ways (hopefully not linear and as robotic and weird as this).
Very interesting technology though. I could see this coming to an smartphone near you any day now. And there will be ways people animate with these tools but this isn't it.
I would have been pretty bummed by my teens if I found out all my life's history was there for the whole world to crawl, collect, train their ad/surveillance NNs on, etc.
Sure you could potentially identify the kid, but nobody would ever have any reason to go through the effort.
Privacy is not a right, it is a condition under which you have different rights. Whether that condition exists depends on social norms - for example a picture of someone in their underwear at a public locker room is very different from a picture of someone in an equal state of undress at the beach. A major factor in whether something is an invasion of privacy is the amount of effort others need to take for it to become public - you can for example have a private conversation in a public restaurant despite the fact someone could theoretically eavesdrop, it only stops being private when you start talking so loudly that there is no need to eavesdrop. Also to be considered is the likelihood of someone maliciously trying to gain information - a bank failing to shred financial documents might be a violation of privacy as someone going through their trash is a real risk; but my grandma doesn't need to shred old post cards. I would definitely consider trawling obscure websites with an ai to be in the eavesdropping/dumpster diving regime.
Coming back to the original point, yes you don't have to explain to anyone else why you are exercising your rights, but freedom isn't free and you need to be able to justify to yourself that what you gave up in exchange for your rights was worth it. Idealistic platitudes might at first glance seem comforting, but they make serious conversation impossible. At the end of the day privacy on the internet is an extremely nebulous concept, and without questioning "what's the point?" every now and then, it's easy to lose perspective.
PS: Privacy is a right, not a mere condition. In country of my residence it's meant to be protected by the constitution, so I think it qualifies.
dain didn't work for me in m1 https://github.com/nihui/dain-ncnn-vulkan
That's a neat use case, and definitely a good way to show off, but what about more than one image?
The overwhelming majority of video that exists today is 30fps or lower. The overwhelming majority of displays support 60hz or more.
Most high-end TVs do some realtime frame interpolation, but there is only so much an algorithm can do to fill in the blanks. It doesn't take long to see artifacts.
I would be more interested to see what an ML-based approach could do with the edge cases of interpolating 30fps video than 2 frames.
> Theoretically, you can do a better job with multiple frames but this doesn't bring much more values beside of some extreme cases.
Edge cases that require more information than is present in two frames are very common. That's why most frame interpolation methods also have an "artifact masking" feature.
But what if we did use the information from surrounding frames? That would probably be too complicated for traditional frame interpolation, but that's not what we're talking about.
What if we used a data set trained on the entire video file - or even a collection of similar video files - to fill in the gaps?