Few-Shot Video-to-Video Synthesis
nvlabs.github.io
nvlabs.github.io
https://i.stack.imgur.com/tqPDO.gif
Actually that's probably better than average. Enough to bring back memories but pretty awful and hard to look at more than a few moments especially from the flicker.
Seems like a perfect job for computational video processing
Generally video models are quite far behind image models because they're so much higher dimensional (much more computationally expensive and expensive to annotate) but I'm sure it's a matter of time before someone releases a DeOldify type model for video (if one doesn't exist already).
[1] https://github.com/jantic/DeOldify [2] https://github.com/idealo/image-super-resolution
If all you have are the VHS tapes you could try getting them converted by a different company, or doing it yourself. The flickering looks a lot like a VHS player not doing brightness correction correctly to me. The original film wouldn't have flicker, The film scanner would be unlikely to introduce it. A VHS player needs to do brightness correction on the signal it reads from the tape in any case since the signal strength may vary with tape age or wear, or even the signal strength at time of recording. A VHS player would normally use the sync pulse to detect how much it needs to amplify the signal, if that mechanism gets out of whack you can get wild brightness fluctuations. Some of the color could probably also be rescued if the VHS conversion was done more carefully.
If you don't have the VHS tapes anymore either (which is a perfectly reasonable thing after having them converted) you could still probably get a lot of the brightness fluctuations out by editing the avi files. I don't know of the top of my head any tool that'll do brightness equalisation like that automatically but I'm sure it exists, it's computationally not a hard problem. You'll probably not be able to recover much color from this stage though, the signal is thoroughly quantized once it hits the digital realm and no amount of fiddling will get you any more color resolution though you might be able to correct it a small amount with some color correction.
In all cases you're best off getting as close to the source material as possible, but I'm sure that even if the digital copies are all you have you can get it to a place where it's at least watchable.
Though this is very impressive, it seems to take longer and longer to make those tiny improvements that make all the difference wrt believability.
N.B. The most convincing TTS I've ever heard (predating Lyre by quite a bit) generated things like this: http://web.archive.org/web/20190803012012/https://instaud.io...
https://drive.google.com/file/d/1zRvJEGJjTpKvvzel-J0agh3fKBn...
I'm integrating it into a "Snapchat filter" type app with lightweight social features just as a means to bring it to market and hopefully attract Facebook or Snapchat or Tencent into buying it. I'm building it to sell, essentially.
I need capital so I can fund my real ambitious start-up of end to end computational filmmaking. Graph-based story language, light field camera optics, tracking and localization in prerendered environments, content-aware shaders, real time storyboard population and automated editing, posture estimation and mistake correction...
With patent protection, I think it could unseat Disney and make more money than they do with Marvel and Star Wars.
I need a lot of capital to build my lab. Optics (good sensors and glass), a modest studio with rigging and tracking set up for experiments, and a handful of engineers.
https://apps.apple.com/us/app/celebrity-voice-changer-face/i...
I checked the Android app store and all the apps used text to speech before vocoding. I had no idea this existed on iPhone.
Thanks.
However, one of the problems of deep learning is that you need to have a good dataset first, but how are you going to build one? Well, the big game companies wouldn't reveal their models and textures that easily. For indie artists to utilize this technology, there needs to be a centralized community project built around gathering and preprocessing data, rather than waiting for someone like Adobe or Autodesk do it.
As the tech gets cheaper, this calculous changes.
Kudos to NVIDIA Corp. for the planned release of this code. It seems to have a lot of commercial potential. Maybe they see it increasing demand for their hardware.
I'm automatically suspicious when someone doesn't, even though I guess most of the time it's nothing nefarious. It's just extra work and effort to bring the code into presentable shape and that effort could be spent on the next paper. Once the paper is published, the material benefits have been reaped, publication count incremented. Of course this is a short-sighted view, because in the long term not only the paper counts matter, but also one's general reputation within the community.
Maybe if you scale up the training data the model is going to learn some better warping, but I think to really get better photorealism there has to be a 3D component, some kind of distance/movement estimation like in https://arxiv.org/pdf/1704.07804.pdf and a shape inference/transfer step.
Can’t wait to see FOX News anchors deepfaked to have the exact opposite views.