VideoGigaGAN: Towards detail-rich video super-resolution
videogigagan.github.io
videogigagan.github.io
I'd say most videos in practice are longer than 200 frames, so lot more research is still needed.
For live streaming it's pretty common to see 2 or 3 seconds (reduces broadcast delay, but with some caveats).
0: https://dev.to/100mslive/introduction-to-low-latency-streami...
It could work for TV or movies if done properly at the scene transition time though.
Moreover I imagine that further research and power will do a lot, smarter, and quicker.
Don't forget people had toy story-comparable games in a decade or so after it was originally rendered at 1536x922.
But idk how someone can write "extremely long videos" with a straight face when meaning seconds.
Maybe "long frame sequences"
Would it kill them to say that the method works best on short videos/scenes?
Source: https://www.nasa.gov/history/115-years-ago-wright-brothers-m....
:) "Our model is proven production-ready using real-world footage from Taken 3"
For animations it's around 15 seconds.
Funny - this is also a good description of I Am Legend.
https://www.cbr.com/i-robot-original-screenplay-isaac-asimov...
I think like the motorcycle chase that they borrowed from the lost world in Jurassic world, they also have a scene with those tiny dinosaurs pecking someone to death.
Another movie of him that crimes of this non stop is The Prestige.
https://en.wikipedia.org/wiki/Children_of_Men#Single-shot_se...
Also, it's less likely that you'd want to upscale a modern movie, which is more likely to be higher resolution already, as opposed to an older movie which was recorded on older media or encoded in a lower-resolution format.
It reminds me of the story about the Air Force making cockpits to fit the elusive average pilot, which in reality fit none of their pilots...
The details changing every ten seconds or so is actually a good thing; the viewer is reminded that what they are seeing is not real, yet still enjoying a high resolution video full of high frequency content that their eyes crave.
I don't see the word "bias" appear anywhere in this work :'(
I guess it will be like people pointing out bird sounds in movies, that those birds don't exist in that country.
The real owl has fine light/dark concentric circles on its face. The app turned it into gray because it does not see any sign of the circles. The real owl has streaks of spots. The app turned them into solid streaks because it saw no sign of spots. There's more where this came from, but basically only looks good to someone who has no idea what the owl should look like.
More seriously, though, yes, the thing you’re describing is exactly what the AI safety field is attempting to address.
Is it though? I think it's pretty obvious to any neutral observer that this is not the case, at least judging based on recent examples (leading with the Gemini debacle).
Who decides what is "societally-harmful content"? Isn't literally rewriting history "societally-harmful"? The black T.J. was a fun meme, but that's not what the alignment's "unintended effects" were limited to. I'd also say that if your LLM condemns right-wing mass murderers, but "it's complicated" with the left-wing mass murderers (I'm not going to list a dozen of other examples here, these things are documented and easy to find online if you care), there's something wrong with your LLM. Genocide is genocide.
> Who decides what is "societally-harmful theft"? > Who decides what is "societally-harmful medical malpractice"? > Who decides what is "societally-harmful libel"?
The people who care to make the world a better place and push back against those that cause harm. Generally a mix of de facto industry standard practices set by societal values and pressures, and de jure laws established through democratic voting, legislature enactment, and court decisions.
"What is "societally-harmful driving behavior"" was once a broad and undetermined question but nevertheless it received an extensive and highly defined answer.
This is circular. It's fine to just say "I don't know" or "I don't have a good answer", but pretending otherwise is deceptive.
Are you stupid, or just pretending to be?
[1] - Are people seriously still trying to argue that it was some sort of weird artifact? It was blatantly overt and explicit, and absolutely embarrassing. Hopefully Google has removed everyone involved with that from having any influence on anything for perpetuity as they demonstrate profoundly poor judgment and a broken sense of what good is.
This is a bad idea.
Equating extremist views with those seeking to defend human rights blurs the ethical reality of the situation. Adopting a centrist position without critical thought obscures the truth since not all viewpoints are equally valid or deserve equal consideration.
We must critically evaluate the merits of each position (anti-fascists and fascists are very different positions indeed) rather than blindly placing them on equal footing, especially as history has shown the consequences of false equivalence perpetuate injustice.
I know nothing of makeup tho, just describing my observations.
The nails have nail polish in the original, and the lips also look like they have at least lip gloss or a somewhat more muted lipstick.
Socrates: I heard, then, that at Naucratis, in Egypt, was one of the ancient gods of that country, the one whose sacred bird is called the ibis, and the name of the god himself was Theuth. He it was who invented numbers and arithmetic and geometry and astronomy, also draughts and dice, and, most important of all, letters.
Now the king of all Egypt at that time was the god Thamus, who lived in the great city of the upper region, which the Greeks call the Egyptian Thebes, and they call the god himself Ammon. To him came Theuth to show his inventions, saying that they ought to be imparted to the other Egyptians. But Thamus asked what use there was in each, and as Theuth enumerated their uses, expressed praise or blame, according as he approved or disapproved.
"The story goes that Thamus said many things to Theuth in praise or blame of the various arts, which it would take too long to repeat; but when they came to the letters, "This invention, O king," said Theuth, "will make the Egyptians wiser and will improve their memories; for it is an elixir of memory and wisdom that I have discovered." But Thamus replied, "Most ingenious Theuth, one man has the ability to beget arts, but the ability to judge of their usefulness or harmfulness to their users belongs to another; and now you, who are the father of letters, have been led by your affection to ascribe to them a power the opposite of that which they really possess.
"For this invention will produce forgetfulness in the minds of those who learn to use it, because they will not practice their memory. Their trust in writing, produced by external characters which are no part of themselves, will discourage the use of their own memory within them. You have invented an elixir not of memory, but of reminding; and you offer your pupils the appearance of wisdom, not true wisdom, for they will read many things without instruction and will therefore seem to know many things, when they are for the most part ignorant and hard to get along with, since they are not wise, but only appear wise."
What I am concerned about is that AI providers will keep wasting time and resources trying to implement band-aid "patches" to address what is actually an innate limitation. For example, exception processing at the output stage fails in ways we've already seen, such as AI photos containing female popes or an AI lying to deny that HP Lovecraft had a childhood pet (due to said pet having a name that was crudely rude 100 years ago but racist today). The alternative of limiting the training data to include only curated content fails by yielding a much less useful AI.
My, probably unpopular, opinion is that when AI inevitably screws up some edge case, we get more comfortable saying, basically, "Hey, sometimes stupid AI is gonna be stupid." The honest approach is to tell users upfront: when quality or correctness or fitness for any given purpose is important, you need to check every AI output because sometimes it's gonna fail. Just like auto-pilots, auto-correct and auto- everything else. As impressive as AI can sometimes be, personally, I think it's still lingering just below the threshold of "broadly useful" and, lately, the rate of fundamental improvement is slowing. We can't really afford to be squandering limited development resources or otherwise nerfing AI's capabilities to pursue ultimately unattainable standards. That's a losing game because there's a growing cottage industry of concern trolls figuring out how to get an AI to generate "problematic" output to garner those sweet "tsk tsk" clicks. As long as we keep reflexively reacting, those goalposts will never stop moving. Instead, we need to get off that treadmill and lower user expectations based on the reality of the current technology and data sets.
GPT4 told me with no hesitation.
I'll take that as supporting my point about the folly of wasting engineering time chasing moving goalposts. :-)
"Hmm… let’s try a different topic. Sorry about that. What else is on your mind?"
We seem to have a culture of completely paranoid people now.
When the internet came along every conversation was not dominated by "but what about people knowing how to build bombs???" the way most AI conversation flips to these paranoid AI doomer scenarios.
They will have huge budgets for compute and the makers of compute will be happy to absorb those budgets.
Cloud production was already growing but this will continue to accelerate it imho
the compute budgets for basic run of the mill small screen 3D rendering and 2D compositing is already massive compared to most other businesses of a similar scale. the industry has been under paying their artists for decades too.
I'm willing to bet that as soon as unreal or adobe or whoever comes out with a stable diffusion like model that can be consistent across a feature length movie, they'll stop bothering with artists altogether.
why have an entire team of actual people in the loop when the director can just tell the model what they want to see? why shy away from revisions when the model can update colour grade or edit a character model throughout the entire film without needing to re-render?
Limited access to the tech added some mystique to it too.
Just like digital cameras created a lot more average photographers, it pushed photography to a higher standard than just having access to expensive equipment.
While of course it isn't impossible for any industry to reinvent itself, movie as an art form won't die....having doubts about where it's going.
Unlikely in the next 10 years or the next 100?
With that said, I’m pretty confident that the movie industry will exist in 10 years (maybe heavily tranformed, but still existing and still pretty big). If it’s still a big part of current popculture by then (vs obviously on its way out) then I’d expect a collapse of it to require a change that is not a result of AI proliferation, but something else entirely.
Realistically, AI being able to replace Hollywood is something that could happen in 20-50 years. That's within most people's lifetime.
I don't think we're far away from models that are able to take video input of an almost finished movie and add the finishing touches.
Eg. make the lighting better, make the cgi blend in better, hide bits of set that ought to have been out of shot, etc.
I think the video of the camera operator on the ladder shows the artifacts the best. The main camera equipment is no longer grounded in reality, with the fiddly bits disconnected from the whole and moving around. The smaller camera is barely recognizable. The plant in the background looks blurry and weird, the mountains have extra detail. Finally, the lens flare shifts!
Check out the spider too, the way the details on the leg shift is distinctly artificial.
I think the 4x/8x expansion (16x/64x the pixels!) is pushing the tech too far. I bet it would look great at <2x.
I believe this applies to every upscale model released in the past 8 years, yet undeterred by this scientists keep pushing on, sometimes even claiming 16x upscaling. Though this might be the first one that is pretty close to holding up at 4x in my opinion, which is not something I've seen often.
Plus, a rather small data set: REDS and Vimeo-90k aren't massive in comparison to what people speculate Sora was trained on.
I remember being able to add a lot of detail to the monsters that I could barely make out amidst the clothes piled up on my bedroom floor.
You'd have to train it to go from a reduced resolution to the original resolution, then apply that to small parts of the screen at the original resolution to get an enhanced resolution, then stitch the parts together.
The difference will be that this time the images will be crystal clear, just hallucinated by a neural network.
Their 128x128 to 1024x1024 upscales are very impressive, but I find the real artifacts and weirdness are created when AI tries to upscale an already relatively high definition image.
I find it goes haywire, adding ghosting, swirling, banded shadowing, etc as it whirlwinds into hallucinations from too much source data since the model is often trained to work with really small/compressed video into an "almost HD" video.
Intergrading this with Adobe's object tracking software (in premier/after effects) may help.
I think it's almost inherently so, because of the care that an artist takes in choosing keyframes, deforming the action, etc.
Since this isn't lossless decompression, the point of having no "real" data is already reached. It _is_ inventing things, and the only relevant question is how plausible are the things being invented; in other words, if the video also existed in higher resolution, how close would it actually look like the inferred version. Seems obvious that this metric increases as a function of the amount of information from the source, but I would guess the exact relationship is a very open question.
The start point. Upscaling is by definition creating information where there wasn't any to begin with.
Nearest neighbor filtering is technically inventing information, it's just the dumbest possible approach. Bilinear filtering is slightly smarter. This approach tries to be smarter still by applying generative AI.
There is plenty of real information: that's what the model is trained on. That information ceases to be real the moment it is used by a model to fill in the gaps of other real information. The result of this model is a facade, not real data.
(And I already have glasses, thank you).
If you just want the picture to be pretty, this is probably cheaper than a bigger sensor.
[1] https://www.computer.org/csdl/proceedings-article/dcc/1992/0...
They used digital upsampling techniques and colorization to make World War One footage into high resolution. Jackson would later do the same process for the 2021 series Get Back, upscaling 16mm footage of the Beatles taken in 1969: https://www.imdb.com/title/tt9735318/
Both of these are really impressive. They look like they were shot on high resolution film recently, instead of fifty or a hundred years ago. It appears that what Peter Jackson and his team did meticulously at great effort can now be automated.
Everyone should understand the limitations of this process. It can't magically extract details from images that aren't there. It is guessing and inventing details that don't really exist. As long as everyone understands this, it shouldn't be a problem. Like, we don't care that the cross-stitch on someone's shirt in the background doesn't match reality so long as it's not an important detail. But if you try to go Blade Runner/CSI and extract faces from reflections of background objects, you're asking for trouble.
I always thought that alias-free convolutions can produce much more natural videos
The core problem is that any single AI will learn how to upscale "things in general", but won't be able to take advantage of inputs from the source video itself. E.g.: a close-up of a face in one scene can't be used elsewhere to upscale a distant shot of the same actor.
Transformers solve this problem, but with quadratic scaling, which won't work any time soon for a feature-length movie. Hence the 10 second clips in most such models.
Transformers provide "short term" memory, and the base model training provides "long term" memory. What's needed is medium-term memory. (This is also desirable for Chat AIs, or any long-context scenario.)
LoRA is more-or-less that: Given input-output training pairs it efficiently specialises the base model for a specific scenario. This would be great for upscaling a specific video, and would definitely work well in scenarios where ground-truth information is available. For example, computer games can be rendered at 8K resolution "offline" for training, and then can upscale 2K to 4K or 8K in real time. NVIDIA uses this for DLSS in their GPUs. Similarly, TV shows that improved in quality over time as the production company got better cameras could use this.
This LoRA fine-tuning technique obviously won't work for any single movie where there isn't high-resolution ground truth available. That's the whole point of upscaling: improving the quality where the high quality version doesn't exist!
My thought was that instead of training the LoRA fine-tuning layers directly, we could train a second order NN that outputs the LoRA weights! This is called a HyperNet, which is the term for neural networks that output neural networks. Simply put: many differentiable functions are twice (or more) differentiable, so we can minimise a minimisation function... training the trainer, in other words.
The core concept is to train a large base model on general 2K->4K videos, and then train a "specialisation" model that takes a 2K movie and outputs a LoRA for the base model. This acts as the "medium term" memory for the base model, tuning it for that specific video. The base model weights are the "long term" memory, and the activations are its "short term" memory.
I suspect (but don't have access to hardware to prove) that approaches like this will be the future for many similar AI tasks. E.g.: specialising a robot base model to a specific factory floor or warehouse. Or specialising a car driving AI to local roads. Etc...
https://paperswithcode.com/task/multi-frame-super-resolution
(It doesn't hurt that a few minor hallucinations aren't going to bother anyone.)
Yes, it is impressive, but it's not what you want to actually "enhance" a movie.
This appears to happen because they begin animating as soon as they finish loading, which happens at different times for each side of the image.
The OP has more in common with the defunct DLSS 1.0, which tried to infer extra detail out of thin air rather than from previous frames, without much success in practice. That was like 5 years ago though so maybe the idea is worth revisiting at some point.
If you have the high res data you can actually compress the details which are there and then recreate them. No need to have those be recreated, when you actually have them.
Downscaling the images and then upscaling them is pure insanity when the high res images are available.
I think it a bit unimaginative to see no use cases for this.
The issue isn't NN reconstruction, but that you are reconstructing the wrong data.
I think upscaling framerate would be more useful.
Down sampling is a bad way to do compression. It makes no sense to do NN reconstruction on that if you could have compressed that image better and reconstructed from that data.
https://en.wikipedia.org/wiki/JPEG#JPEG_codec_example
So you could downscale, then compress as usual, and then upscale on playback.
It would obviously be quite attractive to be able to ship compressed 480p (or 720p etc) footage and be able to blow it up to 4K at high quality. Of course you will have higher quality if you just compress the 4K, but the file size will be an order of magnitude larger.
Are you saying low-pass filtering is bad for compression?
Is blurring good for compression? I don't know what that means. If the image size (not the file size) is held constant, a blurry image and a clear image take up exactly the same amount space in memory.
Blurring is bad for quality. Our vision is sensitive to high-frequency stuff, and low-pass filtering is by definition the indiscriminate removal of high-frequency information. Most compression schemes are smarter about the information they filter.
Consider lossless RLE compression schemes. In this case, would data with low or high variance compress better?
Now consider RLE against sets of DCT coefficients. See where this is going?
In general, having lower variance in your data results in better compression.
> Our vision is sensitive to high-frequency stuff
Which is exactly why we pick up HF noise so well! Post-processing houses are very often presented with the challenge of choosing just the right filter chain to maximize fidelity under size constraint(s).
> low-pass filtering is by definition the indiscriminate removal of high-frequency information
It's trivial to perform edge detection and build a mask to retain the most visually-meaningful high frequency data.
You're also ignoring the part where all lossy codecs throw away those same details and then fake-recreate them with enough fidelity that people are satisfied. Same concept, different mechanism.
Look up what 4:2:0 means vs 4:4:4 in a video codec and tell me you still think it's "pure insanity" to rescale.
Or, you know, maybe some people have reasons for doing things that aren't the same as the narrow scope of use-cases you considered, and this would work perfectly well for them.
Because you can just not downscale them and compress them in the frequency domain and encode them in 200Kbps? This is pretty obvious, seriously do you not understand what JPEG does? And why it doesn't do down sampling?
Do you seriously believe downscaling outperforms compressing in the frequency domain?
Throwing away 3/4 (half res) or 15/16 (quarter res) of the data, encoding to X bitrate and then decoding+upscaling looks far better than encoding to the same X bitrate with full resolution.
For high bitrate, native resolution will of course look better. For low bitrate, the way H.26? algorithms work end up turning high resolution into a blocky ringing mess to compensate, vs lower resolution where you can see the content, just fuzzily.
Go get Tears of Steel raw 4K video (Y4M I think it's called). Scale it down 4x and encode it with ffmpeg HEVC veryslow at CRF 30. Figure out the bitrate, then cheat - use two-pass veryslow HEVC encoding to get the best possible quality native resolution at the same bitrate as your 4x downscaled version. You're aiming for two files that are about the same size. Somehow I couldn't convince the codec to go low enough to match, so I had the low-res version about 60% of the high-res version filesize. Now go and play them both back at 4K with just whatever your native upscale is - bilinear, bicubic, maybe NVIDIA Shield with it's AI Upscaling.
Go do that, then tell me you honestly think the blocky, streaky, illegible 4K native looks better than the "soft" quarter-res version.
This has two major benefits:
1. You cut out the low resolution half your network entirely. (Go check out the architecture diagram of the original post.)
2. Your encoder network now has access to the original HD video, so it can choose to encode the high-frequency details directly instead of generating them afterwards.