Enhancing Photorealism Enhancement
intel-isl.github.io
intel-isl.github.io
These images are within a hair's breadth of being indistinguishable from reality. It's an enormous leap. An enormous leap from the previous techniques and an enormous leap towards photorealistic realtime graphics.
You don't want your games to look realistic? Go play pacman or outrun.
You want to use path tracing instead? You have no idea how much cheaper this technique is. And on top of that the neural network is also fixing up the unrealistic texture work on all kinds of things in the scene.
The first games to use this are going to be from small game studios with less investment in the status quo and they are going to blow the competition out of the water.
The cool thing is that this gives game-designers an extremely powerful additional tool. You could for example easily go for a more retro-look by training with kodakcolor images ...
And I don't buy the argument that GTA doesn't look more realistic because that's what the designers were going for. If this was a game with a clear "comic appearance" that might be true, however I would argue it looks like it does, because the designers did the maximum that they could achieve with their tools and within a certain budget.
What they've done is interesting, and I think with further refinement it is certainly a technique to consider.
Impressive result nevertheless. Improving this with adversarial networks will probably generate images which are indistinguishable from reality in just a few years.
Hot take: The overall tone of the style produced using this dataset comes from capture artifacts.
fps > beautiful images until you get at least 60... which this wasn't.
im pretty bearish on this kind of innovation in games tbh, its never going to be low latency enough to be viable, even if they'd improve their framerate - i dont see it going beyond academic interest because of that
definitely interesting to read about though, even if i doubt its viability in gaming
Their post processing can only start after the full image has been rendered, so whatever photorealism filter they apply will necessarily delay the image until the process is finished.
The blurriness, green tinge, and washed out colours are qualities photos sometimes have, but it's not what we see with our eyes. But they do help hide details that would give the image away as a computer rendering.
The creators aren’t necessarily aiming for photorealism. They want their games to look good enough and be entertaining. Adding more realism could make it look worse. Obviously if AI could solve the hard problems like global illumination that’s great. And I’m sure it’ll be possible to e.g easily add more variation to textures by using larger datasets. But changing the tonemapping to a colder palette, or just replacing the asphalt textures with smoother ones?
The creators of GTA didn’t use the sunset tones in the game by some creative mistake or lack of computing power. They obviously went for non-photorealism deliberately! It’s a stylized view of California. It’s like the burger in the commercial doesn’t look like the real one - it looks a bit unrealistically perfect because that’s actually more attractive than the real thing. Realistic and photorealistic are tangentially related.
The comparison here should perhaps be between frames of the game that already attempted to look like the cityscapes dataset.
GTA obviously doesn’t have a ground texture that isn’t smooth because it wouldn’t be possible, or “too warm” palette because it’s too computationally expensive to have the bleak cityscapes tone mapping.
A GTA frame rendered at its current frame cost, with textures the same size (but different) and a customized tone mapping, would look more photorealistic and make the difference to the cityscapes dataset much much smaller.
The authors could likely have done this comparison since they probably have the ability to modify textures and shaders to set tonemapping. That would have separated out the differences that are only artistic from those that are lacking photorealism because of computational expense.
The only thing I didn't agree is that Rockstar is incapable of improving their engine further. Since GTA-III, this is somewhat a deliberate stylistic choice up to a certain point, like how Source engine is also opinionated about how its games look.
Also, we need to keep in mind that this method is way more surgical than a simple post-processor. They've probably grafted a lot of tapping points to video driver to be able to implement this.
RDR2 is based on further development of the GTA5 engine and is arguably the best looking open world game at this time.
They didn't mention using anything else than output frames and G-buffers.
I'd love to be corrected if I'm wrong, BTW.
That's a silly dichotomy. We don't even want films to look completely realistic, that's why CGI exists in the first place. Imagine that you ran this over something like an Avengers movie, or even better, Blade Runner 2049. The AI would remove all the stylized/unrealistic details, and someone here would claim that this is clearly an improvement, and any criticism must come from luddites!
Styling as done in some of them is an optional bonus if creators strive for some additional artistic message, but otherwise its all about money, money, money.
You even contradict yourself - they tried to look Avengers as realistically as possible within the realm of comic book fairy tale, that's why they changed character of Thanos significantly well after introducing him (arguably a good choice). From cartoonish styling to realism. Nobody would take the epic battles so seriously if they all looked like Ready Player One.
Modern CGI in films is already basically indistinguishability from reality. You only notice it these days when it has to be CGI (e.g. space ships) or is badly done (which leads to the impression that CGI is still bad like it was in the 90s).
There are several, several youtube videos showcasing the incredible work of studios like Industrial Light and Magic creating most of the backdrops you thought were shot on location.
Some don’t want games to look this way. Nonetheless this technique can be applied to give games and movies a different look that would very hard to acheive by standard rendering pipelines, and it doesn’t have to look washed out like this example, that’s all based on the training data set and how much this ML lighting and texturing is blended into the scene, and that’s another creative choice. I can imagine filmmakers wanting to adjust CG lighting and texturing of a scene using a training data set from old TV shows to make a convincing reboot of a classic TV series, or new series done in that style.
What you predict about small studios seems unlikely as ML is a beast to tame, needs highly specialized engineers and huge high qualty datasets for training.
This. I thought those were photos they took somewhere that looks exactly the same as GTA V, but then when I watch the video explaining it I was like holy crap. No Ray Tracing or whatever tech demo was ever this close to photo realistic.
>You have no idea how much cheaper this technique is.
Could someone explain this in greater details? I suppose you will still need to create some decent base image or graphics for the enhancement to work to its maximum?
(also, this is simply one guy's opinion, but I am curious if anyone else feels the same)
For instance, in my view, I tend to prefer a stylized less hyper-realism in games. I think this is because the enhanced color, light, and brilliance is more aesthetically pleasant to look at, even though it is certainly more uncommon in the real world.
I don't believe this is because saturated color, and brilliant light doesn't exist in reality, but these occurrences happen more rarely and in specific, often planned, situations and settings. And some people (trained or gifted photographers, for instance), have learned how to capture this look, or to take something rather ordinary, and capture it to appear extraordinary, much like the difference between your average person snapping a photo with a disposable camera, and Ansel Adams. The former will produce a photo that is far more common in appearance, and thus...boring, or uninteresting. So when it comes to games, if you simulate the "average look of reality" the game certainly appears more photo-realistic, but may also be less aesthetically interesting (to me, at least).
This says absolutely nothing about the technology developed here, which is incredibly impressive. Nor do I think there is a right or wrong way create game visuals. I think that misses the point.
However, I feel I'm not the only person who can be blown away with how realistic something can be made to look, yet ...also find it more aesthetically drab and boring to look at. And in the case of games, as I mentioned above, I tend to prefer to see things that are rare, interesting, and extra-ordinary, and thus less "realistic".
Anyone else feel the same?
Without a comparison to the original pictures, they do look very good, the roads in particular look more realistic, as though taken from a dashcam or medium quality phone camera. However the foliage just looks blurred.
A lot of (interactive) CGI already looks more realistic than this. Take for example, this Australian scene made in Unreal Engine 4 - https://www.youtube.com/watch?v=Pg75bfkegtU
And that's not just pre-rendered, here's someone exploring the same scene interactively - https://www.youtube.com/watch?v=y87QL7IPE34
I want my games to look cool. I dont wanna them to look like streetview or geoguessr.
The result is impressive, but I'm not sure it would be good for video games. Maybe with some added filters it would be a lot better.
1. We do not know the set of pictures that the authors did not show us. So maybe this technique works really well only for a small fraction of cases.
2. Games (especially games that involve moving cars) tend to simplify mechanics so much that the moving end result will never look realistic, regardless how perfect the still picture is.
So yes, I expect great outcomes from such a method, but I don't think its going to be a revolutionary silver bullet.
As well, I’m very excited for the scrappy devs who will turn this technique in to something wildly unexpected.
I think this technique has a lot of potential, but will probably need a lot of work before it can be applied to video games. It might first show up as a filter for 'photo mode' in some games though, that would be interesting.
I can't just play GTA 5 as is?
Then, every AAA game looks dumb, non-interactive and limited even with HQ graphics.
I actually don’t think this looks any less “like a video game”, but its color is more washed-out, and darker. It seems like a net loss to me, sorry. Not for fancy theoretical reasons, I just think the GTA looks better in all the examples.
As others have said, using this directly in the engine (tuned to work with the intended art style etc) could probably produce almost miraculous results, if it can be made to work at a reasonable frame rate.
That would also allow the developers to use high-quality rendered images instead of these green-tinted, low-contrast "automotive grade" camera images as source of truth (there are good reasons these images look like that ... but they don't look pleasant).
Where it noticeably fell short was displaying red lights (tail lights, traffic lights, etc) that lost their brightness.
this would work great in a flight sim tho, albeit blur and out of focus effects are terrible for visual clarity
however, there's one small bit that has not been given enough attention imho, which is on the latter set, under "removes distant haze": this would be a great technique to complement flat LOD reduction of open world games while keeping up the crispness and contrast of distant items.
Get some good cameras and lenses and capture footage of wherever they want to set their next game. Then they can apply some motion picture color grading to that and use that as the input data for the NN.
Of course, this all depends on what the intention behind the game is. If you do want maximum immmersion, visionrealism would be the way to go.
Enhancements like this are going to be the thing that brings VR fully into the mainstream. It will be industry-changing.
This tech is in its infancy. I promise you, the final result is going to be indistinguishable from human vison.
There might also be a misconception that GTAV couldn't make their game more murky/realistic-looking, but more often than not, saturation and contrast is purposefully cranked up in games (particularly driving ones). I will say that the lush background mountains look cool though.
Define 'photorealism'. You limit yourself to the rendering aspect and ignore the content/asset problem and thereby economic factors.
This technique is probably one or two orders of magnitude cheaper than generating geometry and shaders that have the required level of detail in any traditional way.
Either by an artist or programmatic/procedural (Someone has to write that code too/set up that node graph in Houdini or the like). Yes, you can also just 3D scan stuff (see Quixel, etc.) but that has limits too.
More specifically for achieving photorealism by 'traditional means': consider scales.
An asset, however produced in finite time, will only hold up to some range of scales. The technique in the paper allows to get as close or far away as you want and everything will still look real. No level of detail work needed, just a few more training data.
Traditionally, if you use textures, you will reach the limit of the texture resolution. Even if you use procedural techniques – not everything is a fractal. A close up of a brick looks very different from a wall of bricks etc. Again, who will write that code?
My guess is rather that we will see a hybrid of more photorealism by traditional means and more 'icing on the cake' by methods like this.
Disclaimer: in the UK, so grey and murky == realistic.
Ditto in Northern Spain.
The sunny dataset is more akin to Mediterranean landscapes. To me the shinny/sunny/reflective stuff it looked unreal.
It's not just about murkiness and lighting quality, this is akin to adding the last 10% of polish that usually takes 90% of the time. It's also the difference between amateur vs much greater proficiency level. Having access to the assets of scene in 3D max, someone with a basic training can replicate base GTA scene, the enhancements shown here takes years more experience. The amount accessibility and saved labour is tremendous. It's why even VHS video is still more photo realistic than 4K RTX. Real life has a lot of minutiae and interacting details beyond resolution and good lighting models.
Especially the GTAV dataset is often used as a synthetic dataset for research. The appeal is that you can extract normal/material and other semantic information from the game engine for a driving scene.
Basically avoid expensive segmentation labels on real data by working the other way around. Things like this synthetic to real translation can then be used for other downstream tasks (semantic segmentation, distance/normal prediction).
I wonder how much complexity could be saved by letting the application properly annotate objects: "This is a tree located at coordinates X,Y,Z in the scene. Please NN take over and project a tree appropriate for the light conditions and geographical location."
Further, I’d love to also see a tech demo using a Californian dataset as the input: To my originally European eyes, it’s simply making GTA look more like video shot in Germany vs more realistic (there’s a difference!), though I get that any data set can be fed into this :)
[1]: https://www.youtube.com/watch?v=P1IcaBn3ej0 (0:19)
"Inference with our approach in its current unoptimized implementation takes half a second on a Geforce RTX 3090 GPU."
That means 1 or 2 frames per second, so interactive but pretty much unplayable, which is to be expected as it's just research, but considering that they have DLSS working in realtime with tensor cores perhaps something like this will also be achievable very soon.
Edit: Out of curiosity I looked up how long DLSS takes to process and it's less than 2ms, so you'd probably need to speed this up by two orders of magnitude to run it a game.
“We also seek to eliminate artifacts that can be seen in the results of prior deep-learning approaches, which often hallucinate objects. To this end, we analyze the datasets that are commonly used for photorealism enhancement. Our analysis reveals that their scene layouts differ in ways that can explain artifacts commonly seen in prior work. To better align the datasets and alleviate the artifacts, we propose a new strategy for sampling image patches during training. We further design a new adversarial training objective that facilitates enhancements that are geometrically and semantically consistent with the content of the input image.”
The image enhancement here is happening "in camera", like a fragment shader; but it's using a lot more context than the sort of local/cellular kernels that fragment shaders usually are.
Look at e.g. the gutters on the roads, where there are leaves/dirt — they get filled in with tons of extra perceptual texture (texture that exists in the projection plane, rather than texture that exists "on" the road), bringing them up to the same LoD as the rest of the road.
If that texture was "pushed back" to the origin texture, said texture would have to exist at some sort of ridiculous resolution that modern graphics cards wouldn't have a hope of rendering.
So, rather than rendering the road at 8x-16x and then downsampling it, in order to get a pile of leaves to appear to have ~1.5x LoD at the projection plane, you can just achieve that 1.5x LoD pixel by pixel at the projection plane.
Or, to put that another way: when a digital artist is painting a landscape, they don't have to zoom in by 8x, paint every individual blade of grass, and then zoom out. They can just use an artistic technique at 1x zoom that approximates the texture that would have been created by downsampling "photographic" detail painted in at 8x to your screen resolution.
Here, the ML model is playing the role of the digital artist, "painting" on the projection plane.
To do that, you’d need access to the game’s source code and assets, and you would need to rewrite portions of the rendering engine. Training might be incredibly slow if it needs to iterate on builds of the game, or that problem could be worked around but might need considerable engineering to support the training iterations (e.g. data representation and a pipeline with the ability to render changes to the artwork and renderer on the fly without needing a game rebuild).
So I’d speculate the answer to why not is that it’s just not what the authors set out to do, that they don’t have access to the game source nor the time or resources needed to refactor the renderer.
Part of the magic of neural networks is their black box properties - they can do what they do without needing to understand 3d geometry or integrate with a render engine. Throw an image in, and a new image pops out, without the neural network needing to understand what it’s doing or why.
That said, I would guess that coming down the road is examples of exactly what you’re suggesting - games that will use neural networks to drive realism in the assets and shaders and renderer. To some degree there are already tools starting to do this.
Maybe GTA 6 or 7 will include no textures and lower poly models, and instead spend a large budget driving around LA and just film stuff to train neural render.
Obviously the GTA V designers applied a particular style and colour scheme which gets lost, but I'm sure something like StyleGAN could address that.
In addition, the enhancement blurs a bit, which imo is not good; the blur adds to photorealism by hiding mistakes, but on the other hand it makes it more obvious that you are looking at a picture because you can't see details.
I would argue that these kinds of renders are getting good enough that we should be more precise about what is meant by photorealism. Most pictures i take with my phone have alot of blur, don't capture colors well, and are frankly quite crappy. I feel like people expect photorealism to be equivalent to a really well-planned photo by a professional.
I've always wondered how games were going to solve the last mile of photorealism, and an additional color-grading filter seems like the perfect step. Really impressive work.
I think titling it "photorealistic enhancement" is a bit of a whiff, but "photorealistic real-time color grading / filtering" is bad-ass.
GTA uses some surrealist contrasts because it makes the games easier to play, especially for long hours, or when you go fast.
But the tech is so new I'm not going to complain, it's already very good, especially given what it eats in resources.
[0]: https://en.m.wikipedia.org/wiki/Physically_based_rendering
This neural network operates on images - input is image and output is image. Note that it cannot synthesize an image from scratch, it can only take an image and make it look more like the lighting in it’s training set.
PBR is a technique for taking a 3d scene description and synthesizing an image, so for PBR you need models and lighting and materials, and even with physically based techniques it might still come out not completely realistic looking. There’s always room to emulate realistic subtleties of texture, lighting, camera lenses, video/film response, outdoor colors, atmosphere, etc.
So where previously, you needed a lot of designers to make a big world look realistic with PBR, now you need to collect a big dataset with the style you want to have, and apply it to your game. I suspect that will scale differently.
But why does every car look like it's been freshly waxed? The newspaper boxes have a bit of rust and grime on them. Those look really real.
Looks great in overcast settings but has some difficulties with sunny whether.
Pretty impressive.
http://youbentmywookie.com/wookie/gallery/0910_wtf/fast_food...
They made the game look bad with shitty fps.
If GTA developers wanted they could make the game look photo-realistic but that was never the intention.
It's impressive for what it represents technically, but it's visually subtle.
Lol, I guess you’re new? Or maybe you don’t know the difference between 3D rendering software and game engines?
GTA V came out in 2013 and the fact that it is still used as a reference for graphics quality tells you everything you need to know.
The point is GTAV, which was/is exemplar for game engine graphics, and even current gen RTX game engine rendering all look "bad" due to constraints and budgets of real time rendering. Bad in the sense that they're necessarily rudimentary - actual photo realism requires many layers of expensive details. This technique fills in many of those kind of details. The improvement in photo realism demonstrated here is significant and a huge leap in terms of image quality.