Nvidia's Crazy New Neural Engine Is Redefining Realism in Graphics
youtube.com
youtube.com
If you go back to that game today, there's a lot you can nitpick, but I think the biggest problem that remains today is that it's still extremely labor intensive to make assets.
At Planimeter, I worry more about asset creation turnaround time than anything else. I worry about it more than the tech we write, I worry about it more than fiction writing, or audio engineering. Just nothing else compares.
It's the most expensive part of any game development pipeline, and yet industry-wide accessible photogrammetry is half-backed, and it also doesn't help hobbyists who do 2.5D or 2D work.
So yeah, this neural engine stuff is superphenominal, for people who care about PBR workflows and photorealistic pipelines. This is obviously the most computationally complex work that can be addressed today, anyway, which I appreciate.
But the people who can actually access and harness that tech is just so absolutely tiny. I guess I just care a lot about this one particular sector of the industry where you have these world-class bedroom professionals who become studio professionals and Unreal and Nvidia are probably the only orgs in the world who cater to them, but for some reason I have to put in a lot of effort to articulate why what we have today, despite being so much more powerful than what we've had 20 years ago is less accessible and less functional and less empowering than what we had then.
I think it's primarily the labor factor in artwork, but also accessible engine tech today is actually worse than what we were using then, simply because no one really uses the id Tech family of engines anymore besides the most modern incarnations of them, and even to this day no class of engine compares to old versions of id Tech. Not even Unreal, by a long shot, despite having industry leading rendering capabilities. Everything else about that engine is half-baked or unusable.
Generating assets using AI will drastically cut costs and turnaround time. I'm eager to see these tools generate assets in real-time. Imagine being able to interact with an endless possibility of scenes. Games would ship with a generative model, and entire levels could be created on the fly, ensuring every play through is as different as the authors want it to be. Exciting stuff.
Fast content creation isn't, or the tools are sparse and locked behind small, costly proprietary silos.
Frank Gehry used to do really chaotic abstract drawings and hand them off to aerospace CAD operators to turn them into something that could be engineered and machined and assembled. Lots of people could do the drawings and have magic mushroom facades but Gehry understood the resulting interior spaces as well … I think it’s there is a ton of potential for great tools employing the past decade or so of ML/AI but ISTM it should empower the process rather than replace it.
I would have liked to see examples of human skin, and of cloth. That's hard and important. Rendering the perfect cheese grater and glazed pot is nice, but really, not that important.
Most of the hard problems in graphics today involve scale. Epic's Nanite texture compression system is impressive, but about 60% of the work has to be done in CPUs because GPUs don't have the right stuff for it. And Nanite is still just for rigid objects. Crowds of individually dressed people are still tough to render fast.
NVidia put in ray-tracing hardware into GPUs. It's not used much in games. NVidia was angry with reviewers who just ignored their ray-tracing hardware in reviews. The reviewers were not wrong.
And could we please have good hardware support for order-independent translucency? That should be standard. Then we can get rid of depth sorting of faces, which never works right.
Well, it's true that the proportion of games that have ray tracing -- over the space of all games released -- is small. But most triple-A games seem to have ray tracing these days. Of the last four such titles I've played three had ray tracing (Resident Evil 4, Jedi Survivor, Hogwarts Legacy). I think that's only going to become more true in the future, as IMO it does materially make an image quality difference.
We are at the third generation of cards with RTX and there are almost no games using ray traced global illumination, and even if they do they still have to offer rasterized lighting for older hardware.
I think you're saying that id Tech engine, even the old ones, are the best by a long shot; but some words in that paragraph seem to contradict that.
And despite being a fan of John Carmack's works, I understand that the Unreal engines have its own advantages.
Back at the time licensing from iD was basically "here's a cd, never call us again." Meanwhile Unreal, at a much more approachable price, also came with support contacts, a private email list and irc channel, etc. It was like night and day between the two.
In my role as a level designer I vastly preferred UnrealEd. The workflow was more interactive, and the BSP implementation way more robust against leaks. There were flaws to Unreal... the initial versions of UnrealScript were not great. But again, there was simply no comparing the two imo.
What you still can't scale... Art assets... Especially if you're not creatively talented in that way. There's a reason so many AAA games these days have staff that are like 80% artists, 10% developers, 10% everything else.
At least in F2P, the 'staff' are 80% artists because cosmetics are ludicrously profitable.
Let's take for example Second Life. The content is almost all user-generated, and when it originally came out it looked a lot worse than other games that were available at the time. The difference is mostly: since the content is user generated, you don't have some army of artists touching up the textures and pre-baking the lighting to make everything look at least plausibly realistic.
What Second Life was missing was the technology to compute global illumination in real time against arbitrary content so that the user-generated assets look as good as if some artist had specially lit and textured each one. (Maybe even better, because computers are more thorough and exact than humans -- assuming exactness is what you want. Artists sometimes take creative liberties because unrealism sometimes looks better.)
I haven't paid attention to Second Life in a long time; if a few minutes looking at random youtube videos is any indication, it seems they haven't really made much progress on this front. I think eventually companies will figure this out. I haven't paid attention to Facebook's Metaverse project because I expect it to be a flop (for non-technical reasons), but maybe they're banking on the graphics technology being there to make user-generated content look not-fake. I think that part might be achievable now with modest hardware.
As far as this new paper goes, I'm not sure if it makes things easier or harder for content creators. It seems like it makes the rendering stack more complex and harder to understand, but if someone can supply came creators with awesome textures and they don't have to care where they came from or how they work then maybe that's fine.
Theres also non-AI work that can help dynamically fill scenes with world meshes to help create seamless transitions between spaces as a part of overall terrain development.
A lot of this tech will be proprietary in nature, and it may be a long time, think decades, if ever that we get to use it as general creators.
It means that the gap between studios and bedroom professionals who used to create the mods that became full fledged intellectual properties that dominated for years will widen and there will be fewer opportunities for individual authors and small teams to become the next generation of production studios.
This will happen while large organizations enjoy extreme technical leverage while not having many great ideas to execute on simply because there will be fewer fiction writers and fewer studios who have concentrated technologies.
I currently don’t see much room for the next Zoid to create the next CTF gamemode.
I don’t really see anything meaningful come out of community contributions to games anymore ever since the introduction of the UGC model which has superseded the traditional modding model.
Instead of modders using game software mechanisms to create something new and interesting organically, UGC structures in games are explicitly designed so that creators can only do nominal things like create hats and skins for the most part.
It’s shifted the mentally from, “Here are the tools we used to create the game,” to “Here’s what you’re allowed to create and put on a marketplace so we can take a cut of profits.”
It’s gross, it sucks, and it’s another example of rug-pulling. Future generations will completely miss out on the previous paradigm because there’s so much focus entrenched on user-generated content versus modding.
But imagine the impact on rendered art, VFX, etc. How long until you can ask ChatGPT for a model of a teapot with shaders for blue glaze? The labour involved from inception to render is going to be cut down to a tiny fraction of what's required today.
I hope to see more and more bizarre ways of picking the author order on these kind of papers.
> equal contribution, three of the authors are the real authors, the other seven names were generated by an AI
> Order of authorship was determined by the top score each person received across all three levels of the official Pottermore “Back to Hogwarts Quiz.” For information about the exact scores (and houses), please contact the authors.
https://compass.onlinelibrary.wiley.com/doi/abs/10.1111/spc3...
For a while, [StackExchange.Academia](https://academia.stackexchange.com/) (like StackOverflow, but for questions about academic life) seemed to be getting a lot of questions about the meaning of author-order, where it seemed to vary across fields and even then be prone to subjective interpretation based on unstated rules. It was a very silly situation, made sillier by the fact that those doing it seemed to be under the misimpression that it was the serious, respectable way of conveying authors' relative contributions.
Of course, their particular approach to doing this does leave open one potential avenue for confusion. In particular, are they doing this because:
1. they're researchers who want to stress that the name-ordering is arbitrary in a humorous way; or
2. they're hardcore rock-paper-scissors competitors who were primarily motivated by a desire to better advertise the results of their tournament?
(Though, in the spirit of being clear-and-explicit, I'm just kidding about that last part!)
This still has to be bridged with compressed texture representations (and compressed model representations) but I think that's not too far away.
From my perspective this was the mainly the first contact with the research presentation and the renderings, and I think those are undeniably very impressive to most people at this point in time.
But viewing it as a "Youtube value-add", or whatever the jargon for that phenomenon might be, I guess it's close at hand to judge it as poor, lazy work that doesn't do the original material justice.
I will also admit it is something of a funny juxtaposition to have the photorealistic, high-definition graphics along with voiceover that is very obviously artificial.
I've been playing with TorToiSe and other emerging local projects with voice cloning, but ElevenLabs is so far ahead that it's the only one I've considered promoting to clients.
Sooner than you expect, games and virtual reality are going to look as good as movies -- scratch that, they are going to look better than movies.
We live in interesting times.
I say all this as someone with a Quest who tethers to a PC for Beat Saber and Half Life Alyx. Tethered experiences rule - but untethered ones are really not that far off.
I am not sure how you measure this but you can't run proper games on mobile vr. Mobile is limited, you can't fit in a GPU the same size of a desktop GPU. John Carmack has an interesting talk about what happens when mobile chips get too small and crowded. We simply can't battle physics. It would be awesome if we had the same experience tho.
I would argue that your thesis of "you can't run proper games on mobile vr" is wrong. Today, you can go play DOOM 3 or Half Life in VR, untethered, on a sub-$500 headset. That should startle everyone working on tethered systems.
I think there is a case for local ML becoming more popular too, I could see nvidia making a Shield like box at some point with a mobile 4000 series GPU and good thermals, that could bring those GPUs to mainstream beyond hardcore gamers. It would work for gaming, VR and consumer local ML apps (Siri that actually works and doesn’t leak data).
Maybe some day the latency/bandwidth will be good enough to stream VR from edge servers so you don’t even need a local GPU. We’re not there yet, even for non-VR games
I think we’ll see new classes of games/entertainment too where you’ll just describe the (VR) experience you want and ML constructs a game or (immersive) movie like experience of it. Maybe a different one every day.
I can see Nvidia selling a lot more GPUs in the next decade as ML and real-time 3D becomes much more pervasive. At some point other much more power efficient architectures (our brain uses 12 watts) will trump general purpose GPUs
That is indeed how i see it working. A dedicated VR "console" of sorts that tethers via wifi.
However,
*pauses to put on tinfoil hat*
Nvidia does sell devkits that roughly match the compute footprint you're describing[0]. They're ARM SOCs which puts them at a disadvantage for gaming, but the form factor does exist. If you need a lot of high-power AI compute and are willing to tinker with it, you can't beat CUDA on ARM.
Again though - there's a reason these are sold as devkits and not products. Every YC-backed roach from here to Mississippi is going to spend their next half-decade trying to get your data/money for a machine learning product. Nvidia knows it's a losing game to sell hardware instead of services here, so they're arming the entrepreneurs instead of the consumer. Frankly, I think it's the right move anyways. People are going to need scaling compute for decently fast AI inferencing in the future, and Nvidia can keep that scaling curve under their thumb. Hell, I wouldn't be surprised if there are Nvidia execs suggesting that they abandon the gaming market altogether just to focus on more lucrative AI/datacenter customers.
[0] https://store.nvidia.com/en-us/jetson/store/?page=1&limit=9&...
Movies will be made with this stuff. No expensive actors, sets, or long editing times that come with traditional animations.
--------------
Our system is running on Direct3D 12 using hardware-accelerated ray tracing through DirectX Raytracing (DXR). All results are generated on an NVIDIA GeForce RTX 4090 GPU at resolution 1920 × 1080
About one minute in I had a "Turbo Encabulator" feeling. So much so I'm not sure the original article is even real... I'm in no position to validate it without spending many hours scouring the content.
this gets less attraction, but I think it could be the next hit after LoRA.
I mean seriously, it looks like in the next maybe 5-15 years, GPU will be able to render graphics that even another GPU wouldn't be able to distinguish whether its real or fake.
The bulk of the progress will be in human-computer interfaces. We're still mostly interacting with keyboards, mice and 2D touchscreens; voice recognition is still primitive; XR is still experimental and nobody wants to strap a heavy headset to their face, etc.
Realistically, only the biggest triple A studio will be using this correctly and in the nice looking games that would benefit from this
aka procedurally generated textures. Clever for all the mundane repeating pattern/uniform surfaces.