I've never had hardware to run above 1080p at 60fps (monitors, GPUs), so my desires are a bit orthogonal - I wish more games let you customize the graphics settings so that you could maintain a consistent framerate (be it 30 on a low spec laptop, 60 on a regular PC or more fps) at a resolution of your choice.
Things like switching between baked or dynamic shadows, their type/filtering, things like ambient occlusion and tessellation, various post processing effects and filters, model LOD limits and texture resoltion, resolution scaling and so on.
More so, it would actually be nice if games let you download a version with lower fidelity assets (like War Thunder sort of does, for example), so you don't need 100 GB of models and textures if you're realistically only ever going to see 20 GB of those on your hardware.
Thankfully, most modern game engines scale both up and down decently, for example, Unity's URP (though historically a bit half baked and fragmented the community/assets). It's just up to the developers to get over the hubris of wanting "low" settings to still be pretty, that choice should be up to the user.
No. Please don't spread this nonsense rumor. It hasn't ever been true and still isn't.
There are always diminishing returns, but human vision is perfectly capable of noticing the difference between 60 and say, 120. You can literally try this yourself just moving your mouse around on a 120Hz monitor and then capping the refresh rate. In general human vision is really complicated and not tied to any specific "speed".
Also, displays are more complicated. There isn't just a refresh rate, there's differing pixel responses at differing brightness and a variety of different processing delays all the way from the moment you put in some hardware input (like the mouse), to the final result. It's a huge chain.
The end result is that two 120hz displays can behave completely differently in terms of motion clarity.
Do you have sources that show otherwise and back up your claim that this idea is “nonsense”? I will dig up some scientific sources. Are you perhaps reading into my comment and not responding to what I said literally?
You aren’t really addressing what was main point: that fps throughput isn’t the reason for high frame rates in games. The primary reason for this happening is to decrease latency.
We wouldn’t need 500fps for games if we lowered the latency. Or at the very least, the benefits would be much lower. Reducing latency is a great reason to want high fps, but there are other ways to reduce latency.
You replied to the wrong comment, btw. I almost didn’t catch your reply.
Edit: links that discuss the measured speed of human perception:
http://web.cs.wpi.edu/~claypool/papers/fr/fulltext.pdf
(Study is limited to 60 fps, but clearly shows the trend of diminishing returns.)
https://www.healthline.com/health/human-eye-fps
(Argues it’s higher than 60 for some tasks… up to 90 or 100fps.)
https://www.rtings.com/monitor/learn/60hz-vs-144hz-vs-240hz
“there are diminishing returns when it comes to the refresh rate. Most people can perceive improvements to smoothness and responsiveness up to around 240Hz; however, the difference between a 240Hz and 360Hz panel is so small that even competitive gamers might have a hard time telling them apart. If you have a choice between a 1440p 240Hz and a 1080p 360Hz monitor, you're probably better off getting the 1440p option, as the increase in resolution has a much larger impact on the overall user experience.”
https://www.neurotrackerx.com/post/5-answers-to-the-speed-li...
(Mentions seeing a single frame of color at 500fps. This is true! Humans can see a flash of light that’s much shorter than 2ms. For perceiving imagery and tracking motion, the available evidence shows little benefit to going higher than 120fps.)
>my statement that 500 fps is not 5x better than 100 fps to a human
I never said that. I also stated that I know about diminishing returns. You claimed "There’s a good reason to never go above 70fp". There absolutely is a large difference in motion clarity between 70 and something like 120/240/etc. Is the difference from 70->120 as large as 30->60? No, absolutely not. But it is significant and can be seen easily with a cheap monitor.
>You aren’t really addressing what was main point: that fps throughput isn’t the reason for high frame rates in games. The primary reason for this happening is to decrease latency.
>We wouldn’t need 500fps for games if we lowered the latency. Or at the very least, the benefits would be much lower. Reducing latency is a great reason to want high fps, but there are other ways to reduce latency.
While there is a latency improvement and some people care about that, the motion clarity is also significantly improved. (again, just drag some windows around on a 120hz monitor) That's true of movies just as well as games. Movies have pulled a lot of tricks to mask this issue over the years, but 60 is quickly becoming the standard over 30. (And once bandwidth and processing improves, it's likely some day decades from now it will jump even higher)
>You replied to the wrong comment, btw. I almost didn’t catch your reply.
Yeah, not sure how that happened.
When you said “No. Please don't spread this nonsense rumor. It hasn't ever been true and still isn't.”, combined with the downvote, I assumed you were referring and objecting to everything I said including diminishing returns (even though I see you acknowledging it next paragraph.)
We are mostly agreeing violently, I acknowledge that there’s no known hard fps threshold above which nobody can see something. I acknowledge that there are benefits above 60fps, even if they grow smaller.
But it’s still true that the primary reason games are going to 500fps is for the latency benefits, not for the smoothness or high flicker rate. A frame rate that high isn’t generally perceptible, while the latency of today’s games - a latency of multiple frames - is actually well inside the known measurable threshold of response times. The problem isn’t generally the need for more frames per second, the big problem is the time between input and the visible change on screen.
The other topic that would be nicer to discuss is the quality trade offs. High frame rate takes away from other options.
TAA fundamentally decouples the sample generation from the output generation. TAA samples (even native res) have no direct correspondence to the output grid, they are subpixel sampled and jittered. The output grid takes the samples nearby and uses them to generate an output for each pixel.
TAAU/DLSS2 take this further and decouple the input grid from the output grid resolutions entirely. So now you have a 720p grid feeding a 1080p grid or whatever. Thinking of it as "input/output frames" is clunky especially when (again) the input frame doesn't even correspond directly to the input samples - they're still jittered etc. Think of it as a 720p grid overlaid over a 1080p grid, and DLSS is the transform function between these two (real/continuous) spaces. Samples are randomly (or purposefully) thrown onto that grid.
OLEDs are functionally capable of individual pixel addressing if we wanted to. Current ones probably can scan out lines in non-linear order already, so you could scan out the center more than the top/bottom, for example. And OLEDs are already capable of effectively >1000hz real-world response time from their pixels. They just are absolutely critically bottlenecked by the ability to shove pixels through the link and monitor controller quickly enough (give me 540p 2000hz mode, you cowards)
This all leads to a question of why you are still sampling and transmitting the image uniformly. If there's parts of the image that are moving faster, render those areas faster, and with more samples! Maybe you render the sword moving at 200fps but the clouds only render at 8fps. And Optical Flow Accelerator can also allow you to identify movement within the raster output and correlate this with input - if you think about oculus framewarp, what happens if we framewarped just one object? Translate it around against the background, stretch it to simulate some motion aspect, etc. And if you further refine that to a 1x1 region, then you have individual pixel framewarp.
You can also have a neural net which suggests which regions are best to render next, for "optimal" total-image quality after the upscaling/framewarp stages, based on object motion across the frame and knowledge of the temporal buffer history depth in a particular pixel/region (and this is the sort of abstract, difficult optimization problem ML is great at). And those pixels actually might be rendered at multiple resolutions in the same frame - there might be "canary pixels" rendered at low resolution that check whether an area has changed at all, before you bother slapping a bunch of samples into it, or you might render superfine samples around a high-motion/high-temporal-frequency area (fences!). So now you have "profile guided rendering" based on DLSS metrics/OFA analysis of the actual scene being rendered right now.
Another random benefit is that "missed frames" lose all meaning. The input side just continues generating input for as long as possible, as many samples at whatever places will be most efficient. If it's not fully done generating input samples when it comes time to start generating output... oh well, some cloud has less temporal history/fewer samples. But DLSS does fine with that! As long as you still have some history for that region you will get some output, and I'm sure if it was important then it'll be first thing scheduled in the next frame.
You can use this idiom with traditional scanout/raster-line monitors, but ideally to take advantage of OLED's ability to draw specific pixels, you probably end up with something that looks a lot like a realtime video codec - you're sending macroblocks/coding tree units that target a specific region of the image with updates. Or you use something like Delta Compression encoding, like with lossless texture compression.
https://en.wikipedia.org/wiki/Coding_tree_unit
This already is sort of what Display Stream Compression is doing, but that's just runlength encoding, and if the monitor can draw arbitrary pixels at will (rather than being limited to lines) then you can do better.
And again, you can't transmit the whole image at 1000fps, but you can target the "minimum error approximation" so that on the whole the image is as correct as possible, considering both high-motion and low-motion areas. This is sort of an example of the concept - notice the scanline errors in high-motion areas. Great demo but linked to the direct example:
https://youtu.be/MWdG413nNkI?t=176
https://trixter.oldskool.org/2014/06/20/8088-domination-post...
This also raises the extremely cursed idea of "compiled sprites" inside the monitor controller - instead of just running a codec, you could put a processor on the other side (like the G-Sync FPGA module) and what you send is actually the program that draws the video you want. Executable Stream Compression, if you will. ;)
(but there's no reason you couldn't put THUMB or RISC-V style compressed instructions in this - and you certainly could change up the processor architecture however you want, as long as you don't mind doing a Gsync-like compatibility story. It makes upgrading capabilities a lot easier if you control both ends of the pipe, that's why NVIDIA did g-sync in the first place! And there is probably nothing more flexible or powerful than allowing arbitrary instruction streams to operate on the framebuffer (or a history buffer, whether frame history or macroblock/instruction history) or to draw directly to the frame itself. With the bottlenecks that DP2.0 and HDMI2.1 present, this is probably the way forward if you want lots more bandwidth out of a given link speed.)
I said a long time ago that this is basically a "JIT/runtime" sort of approach and people kinda laughed or said they didn't get it. But it's funny that trixter actually used the same analogy there for his demo. But basically DLSS is an engine unto itself already, DLSS is what does the rasterizing, and the engine just feeds samples (that are entirely disconnected from the output, they go into the black box and that's the end of it). And with fractional rendering you can basically view that as a big JIT/runtime. DLSS provides quality-of-service to pixels by choosing which ones to schedule for "execution" and with what time quanta, to produce optimal image quality ("total QOS") over the whole image. And then the resulting macroblocks are individually sent over when ready.
Effectively this works like the biological eye - there is no "frame", changes happen dynamically across the whole frame constantly. Or another analogy would be "what if we could 'chase the beam' but across arbitrary pixels in the frame"? If the motion is predictable then you can schedule the rendering so that the output and blit happens at the exact proper time, down to 0.1ms accuracy. It's Reflex for Pixels.
DLSS knows when a person is running and about to peek out from behind a wall. DLSS knows when someone was sitting there and can render the lowest-latency-possible update when they suddenly pop out of cover and AWP you (can't read minds/make network latency disappear, but it can update it as quick as it can). That already shakes out of the motion vector data and just needs to be generalized to a world where you can render out specific regions/lines at 1000hz.
(I guess technically "frame" still exists as a notional concept inside the game loop, you are generating game state in 60fps intervals or whatever, but rendering is totally decoupled from that and you run the GPU on whatever parts of the image would benefit the most from touchups at a particular moment.)
But again you can see how this whole idea fundamentally inverts control of the game engine - Reflex already tells the game when to sleep and when to start processing the next game-loop, now DLSS will tell the engine what pixels it wants sampled, and handles rasterizing the samples and compressing/blitting them to the monitor. The engine is just providing some "ground truth" visual input samples and calling hooks in DLSS.
(sorry, long post, but I've been musing about this elsewhere and I looooove to chitchat lol)
> you are generating game state in 60fps intervals or whatever
This is also an excellent point that might question my suggestion that high frame rate is being used to reduce latency. Games state updates are already decoupled from rendering in lots of games. Having an extremely high render refresh might not mean that the latency between controls and visuals is reduced proportionally. Or maybe it helps but has a limit to how much.
DLSS is an interesting topic here. Do you see it eventually working for fractional updates? We’d possibly need a new style of NN or of inference? DLSS currently operates on a full frame, and the new version even hallucinates interpolated frames to boost fps artificially. This doesn’t help with control latency at all, in fact it makes it worse.
Yeah I've been playing fast and loose with terminology here, there's several overlapping but synergistic ideas that aren't the same things. To try and clean this up:
DLSS1 required the full image, it actually was an image-to-image transformer that "hallucinated" a full-res image from an input image. This sucked and NVIDIA gave it up (except for the 500mb of DLSS models that live eternally in the driver for the 2 games that opted for driver-level model distribution). Nobody cares about this anymore at all.
DLSS2 does not need a full image, because it is a TAAU algorithm that weights samples using a ML model. If there aren't enough samples in an area oh well, you just get crappy output (like immediately after a scene change). It manifests as either obvious resolution pop/detail pop, or visual artifacts on moving/high-res things. IIRC this can be assisted by drawing invisible (1% alpha) objects in motion/fences/etc to "warm up" DLSS sample history on the (invisible) edges before just popping them into existence iirc lol.
You need at least some samples near that area, it can't render from nothing, but DLSS2 is not dependent on rendering out the full image to work - if some unrelated part of the image doesn't have samples, oh well. (and this may allow mGPU scaling with reasonable correctness for partitioning an image!). Personally I consider this "loosy goosy correctness" attribute of ML models to be extremely desirable for GPGPU programming - if some edge case messes up 0.1% of samples, the ability of ML models to just ride over it and spit out a reasonable output is super desirable. This includes things like camera noise, dead pixels, etc. Extremely tolerant of data ingest etc. Like if 10 threads aren't quite finished with their sample output because you don't want to wait for kernelfence sync when it's time to start rendering the output buffer... just start going. It'll be fine.
Fractional rendering is a separate and unrelated idea, but I think the time is right with OLEDs here, and with everyone searching for a way to extend perf/tr with costs spiraling it makes sense to see if you can render "better" imo.
Variable rate sampling is another concept that builds on fractional rendering. Render some areas at a higher rate than others. And again this is something that DLSS2 plays nicely with.
--
DLSS3 is actually the successor to DLSS2 and is supported by all RTX cards (yup). Framegen is one of the features in this, and that is only supported on Ada. Supporting framegen requires the inclusion of Reflex, which does benefit everyone hugely.
Reflex basically flips the "render+wait" model to be a "wait+render", by adding a wait at the start of the game loop that delays until the last possible second to start processing the frame, so it's as fresh (input latency) as possible. And this does legitimately cut latency significantly (by ~half) in highly gpu-bound scenarios. And that gives NVIDIA some headroom to play with in framegen tbh. Igor's Lab and Battlenonsense both found NVIDIA to have much lower click-to-photon latency than AMD Antilag, by like 20-30ms in overwatch f.ex.
Framegen as currently implemented is interpolation and yeah that does increase latency. But NVIDIA do have some headroom to play with there, in CPU-bound situations (which is different!). And tbh most people who actually have used it generally seem to find it not too bad, it's the "eew I tried it at the store and it was awful!"/"i've never tried it" who are most vocal about the latency. It's at least an option in the toolbox (see again: starfield).
I think it is possible to move to extrapolation and I hope the current framegen is only an intermediate step. And I think Optical Flow Accelerator is a really cool building block for that. The performance and precision has improved a bunch over the gens, and now it can support 1x1 object tracking (which I mentioned above as seeming like a significant threshold/milestone) so it's flowing pixels really. I see that as being a Tensor Core-like moment that people scoff at but has big implications in hindsight. Being able to incorporate realtime image data back into the upscaling/TAA pipeline seems big even beyond just framegen itself, I don't doubt DLSS3.5 will make further progress too.
You don't need extrapolation (or interpolation) at all. but if you can extrapolate per-pixel, the ability to do a low-cost "spacewarp" that accomplishes most of the squeeze of a full re-render (in terms of moving edges/texture blocks) at much lower cost could would be very interesting. And the OFA could end up being a key building block in that sort of thing.
--
Again kind of a topic shift but there's also this issue of display connection (there is never enough bandwidth) and whether it's lines or macroblocks etc. That's a capability that's offered by OLEDs in theory, and could be explored with a similar FPGA approach/etc. If you can do that, it pairs with variable rate sampling concepts (and ML input tolerance for bad data) very nicely - render out the regions you're updating whenever they're ready, or whenever is optimal for that element to be drawn (to get minimum error).
And in fact quite a few of these ideas synergize nicely together. If you put it all together.
--
These are all kinda separate in general but quite a few of them synergize if you put them together. And I think the zeitgeist is ripe on some, brian heemskirk was talking about some similar ideas on MLID's show a few months ago (not the most recent appearance).
https://www.igorslab.de/wp-content/uploads/2023/04/Overwatch...
https://youtu.be/7DPqtPFX4xo?t=727
(highly pronounced in this game but)
Yes I agree that driver overhead+latency reduction in the pipeline matters a lot and that's the tool Intel just dropped (and has probably optimized their own driver for of course). Classic Tom Peterson, lol, just like FCAT.
Of course new games are running at 70 fps. Most games are absolutely not designed for 500hz, and have zero reason to do that. Cs:go wan’t designed for 500fps, I bet the designers of cs:go never intended people to play at 500fps, or even imagined that would ever happen; it wasn’t possible when the game came out.
There’s a good reason to never go above 70fps: human perception experts tend to agree that our visual system gives us quickly diminishing returns above 60fps, we can’t really see things any faster than that, so I’m quite skeptical that 500fps is necessary.
The whole reason 500fps is useful for competitive esports gaming is to reduce latency, not really to keep increasing frame rates forever. Because games are triple-buffering and sometimes monitors are too, and there’s another frame of latency for controller inputs to be recognized, you might still have 10ms or more of latency between controller input and changes on-screen, even if your game is rendering at 500hz. This means you could get away with 100fps if you had zero latency. When your fps is 60hz and there’s 5 frames of latency, you don’t see responses to your controller until almost 100ms later(!).
It’s been about a decade since I was a game dev, but it’s wild to me that anyone would want 500 fps, or that any games would aim for that. This gives you a grand total of 2 milliseconds do to everything, gameplay + animation + physics + audio + rendering. It’s not a lot of time, and you have to compromise your visuals and rendering (by 10x!) in order to achieve that frame rate. Looks like a lot of new gaming monitors are 144hz, and these new 240/360/480hz monitors are pretty extreme and still a bit rare.
> I understand that some games purposefully limit FPS for physics calculations, but there are many that do not.
FWIW, this isn’t the way I’d frame it, it’s kinda misleading. Very few games are purposefully limiting FPS for the sake of slowing it down. All games, however, have a budget. There are always limited compute & render resources, and both game devs and players want the highest quality available for their budget. Until very recently most games aimed for 30fps, and only really fast twitchy games went for a relatively very smooth 60fps. The games that went for 60 had to cut their polygon counts and physics and gameplay in half in order to achieve high frame rate, so you are totally trading away a richer experience in favor of high frame rates. Only certain kinds of games should even try to do that.
The whole "science says the human eye cannot see beyond 30fps :)" is actually a very old meme at this point. You are correct that high FPS/Hz is about decreasing latency, but you underestimate the importance of it. 144hz is a very clear improvement over 60hz and the PC ecosystem is going to move over to it as a new standard relatively soon. You don't need to be playing a hyper-competitive FPS to notice it either, any game with significant movement will do.
> Because games are triple-buffering and sometimes monitors are too, and there’s another frame of latency for controller inputs to be recognized, you might still have 10ms or more of latency between controller input and changes on-screen, even if your game is rendering at 500hz. This means you could get away with 100fps if you had zero latency. When your fps is 60hz and there’s 5 frames of latency, you don’t see responses to your controller until almost 100ms later(!).
You're assuming VSync or something. There's no guarantee that the input will line up with the next frame, so an input latency of 10ms while running at 100 fps equals a worst-case latency of roughly 20ms, not 10ms. That's why the higher fps & hz numbers really matter.
> an input latency of 10ms while running at 100fps equals a worst-case latency of roughly 20ms, not 10ms. That’s why the higher fps & hz numbers really matter.
Yes exactly, I agree and this was the point I was trying and I guess failing to make. Latency is typically multiple frames, so at 100hz, latency can easily be longer than known science-meme perception times. Even with vsync there’s up to one frame to recognize inputs, then another frame to produce new game state & submit the render, then 1 or 2 more for double or triple buffering, then maybe supersampling and/or denoising, then whatever the monitor does, and I might be missing some steps that add latency. So I suspect 500hz has nothing to do with seeing smoother motion compared to 144hz, and everything to do with getting overall control latency down to below, say, 10ms.
Does that really work though? Does CS:GO and do other modern games poll the controller at the display refresh rate? I know it’s pretty common for a lot of games to decouple rendering from game state. That can mean a lot of different things, but one of the implications might be that controller latency is limited to one or two frames of, say, 60hz game state updates, followed by 3-5 frames of 500hz display refresh. Is latency in today’s PC games limited by the game state update, regardless of what the display refresh rate is?
I’m curious where the perceptual limits really are, and what the max framerate we actually need is. Suppose an imaginary world where worst case input-to-screen latency is 2 frames. Then in that case, how high should the FPS be? What if there was zero latency - like if the next rendered frame magically reflected mid-frame controller inputs and magically displayed with no latency - then what should the ideal FPS be? Would there be any benefit to going higher than 144, or would we be wasting electrons?
VSync introduces latency, so it always gets disabled, same with double/triple buffering. I don't think gaming monitors do anything to process the image, they're usually advertised with ~1ms response times.
> Does CS:GO and do other modern games poll the controller at the display refresh rate?
Modern gaming mice and possibly keyboards use 1000hz polling.
> I know it’s pretty common for a lot of games to decouple rendering from game state.
I think all modern games do. Some old games have issues running at fps other than 60 due to this coupling.
> I’m curious where the perceptual limits really are, and what the max framerate we actually need is.
I don't know exactly; there are definitely diminishing returns, like you've mentioned:
- 30fps: 33ms / frame
- 60fps: 17ms / frame (2x cost for -16 ms)
- 144fps: 7ms / frame (2.4x cost for -10 ms)
- 240fps: 4ms / frame (1.6x cost for -3 ms)
- 500fps: 2ms / frame (2x cost for -2 ms)
- 1000fps: 1ms / frame (2x cost for -1 ms)
Just from looking at those numbers, 144fps is still a clear win and 240fps probably makes sense too, if you've optimized your setup for latency. Everything beyond is probably pointless, unless it's your job to play competitive FPS games ;)
Graphical fidelity in games improves until it maxes out current hardware. This is natural. You can always turn the settings down to get higher FPS, but you can't turn down the high FPS you get in older games into more modern graphics.
> Avg FPS are not going up fast enough in my opinion.
FPS isn't the one and only metric. Audiences generally care more about graphical fidelity. FPS just has to be good enough and stable. Typically that sweet spot is 60fps, though we're slowly moving to 144fps being the standard.
A much better benchmark would be to take a game designed for 120 or fewer FPS on a 4090 and try it on a 1070.
I expect that if you have a 4090, you have an Intel or AMD CPU that exposes a core clock multiplier. You could run this benchmark with whatever value it's at, then reduce the modifier by say, half. That should halve your CPU's clock rate, and I'm guessing you'll see the frame rate decline similarly. You can conclude that the game is "CPU bound", then.
Even if your GPU was infinitely fast, you have to remember that game developers are not optimizing for the game event loop to run in sub-millisecond times. If the core loop in the game takes under 16ms to run on a common CPU, such as one in a console: that's better than 60hz and the overwhelming majority of video game players will never see a benefit.
Some game developers, I see someone in another thread mentioned Doom Eternal, pride themselves on that optimization. With a fast enough CPU and GPU, you could probably reach 1000 FPS on Doom Eternal. A quick search suggests this might have been accomplished with a liquid cooled, this was done with a liquid nitrogen cooled PC with a 6.6GHz CPU: https://www.pcgamer.com/heres-doom-eternal-running-at-1000-f...
How does a 1070 fare versus a 4090 in Hogwarts Legacy or another modern game at 1440 or 4K?
My guess is that we'll see improvements at closer to the current rate rather than at an increasing rate.