* deferred vs forward rendering (deferred adds latency)
* multithreaded vs singlethreaded
* vsync (double buffering)
https://www.youtube.com/watch?v=8uYMPszn4Z8 -- check at 6:30 the latency of 60fps vsync on 60hz. It's not even close to 16ms (1/60), it's ~118ms (7.1/60).It's far cry from simplified pure math people think of when they think of fps in games or refresh rate for office and typing. Software is very very lazy lately, and most of the time these issues are being fixed by throwing more hardware at it, not fixing the code.
Some things cannot be 'fixed'. It's always a trade-off. You can't expect to have all the fancy effects that rely on multiple frames and also low latency.
If there was a simple software fix, GPU manufacturers would be all over it and pushing it to all engines. It's in their interests to have the lowest latency possible to attract the more hard-core gamers (which then influence others).
Just look at all the industry cooperation that had to happen to implement adaptive sync. That goes all the way from game developers, engines, GPUs, monitors. Sure that sells more hardware(which brings other benefits), but a software-only approach would also allow companies to sell hardware, by virtue of their "optimized" drivers.
Wah? Deferred just refers to a screen space shading technique but it still happens once every frame.
> * multithreaded vs singlethreaded
Not sure what you're saying here.
And then of course, yes display buffering does have an impact.
In principle, you can push GPU pipelines to very low latencies. Continually uploading input and other state asynchronously and rendering from the most recent snapshot (with some interpolation or extrapolation as needed for smoothing out temporal jitter) can get you down to total application-induced latencies below 10ms. Even less with architectures that decouple shading and projection.
Doing this requires leaving the traditional 'CPU figures out what needs to be drawn and submits a bunch of draw calls' model, though. The GPU needs to have everything it needs to determine what to draw on its own. If using the usual graphics pipeline, that would mean all frustum/occlusion culling and draw command generation happens on the GPU, and the CPU simply submits indirect calls that tell the GPU "go draw whatever is in this other buffer that you put together".
This is something I'm working on at the moment, and the one downside is that other games that don't try to clamp down on latency now cause a subtle but continuous mild frustration.
https://github.com/microsoft/terminal/blob/main/src/renderer...
The first comment in this function (DxEngine::StartPaint), for example:
// If retro terminal effects are on, we must invalidate everything for them to draw correctly.
// Yes, this will further impact the performance of retro terminal effects.
// But we're talking about running the entire display pipeline through a shader for
// cosmetic effect, so performance isn't likely the top concern with this feature.That said, the VR space has a much tighter tolerance on input lag and there's hardware based mitigations. Oculus has a lot of techniques such as "Asynchronous Spacewarp" which will calculate intermediate frames based on head movement(an input) and movement vectors storing the velocity of each pixel. They also have APIs to mark layers as head locked or free motion etc.