You mention one great technique yourself already.
You mention one great technique yourself already.
What does this have to do with anything? Why do you want to restrict yourself to such an austere setting?
I believe their point was that this can't be done, since the "local simulation" is no longer running locally; it has to resolve & hide the latency _before_ it's transmitted to you, and you don't know what latency that frame will observe.
For a slightly silly example, see what mosh is doing to hide ssh latency.
Not necessarily, why?
The generalised idea that I see is: 'what can you do with beefy processing in the cloud (but behind a latency hop) and a very small amount of processing power locally?'
A VR example: you do most of the rendering on the server, but you do all the correction for small head movements locally. The server has to send a bit more (and not just as a flat image), but you get over the nausea.
Something like this has been state of the art for ages. And I only say 'something like this' because I don't remember the details.
As a generalised solution if you have crazy amounts of bandwidth to burn: the server is a latency hop behind, so it could send all possible futures that correspond to the inputs that might have happened during that latency hop, and the client picks what to display.
As written, this is crazy. But if instead of all possible futures you have a probability model of human input, and you send the most likely futures, you can probably make it work. Especially if you don't insist on sending raw video. (And even with raw video, you can save a lot of bandwidth by using a special codec that can compress multiple video streams simultaneously by taking advantage of inter-stream redundancies.)
As an example: with this system in the original Super Mario Bros your probability model would most of the time say that the user is most likely to just keep doing what they were doing (running right, or holding the jump button etc). Crucially, in the long run Mario will behave as if he jumped exactly when you pressed the jump button, but you might see a short visual glitch where Mario seems to be doing something wrong or weird, but it'll correct itself very quickly.
Similar to how mosh's predictions work.
If you add a bit of extra smarts locally, you can even add some cheap interpolation between possible futures to try to cover some that the server hasn't sent.
Mostly, what the local player and what the other players are doing will have differ between 'futures', but eg you'll mostly likely to be able to re-use the complex lighting rendering for the fully realistic leafy trees and blades of grass you see in the background between nearly all your possible 'futures'.
And while you're right that the computation on the client end could be made lower (depending on the constraints of the game), you're now going to need to care about the hardware of the client beyond "can render a video stream". My (regretfully) smart TV has trouble rendering its own menus, it's not lifting a finger to do a game.
And for best effects, of course, you'd also want to get support from the game itself. AI coding will make this part easier, because there's probably relatively common techniques you can apply to your games to get most of the benefit.
Just to be sure: I'm just outlining what could be done to make remote gaming viable, even with the latency constraints. I'm not saying that this will definitely happen, nor that it's even particularly likely.
Perhaps even low end local devices will just be fast enough to give people good enough games anyway. Nowadays even phones are capable of astonishing feats.