How to beat lag when developing a multiplayer RTS game
construct.net
construct.net
What most games that use this technique do is "lockstep" - they send out a chunk of commands called a "frame" several times a second, and step the simulation forward to step n only when they have frame n from every player.
So, in your scenario, if player A cancels before troops are visible to player B, then player B's computer won't ever show the troops, because simulation B will have paused during the 500ms of network lag, since it didn't have the frame from player A.
1000 units in intense combat can be synced with about 50 KiB/s - so why go to the trouble of deterministic gameplay, dealing with de-sync bugs, and making it diffiuclt to late join?
edit: The drm-like whack-a-mole is precisely the result of the architecture described in the blog, the "old school" way makes the most destructive and obvious cheating simply impossible.
The biggest reason is to remove the need for frame synchronization. In fighting games, it's important to get frame-perfect inputs. Even in an RTS, frame-perfect inputs matter, e.g. when kiting enemy units. You also get features like replays for free. Late joining really isn't an issue because the server can reconcile the state, deliver it to a late-joiner, and then continue on from there. You don't need to deliver the entire history of events to a late joiner.
The starting point for any such discussion is (as in this article) where all updates are done on the server. You then have to deal with packet loss and latency spikes but the real problem is subjective: it just doesn't feel good.
Imagine you're playing Fortnite and when you pressed your trigger you had to wait 50ms for it to acknowledge that your gun fired. That bullet has travel time so there may be another update if you hit someone.
So instead the client gives you immediate feedback and proceeds as if that shot actually happened. This may well include calculating if you hit the target. That target's position may be interpolated too, not the location you last got an update for. This feels way more responsive.
What if your target actually stopped moving after you shot. You get into an ordering canondrum. Now imagine if shooting that shot had recoil (which could affect both aim and position) and getting hit moved the target (eg getting hit by an explosive of some kind).
Doing all of this when latency can easily be >100ms and having it feel good is incredibly difficult.
And fake packet loss and latency spikes introduced by cheaters.
For example, when a cheater gets killed they could send a packet from the "past" saying that they killed the other person first.
To hide the fact that you're always lag spiking 100ms when you see an enemy, you'd have to lag spike randomly all the time, quite a lot, and then you're asking for a kick for low ping.
This is much more complex in that the client and server share commands rather than state and they both have to be exactly deterministic in processing commands to get the new states. There's also computational overhead when rolling back then replaying forward which can be a lot if the game state is large. Searching "netcode" in reddit.com/r/gamedev could show other solutions like the one you've come up with. Netcode has been mentioned on HN in the past as well, though I don't remember--only found by searching.
Nice work BTW, on both the game progress and smooth playable sync.
[0] https://en.wikipedia.org/wiki/Netcode
[1] https://nichegamer.com/stormgate-first-rts-rollback-netcode
The other approach(and the one most RTS games use) is called lockstep where you have a fully deterministic game simulation and clients send each other's input's to each other and all of them run in "lockstep" with each other. Generally this adds latency but has the benefit of scaling well to a large number of entities. It also requires your game simulation be deterministic which is extra fun when you have crossplay with different CPU architectures/OSes. If you've ever hit a "sync/divergence" error in those types of games that's usually a gamestate checksum that failed and bailed(some games were smart enough to re-sync the gamestate which is similar to host migration in P2P titles).
Rollback is a bit of a hybrid of both where you run closer to dead-reckoning(I.E. clients simulate all entities) but as inputs from other clients come in the simulation is re-wound to that point(I.E. "rolled back") and then re-run with the new inputs + some smoothing. It works really well for reaction based games which is why you see it widely used in fighting games and an early version of that was used in many FPSes for critical things like hit-tests(with fun byproducts of getting "warped" around a corner in high latency situations if getting hit changed things like movespeed).
Networking in games is an absolute blast, you have the fun technical problems outlined above but there's also an aspect of psychology/game design where a lot of what you're doing is "masking" latency in a way that's not visible to the user. A simple example here is playing hit sounds/effects locally but resolving them server-side. The 100-200ms latency isn't noticable if there are things happening client-side. From the game design side you can have games where it's about predicting where an action will happen 200ms-1000ms+ in the future which is much more latency tolerant. It's how games like Subspace[1] back in '99 was able to do 100s of players with 250-500ms latency and still have a high fidelity game since it was all about prediction instead of twitch reaction.
Well, that's not entirely true for a number of reasons. For example, if you look at a game like Street Fighter II, there are around 4 unique frames of animation for throwing a fireball. So any rollback that occurs there will hardly be noticeable. In newer games though animation frames are interpolated in different ways so you might have 10x more frames of "unique" animation. This makes rollbacks much more noticeable and jarring, especially in games with 3d models. Then you have the much trickier issue of audio rollbacks. This is even noticeable in games like SFIII:3rd Strike, where you might have the "KO!" audio effect play right as a rollback occurs, leading one player to think the round has ended before they realise a rollback occurred. For some extremely bad examples, search for some videos on Street Fighter x Tekken's, terrible RBN implementation. One of the ways developers get around these challenges is to add artificial delay to things like sound effects, adding more startup frames to moves, and adding massive amounts of hit-stun AKA "impact freeze" making games feel less snappy. For some games these things are not such an issue because traditionally these games rely on "slower" inputs and higher impact freeze because of the game's system mechanics, but in other titles the games just feel worse to play "offline" than their predecessors. For the most part, these are just "veteran player" issues. Newer players are not likely to notice them.
In any case, the bottom line is that RBN is not a panacea to online lag. Your game needs to be designed from the ground up to cater for it, unless you're willing to make certain compromises in online play.
Presented solutions include jitterbuffers to account for packets that arrive late or out of order; synchronized clocks; stream timelines... all those also exist in any typical WebRTC application, plus others such as packet retransmission requests, or even detecting available bandwidth on each client (to adjust video bitrate accordingly).
Of course, RTP is originally meant for video transmisison, but I wonder how much of the typical RTP stack (like that from GStreamer or FFmpeg libraries) might be reusable and useful as a baseline for implementing the online part of a videogame.
Regardless, in my above comment I was thinking of a lower level. Not WebRTC per-se, but the plain RTP protocol.
Every RTP (UDP) packet has NTP timestamps included in the header. RTP endpoints (both client and server) have a well defined set of calculations user to synchronize packets based on their timestamps, and also infer network latency, jitter, even clock skew, that kind of things.
RTP packet headers also have a sequence number. This is used to detect out of order, or missing packets. Then the packets can be reordered, or (if it makes sense) a retransmission request be sent back to the sender.
I.e. a lot of the problems that were being described in the original article, seem to be already handled by RTP.
So I was thinking of how realistic it would be to use multiplayer game data, instead of video frames as the RTP payload, to leverage all these network-related mechanisms of the protocol.
Could have saved me months of thought 10 years ago if one had existed. ;)
It's always validating to see someone else arrive at similar conclusions when facing a challenging technical problem. To the author, I encourage you to keep pushing (if you want to), it's possible to replace the fixed synthetic delay with a control loop. :D This makes both the PDV _and_ the system latency approach their minima over time.
This solution is in roughly the same space as one I wrote about a long time ago. https://www.forrestthewoods.com/blog/tech_of_planetary_annih...
I guess some day we can have quantum entangled FPS games so it's all instant and there is no lag? I only sorta half joke.
It's especially tough for big-world systems where not all the users in one area are anywhere near close physically. Sharded systems can put all the users on servers near their location, but big-world systems don't have that option.
[1] https://community.secondlife.com/forums/topic/451190-merrily...
The ping message could have the client’s current time, and the pong message would have this client time and the server time.
As a consumer, my first reaction is fuck activision this is bullshit, bunch of incompetent dimwits.
The flip side however, as someone who knows better, sympathizes and says, wow it’s amazing this works as well as it does.
Awesome article, especially for anyone who plays games and does technical work.