Riot Games: Artificial Latency for Remote Competitors
lolesports.com
lolesports.com
Anyways, by "scenarios where the actual ping was significantly lower than the target latency" do they mean like if the ping from Busan-server is 10 but they want it to feel like 35? I feel like that's the one scenario you'd care about, making the Busan players have 35 ping was the whole point.
Edit - also the other scenario, where the actual ping is significantly higher than the target latency...well in that case your latency service would be some sort of magical time travel box! Adding latency is literally the only thing it can do right? It's not possible to lower it. Maybe the it depends what they mean by "significantly" but idk the target latency was 35ms so like at most they could be 35ms lower than that.
I guess they didn't catch it because the bug was in the latency service itself which gave the wrong value to both the in game display and to the old network monitoring system. Idk in hindsight it's easy to say if you're going to introduce this new ping equalizing service it might break how you test ping especially because it hasn't been tested in soloq. But meh a lot easier to see what happened in hindsight after they write it all up.
(I have no idea what LoL in particular supports or implements, but that's the general idea.)
By contrast, the server rolling back the game state is indistinguishable from the server always being 15ms behind and the clients running the simulation ahead of time (which they do anyway for fluency [1]). The "jumpiness" is called rubberbanding [2].
[1] https://www.gabrielgambetta.com/client-side-prediction-serve...
[2] https://www.urbandictionary.com/define.php?term=rubberbandin...
Either way, there is nothing magic. At the end of the day, the soonest you can see an action is still physically unchanged. Especially since the prediction model that makes the most sense is just assuming no action.
That said, I dunno enough about MOBAs to try and postulate on how much it matters, and I have no idea if LoL is using lockstep or some other approach.
> The second issue with lockstep is that if one player is lagging, all players are lagging
Where are you getting this?
My game Nebulous uses lockstep, and it's a MMOish mobile multiplayer game. Plenty of clients lag, but one client lagging doesn't cause any other to lag. Furthermore, clients do not have full game state, as an optimization and data transfer reduction first, anti-cheat second.
Path of Exile actually lets you choose lockstep or client-predictive. "Serious" and "competitive" ladder players choose lockstep because, naturally, you want your client's representation of the game state to be as close to the server's representation of the game state as possible, UX be damned.
Prediction could mitigate that (and possibly does, even in sc2), but if one client is suspended and the other is running fine, either every thing that the running client does is moot (and will be reverted) or it gets suspended, regardless of lockstep or prediction.
My speculation: this almost exclusively to protect the private IP of one client being exposed to another client. There is no server processing happening beyond connection verification to help mitigate desync attacks. Deliberate desyncs are still possible though, since I've been the victim of it.
Team Fortress 2 is particularly bad where high latency players can teleport around or you can shoot and hit a laggy player but it doesn’t register by the server due to a rollback. Some players purposely exploit this function by artificially increasing their ping values to absurd levels (700ms or more) which makes them very hard to hit. Many community TF2 servers enforce maximum ping limits for this reason.
Most notably high pings are used to exploit when playing as the Spy class due to the fundamental purpose of the class is to get behind you to backstab you. If you have 700 ping can abuse the fact that you you can get behind players due to the game being out of sync and if you can backstab a player the server will roll back and kill the non-laggy player even though you were no where near behind them from their non-laggy perspective.
I make this distinction because I actually think backwards reconciliation is table stakes for any kind of precision game and is usually pretty easy to implement. It can also help solve teleporting: you run backwards reconciliation on movement commands also (not just shooting or other actions), and combine with a little player position extrapolation.
Nothing fixes a 700ms, actually I would guess the ceiling on competitive play is somewhere around 200ms with incredible netcode. But 200ms gets you pretty far: across the US or the Atlantic Ocean for sure (packet loss over those distances is another story though).
also players only teleport around on your screen in tf2 if you set your cl_interp very low
I think the key word is "significantly". The post states that the tool was originally developed for matches where both teams would be remotely connected, so neither team would have a significantly different latency (you can usually enginneer this by the first approach they suggested, of picking a server in between two locations).
-listen to the players
-investigate past what your own analytics are saying
-find and fix the bug under time pressure
-and disclose the (potentially embarrassing!) initial mistake
I think the article itself could have been worded better though. It didnt feel very clear. As a reader, I want to know what went wrong, how it was handled, and feel confident that you are on top of the issue. All the stuff about "why we chose our network topography" reads as details at best, justification/excuses at worst.
They sacrificed the competitive integrity of their sport because they didn't want to exclude RNG from the tournament, because they wanted the audience of the Chinese market.
This article goes into more detail, giving a larger context for why this situation shouldn't have never happened in the first place.
https://www.invenglobal.com/articles/17207/montecristo-its-e...
Or maybe because the tournament would've been shit without them? Who wants to watch an international tournament with only 1 of the clearly top 2 regions represented? WHole thing would've been a forgone conclusion at that point.
As the article states, increasing the ping doesn't make it more fair for the other teams because they actually traveled to the event. They're physically there. It only makes it more fair for RNG who can't attend because of Covid restrictions.
To paraphrase, it "punishes the teams who are there for a team that isn't."
From a business perspective of course, I agree with you. It's definitely better to sacrifice the level of play for viewership (and it does sacrifice the level of play because artificially increasing the ping at a LAN event makes the game worse).
But if you're running a competitive eSport and you want a fair sport, it's a terrible decision.
Can't agree. Those teams had an advantage when playing RNG. Taking away that advantage does actually make the event more fair for them, even though they are being disadvantaged.
There's a reason this Frankenstein ping is unheard of in bigger esports.
Can you provide any clear evidence that indicates this is an advantage?
I don't think there's anything bigger than League, at least in terms of viewership
Unless "ability to travel during a pandemic" is considered a skill for this game, it should not be factored into the competition. So, if teams who can't travel to be game are going to be included, then the host must find a way to provide a fair game. This means that remote and local players should have the same latency.
Could you clarify what this means? I’m not entirely sure I understand what you mean. I think I just need a little more context.
> if teams who can't travel to be game are going to be included, then the host must find a way to provide a fair game.
I agree with the logic but disagree with the premise: teams who can’t travel to play the game shouldn’t be included in the tournament in the first place, because it places an undue burden on the other teams and lowers the quality of play.
> Unless "ability to travel during a pandemic" is considered a skill for this game
I don’t think it’s a skill; it’s a basic obligation individuals and teams must fulfill.
> Could you clarify what this means? I’m not entirely sure I understand what you mean. I think I just need a little more context.
You have described fairness in terms of punishing or rewarding players for certain non-game behavior. But, this is not the right way to think of fairness, in the context of a competitive game. A competitive game has a set of rules, which the players compete under. In a fair competitive game, those rules don't favor a particular set of players. Regardless of whether they deserve it or not, players who couldn't show up in person would be playing under a handicap if there was no lag correction.
>> if teams who can't travel to be game are going to be included, then the host must find a way to provide a fair game.
> I agree with the logic but disagree with the premise: teams who can’t travel to play the game shouldn’t be included in the tournament in the first place, because it places an undue burden on the other teams and lowers the quality of play.
>> Unless "ability to travel during a pandemic" is considered a skill for this game
> I don’t think it’s a skill; it’s a basic obligation individuals and teams must fulfill.
Sure, it is an understandable position to hold, that players who can't show up shouldn't be able to play. Riot clearly disagrees with you(for business reasons probably, mostly). I sort of disagree with you, mostly because I think that while we're in this weird liminal stage of the pandemic, I think we should try to accommodate people who want to engage from home. But this is a soft disagreement, I think there are good arguments either way.
However, once the decision is made about which players will be included in the tournament, if the tournament is to be taken seriously it must present a level playing field to the players who've been let in.
I do agree with you the host should make a decent effort of accommodating players who can't travel though, but I thought this was pushing it a bit too far for my taste.
What's funny about all this is that technically because the Asian Games aren't going to be held anymore and because it was one of the mitigating factors for RNG not travelling, technically they could physically travel to Korea now.
Funnily enough, all this hoohaa was for nothing as the Asian Games got postponed anyway.
Instead of processing orders instantly, wait between 200ms and 500ms and then process all orders that came in that window in random order. Then being 5ms closer to the server wouldn't matter.
And honestly the wire thing probably isn't real. Light moves 30cm in a nanosecond. The wire lengths could be 3 meters different and only make a 10ns difference. Not sure that would matter all that much.
If an HFT manages to exploit a correctly implemented entropy-pool random number generator using AES-256 to extend the stream as needed, they're welcome to pick up a few more bucks as far as I'm concerned.
Yes, problems have existed before but I'm sure in this case we can assume high-assurance and careful programming, not some fresh grad being assigned the problem and thinking linear congruential generators sound really cool.
Wait, come back, I'm serious! Finding an arbitrary hash of adjustable complexity is a scalable solution to batching transactions across multiple servers with a consistent throughput.
it's part of the exchange's license
and they have several videos of the rather large wire
there's even a tom scott video on it: https://www.youtube.com/watch?v=d8BcCLLX4N4
• Pre-publish, for each time batch, a public key. You could publish lists of these well in advance.
• Let everyone submit, alongside each order, a number arbitrarily selected by them. It does not matter how they select the number, but it would be simpler if everyone chose distinct numbers.
• When order processing is done, do it by the order of closeness of the submitted number to the private key. Publish the private key after order submission is closed.
• Everybody can now verify that the order of processing is indeed by the order of the previously secret private key, and everybody can verify that the private key corresponds to the previously published private key.
Edit: see https://en.wikipedia.org/wiki/Commitment_scheme#Coin_flippin...
Say I know exactly when the window ticks over. I still want to put my order through immediately anyway since it's to my advantage to get it in the earliest possible window. There's no advantage in waiting for the next window to start.
So I put in my order, and my closer competitor puts in their order at the same time, and they get in the current window, while I have to wait for the next one. And now my disadvantage is even bigger because my order is delayed on top of my base latency by the remaining time until the next window ends.
-----
Even if we remove the window idea entirely and just add a random 100-500ms delay to every command immediately, being closer still has an advantage since 10 + Rand(100,500) is still lower on average than 100 + Rand(100,500).
(though the matching algorithm varies)
And this is a little different, but Riot found that playing on ethernet instead of wifi made you about 1% more likely to win games. https://web.archive.org/web/20160814131032/http://na.leagueo...
In what situation can this happen?
> we realized that there was a calculation error that only manifested in scenarios where the actual ping was significantly lower than the target latency. In this situation the actual latency would be considerably higher than what is displayed on the overlay on the player's screens.
They added too much latency because they forgot to account for the latency they added in the client. You can see from their architecture diagram (Figure 4) that the latency measurement didn't include the client delay.
https://images.contentstack.io/v3/assets/bltad9188aa9a70543a...
The blog post states: "The existing network monitoring system measured the latency at the networking layer as shown by the green arrow."
But the blog post states: "a calculation error that only manifested in scenarios where the actual ping was significantly lower than the target latency".
Because realized ping is over by as much as original ping was under, so if original wasn't significantly under...
There are videos floating around that show drawing on a tablet surface with various input latencies (perhaps someone has a link, I can’t find them at the moment). 35ms latency is very noticeable to anyone never mind professional competitors.
Because of the tricks game designers must pull off in lieu of proper distributed consensus which has hard requirements bound by the laws of physics, it is completely likely there are lots of bugs in the system. I think Riot did the best anyone could reasonably have expected of them, and the write up is particularly informative and helpful.
In fact neglecting the impact of distributed consensus is one of the biggest challenges to mitigating it.
They built a complex system to tweak tens of ms (at most) of ping to equalize. And they had a bug and it disadvantaged on team. Thus, they have to replay due to unfairness introduced by Riot itself.
They could've gone Option 1 - all teams at their natural ping. While it wouldn't be as perfect in theory, it wouldn't have resulted in replaying matches. Plenty of other esports and fighting games play at natural ping and have real results. There is unfairness due to location, but only scrubs blame 5-15ms ping advantage for a loss in a game as strategic as LoL.
People play Melee with ping differences bigger than that and it works fine lol. Waaay more technical game shows that there's tolerance.
Not to mention that the server option could've had some system to distribute across advantageous servers over the course of a series.
Of course 35 ping is not THAT bad and I agree, while unfair, it wouldnt be the end of the world, it's not even Worlds, just MSI, an invitational event.
imo it's very serious tournament
Sure, people play with ping disparity in tons of competitive games, but that's because there's not really an alternative for online play. It's "play with a ping disparity" or "don't play at all".
Once there is more widespread reliable alternative, which Riot seems to be working toward in their own game tournaments, maybe we'll see if this tolerance persists in competitive tournaments.
I dunno; I'd imagine people would be mad if there was a 15ms ping differential in Melee. I think the online/netplay tournament results are already looked at different than offline, console results.
Moreover, I'm pretty sure that you can't have a differential ping in a P2P game with rollback, so the competitive issues are different between the two scenes. I understand Riot wanting to equalize them. In Melee specifically, I think that the players have to decide on a common frame delay to play, right? If so, then the Melee community has come to the same solution for competitive netplay as Riot has.
Low ping is part of players' muscle memory in how they react to visual signals. Changing that slightly creates a big difference, especially in the early and mid game where the outplays are less strategic and can often be more mechanical.
Even in later game teamfighting, the familiarity with latency between key/mouse input and actual in-game movement is mechanically intense (ex. auto-spacing in a close teamfight)
Modern fighting games, as well as older ones that have been modded to included this (such as Melee), have a netcode model called "rollback" that can negate network-induced latency on local inputs. The trade-off is that what you see is usually inaccurate until you receive inputs from the remote player's machine.
Despite fighters often requiring fast reaction speeds, this downside is not a big deal granted that the ping is low. The first few frames of many actions are generally quite subtle, enough that they are quite hard to distinguish from each other; it's usually specific keyframes/poses that people will actually be reacting to.
Melee is probably hurt the most though, as movement is extremely fast and very important, and the difference between wavedashing left or right is not masked by subtle animations.
Another interesting scenario is projectile-based weapons. For reasonable Internet latencies and reasonable in-game distances, the projectile can travel faster than the packet announcing that it has been fired. The game state that gets sent from the shooter over the network is "I fired my weapon, I killed them". At the time you die, your client has none of that information, and since the shooter is favored, you just die out of nowhere. (This is less annoying, because I don't think you could move out of the way fast enough for it to matter. "My opponent's WiFi is bad, so I die half a second later for no reason" still feels pretty bad though. In the intervening milliseconds you had grand plans to win the game!)
All of this, to me, is a great lesson in eventual consistency in distributed systems. People break their keyboards over this stuff. Thinking it will be fine for your more-critical-than-a-casual-video-game system means this stuff will be happening to you multiple times a day, and you can't break your keyboard because you can only blame yourself for brining that system into the world. Tread carefully! Tread very carefully. Shit gets weird when your system has a split brain.
Second surprising thing is artifical latency added on the client. Anything on the client I would avoid, if possible.
Third, this all looks like magically connecting two ends. But it is packets flowing from one to the other, and the other to the one. Latency could be added to each independently, even on one end, by artificially buffering to-be-send and/or received packets.
What are the implications for game-play with an asymetric latency, e.g. theoretical 0 delay to receive vs. long delay to transmit, or vice versa?
If you only add latency on the transmission, then 2 actions taken at the same time by players with different amounts of physical latency will be processed at different times. If we assume the game clocks on each client and the server are in sync, then this creates a competitive advantage for the player with lower physical latency. For instance, if we have a timing battle (ie: Zhonya's hourglass), where both players know that at time 'T' they must each take an action, and the first to do so comes out on top, then the player with higher physical latency is at a disadvantage.
For all on the server, you could add latency right after receiving data from the client, but before game code processes it, and before sending data to the client.
That said, there likely were technical considerations involved in their decision to do a mix of both.
I now wonder too what those technical considerations are, since having the client add latency requires trusting the client to add the delay and not cheat.
With two cities, a dedicated common connection with measurements for latency in both directions, and real-time tit-for-tat latency adjustment would be appropriate (if remote city A gets latency X, ensure local city B approximately gets latency X on next game tick)
Without that:
0. latency fluctuations could penalise the remote team even though average latency was kept consistent over longer periods,
1. local players could advantage themselves by injecting delays or jitter into remote players during critical periods of play (DoS like network traffic between local and remote cities),
2a. local players could get server updates 10s of milliseconds before remote players (if all latency adjustment added on transmit from player to server), or
2b. local players could get their events updated to the server 10s of milliseconds before remote players (if all latency adjustment added on transmit from server to player).
https://images.contentstack.io/v3/assets/bltad9188aa9a70543a...
The blog post states: "The existing network monitoring system measured the latency at the networking layer as shown by the green arrow."
"The reason we did not find it sooner is that the cause of the issue was a code bug that miscalculated latency, which meant that the values in our logs were also wrong."
And from later in the post:
"Our logs were not displaying the issue because the calculation was wrong. It explained why the latency was worse in the venue than on the internet servers"
And they also don't show that their dynamic adjustments are more fair than a static adjustment. I think it's likely that their system penalizes local players for too long after a short remote latency spike.
Now consider on top of the fact that there are more buffers than just the `tc` buffers involved. This leads to non-linear, non-deterministic fluctuations in latency performance when bursting occurs.
tc is purpose-built to shape throughput, not for tuning precision latency.
(for some context: traffic shaping is a big part of what I work on in my day-to-day helping build and run a wireless ISP)
they didn't let anyone proficient at LoL even try using it?
in games you really feel ping differences
if you were playing e.g for a year on 5ping and then jumped on 30, then you'll feel that in games like LoL and probably CSGO? idk about the second one.
Some problems cannot be solved with clever tricks.
# hping -S -p 1234 somehost.com
HPING somehost.com (ens3 1.2.3.4): S set, 40 headers + 0 data bytes
len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=0 win=26883 rtt=0.9 ms
len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=1 win=26883 rtt=0.8 ms
len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=2 win=26883 rtt=0.7 ms
# tc qdisc add dev ens3 root netem delay 25ms
# hping -S -p 1234 somehost.com
HPING somehost.com (ens3 1.2.3.4): S set, 40 headers + 0 data bytes
len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=0 win=26883 rtt=25.9 ms
len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=1 win=26883 rtt=25.8 ms
len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=2 win=26883 rtt=25.7 msand they did it basically after the network stack, so if a paket took 40ms they added no latency if it took 20ms they added 15ms and so on.
so EACH paket was evaluated and of course this is not an easy feat, because lol is serverside authoritive so the latency must be correct inside the server and the client for each packet.
Simulating this stuff is pretty simple. Model each incoming stream as an MM1 queue and play around with your PID algorithm until it gives you the results you need.