MMO Architecture: clients, sockets, threads and connection-oriented servers
prdeving.wordpress.com
prdeving.wordpress.com
Any time you add more servers to spread load, you're increasing latency because for each hop you're traversing the entire software network stack twice plus hops through hardware switches.
Nobody uses dedicate thread per client anymore (if they do, its a poor design).
As an internet backseat network performance person...
Have you considered one thread per core/NIC queue receiving packets, with RSS (receive side scaling)? If your bottleneck is network I/O, that should avoid some cross-cpu communications. Otoh, if you can't align client processing to the core its NIC queue is handled on, then my suggestion just adds contention on the processing queues; although maybe receive packet steering would help in that case. But, I'd also imagine game state processing is a bigger bottleneck than network I/O?
also worth mentioning, we wrote everything in c++. anything else is too slow.
openssl speed -evp CIPHER for example: (ECB only for reference, don't use)
type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes 16384 bytes
AES-256-GCM 655193.56k 1747223.44k 3399072.36k 4490100.61k 5033129.30k 5108596.96k
AES-256-OCB 631357.15k 2268916.74k 4794610.30k 6492985.36k 7174274.14k 7301540.60k
AES-256-ECB 997960.96k 3972424.35k 8096120.70k 8105542.89k 8179659.94k 8188882.64k
AES-128-GCM 747463.90k 1856932.41k 3762591.34k 4700335.25k 5224533.22k 5157661.06k
AES-128-OCB 702729.38k 2479151.86k 5919529.91k 8719316.46k 10545305.23k 10536095.61k
AES-128-ECB 1291715.10k 5180010.63k 11093258.14k 11815558.81k 11913592.61k 11947322.76k
OCB also won the CAESAR competition for "High-performance applications" portfolio. It is much older than the competition and is no longer patent-encumbered.LMAX disruptor?
Could you explain more about your lockless queues?
I recently wrote a lockfree ringbuffer inspired by LMAX Disruptor but it is only thread safe 1-thread to 1 thread. SCSP. It has latency between threads on a 1.1ghz Intel NUC of 80-200 nanoseconds.
I have ported Alexander Krizhanovsky's ringbuffer to C but I haven't benchmarked it.
https://www.linuxjournal.com/content/lock-free-multi-produce...
Second Life has 50,552 users connected right now. Typical Saturday. They're all in the same world. But they're not near each other. The biggest crowds are about 130 users. There are a lot of filters. Each client is connected to the server for the region it is in, plus a few nearby regions, so it can see past region boundaries. Within a region, the server sends updates to each user as objects move. Updates are frequent, up to 45Hz, for nearby objects in the viewing frustrum. Lower for distant objects, and much lower outside the viewing frustrum. That avoids the O(N^2) load problem.
The clients overload before the servers do, incidentally. The classic clients are single-thread and do too much CPU work per visible avatar.
He was the 2nd employee at Blizzard, did the netcode for their games and also for Guild Wars.
His blog is great https://www.codeofhonor.com/blog/tough-times-on-the-road-to-...
Anyhow, both talks are fascinating and surely worth a watch.
That lines up with Pat's assessment:
> Initially Collin Murray, a programmer on StarCraft, and I flew to Redwood City to help, while other developers at Blizzard “HQ” in Irvine California worked on network “providers” for battle.net, modem and LAN games as well as the user-interface screens (known as “glue screens” at Blizzard) that performed character creation, game joining, and other meta-game functions.
With a 2 tier setup, you could theoretically have a central server processing batches of events that are aggregated across multiple players/regions/etc. A single thread can service upwards of half a billion events per second, and if each of those events covers multiple potential players, then I'd argue we are in a pretty good position.
For me, having truly 1 consistent global universe is the only thing that would get me to consider an MMO in 2023+. The technical limitations forcing sharded worlds were excusable when WoW was released, but I don't have patience for those arguments anymore. Not if you want me to expressly burn my time and get paid for it on a recurring basis.
Can you explain/link to some?
I suspect MMO servers, particularly when server authoritative end up having to do a lot more processing per event and tick. Distributing world state updates is also an extremely thorny issue with lots of inter-dependencies.
Whatever is good for low-latency trading systems (i.e. where you are paying contractually for microsecond-level guarantees), is also maybe good for gaming.
The whole concept with these patterns is that there is only ever a single writer at any given moment. Single writer principle w/ batching is what gives you enormous uplift in throughput, even when you are necessarily constrained to fully serialized business semantics. I'd argue that an MMO does not need to be as strictly serialized as whatever CBOE, et. al. are doing.
Aeron on the face of it looks very similar to the architecture diagrams you'd get popping up designing a networking layer in a game already. I suspect the cross-pollination in both directions could be interesting.
Neither seem to particularly help with the intensive/interesting bit of an MMO which is running the actual simulation. For example big fleet fights in EVE as far as I remember have never been IO bound rather single-core performance bound. It's common for the bigger Corporations to contact CCP ahead of big fights so the node it's likely to happen on can be moved to a bigger machine. It's actually quite interesting because a lot of it is written in Stackless Python so not the most performant of languages and trapped in a single-threaded context so there is (nominally at least) quite a lot of headroom.
For big seamless worlds how you distribute the simulation is the tough bit. Hence companies like Improbable trying to solve that generally.
It could even be multiple games connected with sufficient break in context. Eve for space and another game for ground assault. But they actually interact live.
Then throw in next gen VR and you basically have ready player one.
EVE is by the way also working on an FPS, Vanguard. Which is supposed to interact with the game world of EVE.
I worked with Oddur and Ivar there quite a few years ago now so am looking forward to SEED!
I had some hopes for Improbable. Improbable has a big seamless world metaverse back end system. It involves a lot of remote procedure calls between multiple servers. The result seems to be excessive server costs. Five indy games, some quite good ("Worlds Adrift" was one) tried it, went live, and went broke. Improbable, a VC-funded company with about $400 million, has been thrashing around ever since. They tried setting up their own game studio, and produced "Scavengers". That was a game where a few thousand players all charge the same goal. More of a crowd tech demo than a game. After that flop, Improbable pivoted to making simulators for the British military, which apparently worked, the military not being too concerned about a few dollars per hour server cost during war games. That job done, they tried hooking up with Yuga Labs, the Bored Ape / Otherside people. Did another zerg rush demo for them. But cost per user per hour was so high they only ran that twice, for a few hours each time. Now they're pivoting again, to servicing baseball ("MLB", as the baseball industry calls itself) so that fans can watch games in VR, or something like that.
So what went wrong? The business problem was that ad-supported metaverses don't work. There is no role for "brands" in a highly immersive world. Quite a few companies have now figured this out the hard way. Ignoring that, though, what are the technical problems with scaling?
Having spent too much time inside Second Life client code, and written my own client, I can answer that. First, the user needs a "gamer PC" and serious network bandwidth to deal with a highly detailed dynamic world. With user created content, a key metaverse feature, there's far less instancing, and you need about 3x the GPU memory of a curated game. So you need roughly an NVidia 1060 and a few hundred Mb/s of network connection. The average Steam user has that, (see Steamcharts) but the average Facebook user does not. If you want to support WalMart $99 PCs and phones, there's a big problem.
"Cloud gaming", where the GPU is in a data center, lets anything that can play NetFlix play AAA title games. The problem is cost. Most of the cloud gaming hosting services gave up. Even Google gave up. NVidia GeForce Now remains, but they've raised their prices several times. Each user is using a dedicated PC-class system with a good GPU, so this isn't cheap. So trying to solve the user cost problem with cloud gaming doesn't work.
Server side is actually less of a problem. The Second Life server architecture isn't bad. Second Life was supposed to be the "next Internet" when it was designed over 20 years ago, and as a result, it's overdesigned compared to most MMOs. User-developed clients are not only encouraged, most users use a third party viewer, rather than the open source Linden Lab viewer.
The networking architecture to the client is not too unusual. The time-critical stuff is on UDP, and the bulk data transfers are on HTTP. The UDP system supports out of order delivery to eliminate head of line blocking. (Reliable delivery, in-order delivery, no head of line blocking - pick two.) The trouble with allowing out of order delivery is that higher level operations with state can get into trouble. Hence discussions like this.[1]
The most unusual thing is that each client talks directly to multiple region simulators in the same area. This is what produces the seamless world illusion. The simulators also talk to each other, but mostly about objects crossing the boundaries between regions. Inter-simulator traffic is thus manageable. It's more of a state locking problem than a bandwidth problem. The design predates the theory of eventually consistent systems and conflict-free replicated data types, leading to immersion-breaking out-of-sync problems, including teleport failures and getting stuck crossing the boundary between regions.
With this architecture, there's no major limit to the size of the world. The internal traffic within the data center per region simulator does not grow with the size of the the world. Nor does the per-client traffic. Traffic to the asset servers does grow, but that's cached (AWS + Akamai). Shared services (login, billing, etc.) have multiple servers with load balancers.
So that's the successful path to a big, seamless world.
There's much trouble with seemingly random sluggishness, but that turns out to be due to various specific problems, some of which are being fixed.
[1] https://community.secondlife.com/forums/topic/503010-obscure...
Have any of them tried to add online shopping via affiliate links?
I'm building the server in Go. I use goroutines a lot. I use a shrinking/growing ring buffer for requests and responses. Uses UDP and I just receive a request and pop it onto the queue. I keep the address of the client along with the packet request. The queue is read by another goroutine and gets processed and sent to a packet handler. Which does game state work, and queues a response, or passes it off to another handler. Then the go routine for the responses picks it up, reads the address and port of the client and sends it the packet.
All I have maybe 5-6 threads running in parallel and I already can log into the game and walk around with others. There's tons of work still needed. But it's nice to know my underlying design is sufficient to maybe handle the dozens of people playing for nostalgia.
I have client sources as I am working with the person that owns the rights.
Protocol is not documented, so reverse engineering, but having the clientside helps.
The server software was lost to time and accidents. But we have dumps of quest data, times, mobs, etc and all assets in original form. So as soon as I get the server in a working state. We have the full game.
We plan to open source/creative commons everything when we get something playable.
I've done 4 massive redactors as the code grows, to accommodate changes in understanding and design flaws.
We had reached out to him when we discovered his work back then. But he refused to join our effort unless we 100% open source everything then and there. Due to legal reasons at the time (related to the T4C owners and T4C using similar code to BMC), we had to stay closed source. He still refused to work with us even with a promise that we plan to open source eventually.
We shared with him all the packet data models and such but he was kind of unresponsive after our initial conversations.
He is also working 100% blackbox while we have sources for the client and assets/server data.
Hopefully all that makes sense. I haven't really paid attention to his project since then and it seems he stopped all work efforts on it 2 years ago.
But to answer your question, it's quite a bit far behind what we have. But there is so much work that goes into an MMO. We have solved the network stuff, we are just working now on world state as we have the ability for multiple people to log in and do stuff with dummy data. I have been writing tools and importers for the server data. Like zone files, creature, NPC, item, etc data. All the monster spawners are in the zone files for example and then needs to be cross-referenced with other data, and then assets to actually draw spawn mobs.
I also think there is a compelling design space where cheating doesn’t really matter or make sense where trusting the client is fine.
2023 and now we have virtual threads that replace blocking IO with non-blocking automatically! Haven't tested it yet though! But I am reviving the 2001 MMO now!
In between I made my own MMO backend that uses NIO, conclusion is event-based network and validation is the only way to scale action MMOs past 100 players.
Also you need to share memory atomically between cores, so Java or C.
Other than Erlang, Pony[1] might be an interesting choice. It allows sharing memory, even mutable memory, but it tracks the sharing in the type system. It's been a long time since I looked at it, and it definitely had a bunch of rough edges, but I really liked where the things were going. I hope it got even better since.
I use some C++ features and dabble in js when I make HTML but really since Applet and Flash has been removed the browser is just a bloated waste of time.
Make your apps fast by allowing them to share memory without latency!
I have also coded Perl, VisualBasic, php and C# and I'm pretty sure this is it.
What we need now are VMs that can take any bytecode/instructionset and translate it to all others. For true crossplatform development. But that requires us to dump dynamic allocation. Arrays of 64 byte structures FTW!
You're gonna love this: https://en.wikipedia.org/wiki/WebAssembly
Still waiting for a Windows release of this: https://github.com/bytecodealliance/wasm-micro-runtime
Don't hold you breath.
Also needs Risc-V and ARM 32/64.
The whole point of games is real-time action.
It does not need to be violent, but it needs physics that give you the feeling of being there.
Multiplayer > VR
I haven’t updated it in some time because busy from work but take a look:
PS: now that I actually RTFA, the same concept is already mentioned there, it's just called Frontend Server.
Basically, learn how an OS works first before trying to reimplement it badly.
Whether that is actually true probably depends a lot on the specific operating system and async/await runtime (I think there's a lot of mysticism involved seeing async/await as some sort of silver bullet instead of relying on cold hard performance numbers alone).
(Async/await was invented in interpreted languages to circumvent their global interpreter locks, which is a whole different problem not related to context switching at all.)
Event loops were old hat in GUIs, and web servers like Nginx adopted them so they could service thousands of requests concurrently without allocating stack space for thousands of threads.
And the context is not the same. Switching threads requires you to go through the kernel scheduler, so it has to change page tables and stuff in the CPU, right?
It lets you write simpler code too, because some events can be handled in the loop without involving mutexes and thread safety.
If async/await is, as you claim, a pointless hack for interpreted languages, why did Nginx, written in C, get so much traction with essentially the same architecture?
There is a lot of great info about operating systems - specifically ring0 vs ring1 for the context switch overhead.
Meanwhile, only Python and Ruby had GILs. Async/await became popular through JavaScript and C#. Now very popular in Rust as well. None of these runtimes have a GIL.
Async/await or Fibers (C++, Go, Java 21) are entirely about context switching overhead.
To be fair, while typical JS implementations do not have a GIL, it's usually unnecessary due to the runtime being single-threaded.
In case anyone else is curious when the above languages grew this feature:
(F# piloted this in the 1.x days, back in 2009 or earlier.)
It seems like C# developed this feature first in 5.0 in 2012, followed by JS (Dec 2016 in Chrome, widely supported across other major browsers within a few months). Python 3.7 came out in 2018, Ruby 3.0 came out in 2020.
Async/await landed in C++20 via coroutines, and Java 21 landed last month. Rust grew this feature in 2018, which landed in the stable channel in 1.39 (Nov 2019). Golang has had goroutines / channels since day 1.
Yeah, "we avoid a GIL because we don't offer threads at all" is, like pre-1.9 Ruby's, "we don't have a GIL because are threads are 1:N green threads" both technically "no GIL" but also worse for thread-based parallelism, then "we have native threads, but there is a GIL, which native code can release".
> Ruby 3.0 came out in 2020.
Ruby 3.0 doesn't have async/await, though. There's a third-party colorless Async gem, and built in constructs (that require a scheduler implementation) for lower-friction asynchronous programming (async as an approach has been popular for Ruby for quite a while), but not with async/await syntax in either case.
(I haven't written anything in Ruby in roughly a decade, so I'm happy to defer to anyone who's actually followed the language the whole time)
That's fair. I had skimmed and found some reference to the async gem relying on Fiber::SchedulerInterface which was introduced in 3.0, although it apparently existed back in 2017.
There's similarly the async-await gem by the same org (also first released in 2017), which at least partially supports that syntax. But in either case, it's not a native language feature.
P.S. C++ async is entirely about running on embedded hardware, not about "overhead".
P.P.S. All garbage-collected languages have a GIL of some sort. This is the real reason GC languages aren't used in systems programming, not GC pauses.
Stackless (i.e. async/await as opposed to green threads or stackfull coroutines) context switches have the additional advantage that they reuse most of the stack and only suspend a single stack frame. This means that the rest of the call stack (shared between contextes) can stay hot in cache.
Whether any of these costs matter depend a lot on the application, the amount of context switches and the amount of work done between switches.
Say a user does two actions in quick succession - action A that is something that inherently takes a bit longer to process on the game server (ie a complicated transaction of some kind) and then right after does action B, one that requires little to no processing.
How is it best handled to make sure that the order of actions by the player is maintained with processing? With this scenario, it is possible action B will be processed before action A, leading to the game server considering them in the wrong order. I'm assuming the solution is to have some additional logic to make sure you are only processing one action per player at any given time, but I was wondering if anyone knows any better ways?
Critical actions should also have a sequence counter to make sure they are enqueued in order and to detect whether actions got lost somewhere (especially when using an unordered networking protocol like UDP) - message ordering usually already happens on a lower level than the game code though in the code layer that sits on top of UDP sockets, but a higher level per-player action counter can't hurt either I guess, if only for debugging)
Latency-sensitive games also might want to speculate ahead by running the local player's task queue on the client, and roll back when the authorative state coming in from the server disagrees with the speculated client state.
The server also doesn't need super precise time slices. Some of that can be hidden on the client with animations, cooldown timers, and other visual tricks. An MMO will have different timing expectations than say a twitch shooter.
AFAIK in the rest of the world, Photon Server is the most popular off-the-shelf solution (not necessarily for true MMOs though). Maybe that's good enough for most games to not wanting to write your own server backend (https://www.photonengine.com/)