Thundernetes makes it easy to run your game servers on Kubernetes
github.com
github.com
Also is Hesky still around these days?
Seems well and good. I'm not sure if you lose any real DDOS protection by going through a load balancer.
I do have to wonder if its worth running game servers in k8s or just use autoscaling instances without the pod concept. You're really fighting against what k8s does well, for what I assume is no benefit. Thoughts?
Alas, it still does; at least for the Java client.
When it comes to gameserver performance what tends to be the biggest issue is predictable performance. So, kubernetes is not inherently antithetical to running game servers as long as you don’t share the cpu core(s) with anything.
In fact there is a project that (while I was not a part of) was spearheaded by ubisoft in collaboration with Google for running game servers on Kubernetes (called Agones)[1].
Ubisoft also used “Thunderhead” (which this project seems to be a variant of) on Azure with rainbow 6 siege. So it’s definitely not considered a show-stopping problem.
When it comes to containers/VMs the biggest hit in terms of performance is in this order: disk/network/memory access/CPU ;; when talking about game servers the most important things go in the reverse of that order, which means that the biggest performance losses are not as important. Network latency might seem like a big deal but you’re limited to frame times in the best case and geographical differentiations will eat much more than the combined weight of frames and Kubernetes networking.
Agones itself bypasses a lot of the Kubernetes networking latency additions, though, not all.
It’s quite old now but interesting for anyone whose experience is purely running on their own hardware.
In this case they seem to be NAT'ing packets from a host port on the correct Kubernetes node (the one running the container) to a port in the container, which can be done fast enough with iptables (or similar mechanism used by Kubernetes).
iptables is a good example -- it can scale rather poorly! Packets are run across the chains at length until a matching rule is found.
For most configurations this isn't a problem - the rules are filtered against quickly.
If density reaches the point to where you have thousands of forwards, it'll slow down a lot!
You'll want to look into optimizations (eg: ipsets), offloading to hardware, or simply going to host networking
Mostly reliant on (hardware offloaded) network latency, memory, and storage. Consistency being the thing of most importance. A spike is felt - interpolation can only go so far
CPU virtualization extensions make the latency cost of VMs hardly noticeable, containers are technically better but it's a split hair.
On an unloaded host and a single VM you can achieve essentially identical performance as bare metal - the CPU extensions, pinning cores, and huge pages are key
Some performance metrics even improve! Disk operations tend to do better from an added layer of memory involvement
Edit: there are options for multi-tenant fairness whichever way you go -- Bare metal, VMs, or containers
Performance isn't much of a worry, more so the 'handles'
There are exceptions. For years you could only run Satisfactory servers by the server acting like a client - rendering and all
Or perhaps more realistic you could provision a tier of very low latency VMs.
Not being snarky, just realistic - I feel like cloud providers sell "kubernetes" as the only way to properly develop apps, because you can be cloud-agnostic, without lock-in, you have full control over your infrastructure.
Without realising, that for most projects, running a kubernetes stack requires a person just for tuning and maintainance. Kubernetes is great - when you hit massive problems of scalability, but most people aren't going to be facing the problems kubernetes solves.
I see no reason to not use it, even if your deployment cluster lives on a single vm. The API is worth it.
- Very elastic demand curve
- Generally ephemeral with any necessary persistent state stored off-host
- Benefit from geographic distribution of cloud regions because they're generally quite latency sensitive
There are some things that can run on cloud, but the heavy hitting game servers should really be bare metal for the baseline.
I gave a talk about this in Stockholm once, unfortunately it wasn’t recorded but I can share the slides if you want.
It certainly depends a lot on how sophisticated your autoscaling is and how closely you're able to follow demand to limit waste, how well you can manage per-host utilization, whether you are CPU or memory bound, and lots of other factors. But even at truly massive scale the cloud hosting option can be very competitive without nearly as much management overhead.
I wonder why our experiences are so different.
For context I was running 2,500+ instances at peak with 40vCPU and 256GiB of memory: the most expensive of those of course being in the regions with low density of players like South America and Australia.
We also had a predictive auto-scaler for our cloud components, I would estimate that our waste was 15% at any time, but if we ran bare metal only it would have been cheaper (except for the operational cost and the fact we needed lengthy commitments for hardware)
We're also seeing some pretty compelling results from ARM based instances: https://www.wired.com/sponsored/story/changing-the-game/
Our waste was the inverse of yours, less waste on the upswing and more waste on the downswing, due to the long lived nature of our instances.
One instance takes roughly 1,000 players, but if you as a player decide to just keep playing, then we won't kick you out of the game for some hours. Some percentage of players will just continue playing as long as possible without matchmaking or map transfers.
Managing bare metal at scale is challenging.
The cost difference would have been enough to hire 2.5x my team as additional resources, but that differs with scale (and so do the discounts)… So YMMV.
The larger concern here is the egress bandwidth costs from large clouds, which if you don't have Fortnite scale and favorable deals, will cost you ~9c per-GB. If you use bare metal, you usually don't have to pay that, the cost is baked in to the instance.
Cons of bare metal being difficult provisioning, generally having to purchase instances ahead of time whether you need them or not, and having them sit around and not be used when you are off peak time.
Best to write your own service provisioning layer that can support both bare metal and cloud instances, that way you fully understand it, can make changes to it, and it becomes a core competency for your team.