Testing a 1,000 player Minecraft server with Folia
cubxity.dev
cubxity.dev
> During the period when 1,000 players were online, we reached a maximum of ~7.9GB/s heap allocation and our GC was hovering around 2-3GB/s when averaged over a minute.
As someone who dealt with soft-realtime telephony stuff, this makes me want to scream in horror. It seems like the platform really hurts the performance here. In an average second with that many players, most (all?) of them will not do any action apart from changing their position. (Just moving around would ideally only use preallocated per-player space only; extra in-game events, npc AI, etc. could still use some allocs) It's great Folia can get more out of the bad situation, but... oofff.
(not a criticism of Minecraft by the way, it's better to be popular and inefficient than not exist - I just didn't realise quite how much overhead there is due to the runtime)
But then the new APIs in JDK itself are designed such that you have to allocate loads of short-lived small objects. I was told that HotSpot does deal with them reasonably well to avoid them degrading the performance, but apparently it isn't very good at it?
So while it makes sense to avoid unnecessary allocations by using different APIs (e.g. not creating Streams in hot paths), pooling brings far more new problems with it. It might make sense for large objects, but generally requires in-depth analysis to make sure it actually helps.
Also, when plugins come into the equation, you need to make sure that those can't modify objects they aren't meant to modify, which involves copying of objects. Additionally, some objects have different representations in the API (what's used by plugins) vs the implementation (what's used by vanilla minecraft), so converting between those representations is another source of allocations.
No idea how to do it in java but a few pointers ought to be enough. You can also omit clearing the memory area between allocations for things that aren't security sensitive, if that is done in java, which I would assume.
There definitely are scenarios where pooling might make sense, but basically the low hanging fruits in that area in minecraft are already reaped.
For sure there are other tradeoffs, and whether it is worth it. But when we actually see many GB/s that is not cheap even if you are able to offload to other cores.
The article also concludes:
>It is funny to consider that having TLABs is the way to experience more frequent GC pauses, just because the allocation is so damn cheap!
Sounds like a nightmare, the problem just snowballs, because now you'll be tempted to get into GC tuning.
I worked on optimising a java library a few years back and one of the bigger speed ups was removing all the object pooling code.
1. by using finalizers
2. by using object initialization blocks
If you avoid these two problems you get nano-second performance, because the JVM does not need to run code during object creation and GC.
The world is not static though. Each player loads in a lot of living entities (monsters/animals/...) and block entities (furnaces, redstone, ...) that all need to update. There's some overlap of course, though in a game like Minecraft where there's a near infinite world to explore that overlap is smaller than you would think.
(note: armchair remark, I honestly haven't a clue)
Block entities (those with complex data, like inventories) can be read/interacted with by redstone (a regular-ish set of blocks in terms of data), which can then trigger a piston or dispenser (which changes the world), and said piston or dispenser can then interact with the players movement, pushing them, or hitting them with an arrow, or dispensing water that changes their movement.
Same again for non player entities, but also factor in their AI having to adjust pathing based on the world changing.
It really can't, though. Parallelization is the enemy of consistent game mechanics, especially in a complex sandbox game like minecraft. There are tons of things that interact with each other in minecraft's world - if you just tick everything at arbitrary times, those interactions cease to be deterministic and might break entirely unless you write a ton of spaghetti code to deal with every edge case. Redstone in particular is the best example of a mechanic that heavily relies on the synchronous nature of the game logic loop.
Sure, but that list of entities in the area is close to static (apart from crazy redstone magic). One would expect them to be pooled and not have many allocations for each tick.
~Normally, yeah, but Folia is based on Paper which, AFAIK, is known for its unreliability when it comes to the technical aspects of Minecraft (ie redstone and farms.) Even your bog-standard simple item sorter is apparently a bit hit-or-miss on Paper.~
Edit: Possibly outdated info, things sound better re: technical Minecraft support on Paper now according to 'ocelotpotpie.
Ah, good to know - been a long while since I last tried Paper.
My simple sorters have been working just fine(tm) on Paper in the last (>5) few years.
I do hear co players having occasional problems with fancier set ups, but most automation just works.
But that's not even what we're talking about here. This was a short-lived test server with new environment, not huge farms. The super high GC stats we see here are the baseline behaviour. More complex scenarios will push this even more.
It might not inefficiently allocate all that RAM every tick, but it still has to scan a lot of RAM, and write to a bunch of locations to update the game state.
It loves CPUs with huge caches because of that.
There were 113 regions in one of the pictures- call it 9 players per region, with ~927 live unique chunks (8x8 chunks for 9 players ~=566, plus a 19x19 area around the world origin thats always loaded). From what I can tell a Minecraft block is a bit over 10 kB. So ~10 MB per region.
If each thread is allocating 10 MB for each region on each tick (which the screenshot shows is 8 ticks per second) *that would work out to 8.6 gigabytes per second*. IMO, that's too close to be a coincidence- I'm thinking there are a lot of shared chunks per region, but the threads are copying significant amounts of data by value or they just have an absurd overhead.
But then again, Notch never claimed to be a great programmer.
They've upgraded to OpenGL 3 and shaders now; but the game's memory and CPU usage is through the roof with inconsistent performance and huge lagspikes.
Most of this is caused by utterly spurious and unneeded allocations; stream code on hot paths and complete overengineering and abstractions totally unfit for purpose. By that I mean heavy usage of inheritance and indirection in rendering code, and horrible memory wasting with vertex data. (Most things are full floats even if they could be packed and in some cases, even duplication of the same data due to bad data design)
With that being said, I'm happy to investigate in detail if you can tell which parts you are interested in.
You are right, the JVM has many optimisations built in to de-virtualise calls and avoid allocations and the such, but if you abuse it, the JIT will give up. I don't think any JIT will de-virtualise or perform escape analysis on 5-6 layers of indirection with branching.
The problem is code like:
while(game_running) { position = new Box(x,y,z) }
instead of reusing objects:
while(game_running) { position.set(x,y,z) }
plus, Minecraft is not static! There are many thousands of other things moving and changing state in the world.
Right now they're in a bit of a transitional period. Since a few weeks ago the most obvious thing to try is using the latest Oracle GraalVM with ZGC. ZGC is a pauseless GC like Shenandoah and the Graal compiler is a lot better at escape analysis and removing allocations than the stock C2 compiler especially now Oracle made the enterprise compiler free to use.
The main problem is going to be that ZGC Generational is not quite launched yet. Without generational GC it may not be able to keep up with those very high allocation rates, even with a better compiler reducing them down again. Generational ZGC should be in Java 21 but then you'd have to wait for GraalVM to catch up. So, probably it's worth trying that combination in the next six months or so.
Whilst I don't know about ZGC vs Shenandoah, the performance improvements from using Graal EE over C2 can be large even for Java (the wins are much bigger still for other higher level/more dynamic languages). But the Minecraft community isn't really known for adopting the latest JVM tech. They're pretty conservative.
Edit: someone downthread linked to https://github.com/brucethemoose/Minecraft-Performance-Flags... which talks about all of that so I guess they're getting more experimental!
Yes. But no. But yes.
The code couldn't been written better to allocate less. But the runtime encourages cheap allocations you don't think about. But the runtime could provide/encourage tooling that makes it hard to make that mistake. But...
It all overlaps. Sure, JVM is cool and handles it, but also what JVM is influenced how people instinctively used it.
Games often use their tick rate (iterations of the event loop per second) as a quality metric, and an rt OS should favor preemption and low scheduling latency over throughput... so maybe that should give a more fluid experience?
But the bottleneck for minecraft seems to be in memory allocations, so an OS that can schedule threads rapidly may not change much after all.
But Java is heavily used in high-frequency trading, is the backbone of many cloud infrastructures and it is such a large platform that there are plenty of small niches where it is used and performance is important.
I'm glad you said that, it's a very wise statement. It's the origin story of so many working-but-inefficient projects. We just don't see the graveyard of perfect-but-never-finished projects on the other side of the scale.
At Hypixel we ran (and I believe they still do run) a custom fork of Spigot from 2014-ish, with features from each sequential update to Minecraft being added to our fork via our Spigot fork. This let us diverge greatly and customize the Minecraft protocol to our own needs, saving hugely on internal bandwidth and letting us optimize crap out of the L7 "BungeeCord" reverse-proxy we used.
Just don't expect that innovation to come out of Mojang/Minecraft
Minecraft would be dead without the community devs and artists.
https://web.archive.org/web/20120713082903/http://www.mojang...
That's a big claim. Got a source?
If anyone is curious why you’d want to have 1000 players in minecraft, this video documents a pretty amazing example scenario:
I am interested in multithreading and parallelism. So I have a journal entry to explore which is about deliberately desynchronizing and resynchronizing game loops for performance.
* For the number of clients can use epoll or liburing.
* Can multiplex multiple sockets per thread my epoll-server does this.
* Can split out recv and send across threads so you can send while receiving and receive while sending
* For the game loops can divide the territory covered to a different game loop.
My question becomes how do you synchronize game loops across threads, you could latch on each thread for partial causality between game loops.
> latch on each thread for partial causality between game loops.
This sounds like a recipe for misery. You have to start thinking about ""light cones"" if causality propagates at finite speed. Do all observers in the game universe witness events in the same order? What if they don't?
You get a little bit of this with rollback netcode, on the level of a few frames, but in that case there's always an authoritative causality and the client's guess is subordinate to it.
It also sounds like fertile ground for exploits, both of the duplicating-object type and the glitch-through-geometry type.
Your comment made me amused because it made me think of a game world where retrocausality was in effect, it would be absurd.
Is the bottleneck for networked games the network of broadcasting updates? Or the CPU usage of updating object states?
Sharding or instancing means you have fewer clients to broadcast stream updates to.
Different game loops interacting with eachother, that's interesting.
How else do you scale a game engine across threads?
Bit of both depending on precisely what's happening.
Ultimately the problem is that, for N "agents" (players or NPCs or active blocks or monsters, etc) in a space, if you're not careful you end up with O(N ^ 2) checks of the form "has X collided/interacted with Y". Octree systems can spread this out, but you're still vulnerable to "what if all the players go to the same spot?"
Instancing lets you back off the worst-case situation by limiting the maximum number of players in a particular spot.
>The server was prepared with a 100k x 100k block pre-generated world. Our custom plugin distributed new players to the least-occupied region.
Minecraft chunk generation is notoriously though, and even a few players generating new chunks will bring the TPS down to an unplayable state very fast.
And that's knowing a big part of the chunk generation already happens off the main thread.
This is true to an extent, yes. Chunk generation can be real tough. That said, chunk generation has been re-written in Paper, so it's substantially faster than vanilla Minecraft.
When we did a large scale real player test on a much earlier build of Folia we had ~327 players at peak and we did not generate chunks because we specifically wanted to see how it ran. We didn't max out at ~327, we just didn't have more joins than that.
The project improved a bunch already so that number without chunk generation is very doable to beat if you didn't pre-gen any chunks.
Generally it's free performance to pre-gen your world though. Highly recommended for regular players.
That's not happeneing, it says "pre-generated world"
Not many systems try to do this. Improbable, Minecraft, Second Life, and Roblox are the ones I know about. Almost everybody else shards.
Any other examples of really big shared seamless multiplayer worlds?
Time Dilation came in to run at 1/10 game/real time ratio, allowing fights to scale better and be a less frustrating experience for participants
Not sure how it is now, but back in 2014-2016 you could inform them of big battles ahead of time for a specific system, at which point they would, in their own words, reinforce the node (moving that particular system to its own allocation of resources). This often was the difference between being able to duke it out in an epic space battle or wait for the system to load while everyone lagged to death.
But there's the bottleneck, because they can only do the calculations of ship movement & actions on a single node. I'm sure it's been optimized to no end as well. IIRC it's written in Python, but that's not going to be the main performance bottleneck.
I have only the smallest of clues about distributed systems, the only way they could scale it up is to somehow make it so they can run a single solar system or cluster of ships on multiple servers, but for that you get the overhead of inter-server communication or you need an asynchronous, eventually-consistent game instead of something realtime.
largest multiplayer video game PvP battle (8,825 players) most concurrent participants in a multiplayer video game PvP battle (6,557 participants)
Planetside 2 got up to 2000 or so, and may hold the record for a seamless land world MMO.
[1] https://www.reddit.com/r/ultimaonline/comments/tr1r6j/uo_atl...
Edit: And there are apparently private shards that still have that player count per shard: https://news.ycombinator.com/item?id=36490436
largest multiplayer video game PvP battle (8,825 players) most concurrent participants in a multiplayer video game PvP battle (6,557 participants)
How is that different from sharding into mini 20 player servers? I was hoping for 1000 players all in visual range :)
Eve Online record is >6000 players all in visual range, but they cheat by lowering server tick from 1Hz down to 0.1Hz https://wiki.eveuniversity.org/Time_dilation
[Speculation, I don't have the data and wasn't involved in the test] The test placed each 'team' 10240 blocks apart, which is close enough for a ~30 minute walk to get to another team (~4 minutes if you take the nether), so I assume a lot of bubble merging/splitting happened. It's also far enough away that, if the regionizing logic allows for it, you're split off onto another thread relatively early into your journey to visit another 'team' spawnpoint. This is a relatively realistic scenario for an SMP, where towns are usually 2,000 to 20,000 blocks apart.
And inevitably, people still find them, usually through advanced techniques -- hacks that tell you whether a chunk is newly generated or part of an existing 'chunk trail' can be used to hunt down players that have taken great effort to hide their location.
Anyways, a distance of 10240 is just how this particular test was set up. The regionizing is still useful in a much cozier world. As I understand the only way for it to be guaranteed that everyone is inside the same bubble (and therefore the ticking is all on a single thread) is for everyone to be in the same 768 block radius, or for there to be a line of players spaced this distance apart. It's rather atypical for a server to organically develop this way, since people really like to explore for hours before settling in the perfect spot. But some heavily planned/curated server are like this.
EDIT: Correction, as doctor_phil points out that's a quadratic scale not an exponential scale. Derp.
A small correction - Folia regions are not Minecraft regions. They don't necessarily align with the 512x512 region grid, and can take different shapes and sizes depending on what the regionizing algorithm is doing. The term "bubble" is used in some of the documentation and may be a more accurate description.
AgentK20 explained the technical bits along with the other reply to you, but basically the "actually all in one world" is the big part. For sharding you're handing players off to different instances. With Folia someone can just walk from one person to another without any issue or lag.
Folia dynamically groups people into regions depending on their distance. So it was designed to have people spread out to be able to support many regions across many threads. You can put 1000 people in one spot, it just gets very unhappy and you're now not really taking advantage of the whole point of Folia.
It's not a solution for every server or person but it's definitely cool because it's another tool in the Minecraft tool box. And the API is very similar to Paper, whereas the sharded options start to get kinda tricky quickly.
Also what's going on with the generational Shenandoah GC spikes in that graph? There's some pretty crazy spikes that completely dominate the pause numbers and make it pretty hard to read the graph. It looks like GC pauses kept increasing up until the server crashed?
Mineplex has won the Guinness World Records award on January 28, 2015 for having 34,434 concurrent players, the most on a Minecraft server at the time.
Mineplex had the most concurrent players but they weren't on a single individual bare metal server or single server jar instance.
A lot of large "servers" like Mineplex, Hypixel, etc run a proxy which sits in front of a bunch of other servers. The concept of a "server" can have many meanings in Minecraft so it gets fuzzy quickly.
I'm not sure if this 1,000 person test is a record but I'm not aware of any other test running so many concurrent players on a single dedicated box with a single instance of Minecraft running. Folia takes advantage of more CPU threads so everyone is in sitting in the same server instance.
Hypixel beat out Mineplex later on, BTW.
You generally want good performing (not VPS or "cloud" instances) but you can run a bunch of different worlds and have people with cross server talk, the ability to warp between worlds, etc. Most of the larger networks and nearly all of the smaller networks use a variation of this kind of model.
* https://www.spigotmc.org/wiki/bungeecord/ - Older, less-performant but still used by some teams
* https://papermc.io/software/velocity - Newer, more performant, maintained by the team that makes Paper, one of the leading performance MC server implementations.
For context of scale btw, when Hypixel had 200k+ players online at peak we had something like 2,000+ 1U E3-1271v3's each with 32GB RAM, all colo'd in a single DC in Chicago. Egress is 70-80gbps or so 95th percentile, with most months (at high peak) egressing 10PB/mo+ of real-time, uncacheable data.
I thoroughly enjoyed creating different game modes, similar to Hypixel. Whenever I would play on Hypixel I’d think about how the games had been implemented.
Seeing this thread brings back a lot of nostalgia. Reading through these comments makes me realise that a lot more work went into these servers than I had ever imagined. Naive me thinking it was just a bunch of spigot servers with bungeecord thrown on top.
I’d love to revisit and get back into it all. Alas, I’m stuck working on software nowadays instead.
You "win" Guinness World Records by paying for them.
GWR is purely just marketing, none of their records should be taken with any merit.
Is it safe to say that this isn't possible? Which flavor of minecraft server is the best documented to write mods for, Spigot?
"Plugins" are for bukkit-based compatible servers, like Spigot and Paper (and Folia, with some modifications).
"Mods" are more for forge, sponge, etc servers.
So you can kinda pick which direction you want to go. But the Paper project has a ton of API documentation and a few Github repos to even get you started writing your own plugins. The Discord server has a dev help channel with people happy to answer any questions.
The Sponge team has a similar setup as well if that's the route you'd like to go.
My understanding is server admins will spend money on mods / plugins. Are there any server types more likely to garner paid license or plugins?
People do buy plugins but the vast majority of the plugins you'd want to use are free open source plugins. A lot of the paid/closed source stuff is kinda crappy or misleading anyway.
You can run a great server using free plugins and software. It might not be a money maker, but if your goal is to have people play and have fun and not buy stuff it's super easy and doable.
People do still donate to projects to support the authors who make the wheels spin.
- Fabric (and Quilt), which uses a mixins system, a really flexible modding API that can do client and server modification
- Paper (and Purpur, Pufferfish, Folia), which is similar to the old-school Bukkit/Spigot ecosystem and supports an intuitive API for server-side plugins. If you're aiming for >30 concurrent players you really need the sorts of performance patching that Paper comes with out of the box, or some custom-developed equivalent to it.
- Forge, which is kinda a general purpose modding API, good for heavy client+server modification like FTB packs
I'd love to green-field everything in Fabric, but the ecosystem is not quite there yet for more serious server setups.
Fabric has quite a few performance mods that can achieve the same thing. I launched my new 1.20 map last week, and TPS was hanging on at 40 players without most of them even enabled.
Which one is the least headache?
But other things you should also be looking into:
Skript provides an interesting easy-to-use DSL language for Paper/Spigot servers - https://github.com/SkriptLang/Skript
Scarpet provides an interesting easy-to-use DSL language for Fabric servers - https://github.com/gnembon/fabric-carpet/blob/master/docs/sc...
Opencomputers and Computercraft are mods that adds computers to Minecraft, running Lua. OpenGlasses2 lets you code your own augmented reality glasses in Minecraft, because why not? https://www.curseforge.com/minecraft/mc-mods/opencomputers https://tweaked.cc/
Pneumaticcraft, Psi, and ProjectRed are three other very logic-heavy Minecraft mods
Minecraft Pi Edition was an ARM-only release of Minecraft Pocket Edition that is scriptable using TCP sockets, with APIs for various programming languages. Unfortunately, it was dead by the time it was finished, as it never got any updates after the initial release, doesn't have ARM64 support, and is a very limited version of the game https://www.minecraft.net/en-us/edition/pi
Minecraft Education edition has a Scratch-like programming system, but it's only available to teachers and is another thing they sorta dropped the ball on https://education.minecraft.net/en-us/resources/computer-sci...
That's an interesting and legitimate strategy to the scaling problem, but good to know the downsides with the upsides when trying to take lessons from this.
Paper (and Folia) have differences from "vanilla" (official) Minecraft, which is part of how they improve performance to begin with.
So it is true that there are differences from vanilla. Most "in-game devices" - if you mean builds, redstone machines, mob farms, etc - have versions that people have found will work on Paper and Folia. They might require some tweaking but complex redstone and mob farms are definitely very doable.
The goal is to have a close-to-vanilla experience while fixing bugs, patching exploits, and improving performance to run more people on a given piece of server hardware.
The biggest issue with Folia right now is that it's very new so there aren't a lot of plugins that support it just yet. That's changing every day! And of course it still has some crashing issues because it's still in development. :)
It's just.... a very sanitized and corporate experience. And for Java edition players just different enough in mechanics / ux / feel to annoy them.
Very much pushes you to stick to official servers or small friend groups, and is quite aggressive on its monetization.
Folia's main trick is to dynamically split the "main game tick loop", where logic and actions are processed into several threads based on location.
Such that processing for say a mob farm in the South of the map does not impede on processing for a hopper based storage system in the North West etc.
This doesn't solve the everyone's at one place problem (or "Jita problem", in EvE online parlance) but does solve some issues with populated servers where player density isn't extreme.
This isn't likely what I'm about to suggest, but you could move to a broadcast type system where you have some sort of UDP broadcast where you simply tell everyone about all the movements, and a alternative negative-ack channel where clients can ask for missed broadcasts. That way only O(N) messages are sent (the magic happens in the network devices). The issue with that is the internet isn't great for packet loss or multicast routing, so if you have O(N) resend requests you're not much better off. If you had all your clients in one datacenter it might work :)
No comment.
I didn't watch the stream because I was offline when it happened, but I just want to note that the test was run by someone not on the Paper team who got tubbo to stream it, which is how they got so many players. It can be tricky to get enough players for such a large test, so it made sense for cubxity to pair up with someone to get more players.
Twitch can get pretty funky, especially twitch chat, so hopefully there's nothing "bad" in that video, though.