How we built an auto-scalable Minecraft server for 1000+ players
worldql.com
worldql.com
The Minecraft setup will be available and useable by anyone Wednesday of next week. More documentation and a roadmap will follow shortly after.
While developing this did you come up with any best practice recommendations for individual nodes?
For vanilla, I’d try Paper (a high performance spigot fork) and see if you still have problems. If you’re lagging with 6 people while running Paper, you simply need a better CPU.
Edit: For a little more a month, you can get a Ryzen 3600, 64GB of RAM and much faster NVMe SSD: https://www.hetzner.com/dedicated-rootserver/ax41-nvme
To be clear, I'm taking about racking your own servers, not cloud "dedicated servers"
You do need to go a bit up in price for proper server hw - although there are a few aging xeon boxes with ecc ram on the low end now.
https://www.oracle.com/cloud/free/?source=:ow:o:p:nav:0916BC...
> Infrastructure
> 2 AMD based Compute VMs with 1/8 OCPU* and 1 GB memory each.
> 4 Arm-based Ampere A1 cores and 24 GB of memory usable as one VM or up to 4 VMs.
> 2 Block Volumes Storage, 200 GB total.
> 10 GB Object Storage.
> 10 GB Archive Storage.
Just to compare with Hetzner - you would typically be able to get 16 or 32gb ram, an i7 with 4 cores and 2x1tb disk (no ssd). I'm guessing single core performance might be higher than the arm offering - and a better fit for minecraft.
Ed: although maybe core count would win out with PaperMC?
In the cheapest offers right now, there's a Intel Core i7-4770, 2x 2 TB Ent. HDD (spinning rust, not ssd) with 32 GB ram. And a xeon box with ecc ram (same, low price).
Either way, it's auction page, may not always have the deals you're looking for in this price range
You can spin up a 4 core 24GB instance, the CPU is 4 dedicated cores and quite fast.
It's not perfect but it ended up being a better experience than we were having with minecraft realms.
A month ago I ran a Paper instance from a Digital Ocean instance with about 4 GB of RAM and OpenJ9 JVM and it never dipped below 20 TPS even with a larger render distance. This was on Vanilla 1.17.
I am somewhat irrationally biased against J9 because they made us stick it in everything at IBM, but I'm willing to reconsider for better Minecraft performance.
I run my server on a 32GB Ryzen5 PC with the world on an SSD and it often struggles with just 2 players no matter what options I tweak. Keep dreaming of the day the Java parts get the same performence as the Bedrock parts (but I know it'll likely never happen.)
WorldQL uses Postgres under-the-hood to store permanent information. I was inspired by companies like TimescaleDB which build new functionality on top of Postgres’s rock-solid base.
Any high-availability solution for Postgres will also be available for WQL.
For the Minecraft world catch-up example, it’s as simple as querying Postgres for block records in a certain chunk after a certain timestamp. No fancy spatial stuff happens on the Postgres side.
Distributed physics is a well understood problem with no "solution". Just different trade-offs. This solution is cool, but it's not novel or anything. Minecraft is particularly challenging because the entire world is highly mutable.
There are no general solutions to any of this. Just a bunch of custom one-offs. SpatialOS is trying. My opinion on it is extremely, extremely negative. And I'll leave it at that.
Unity and Unreal both have primarily single-threaded gameplay. Writing multi-threaded gameplay code is extraordinarily difficult. Unity has been working on their DOTS/ECS system for years, but it does not appear to be close to ready for the mainstream.
So I think the short answer is "it's really really hard. Like, radically harder than you are imagining. Even if you think it's hard. The value of that work is likely not worth the cost. If most players WANT to play on small servers with close friends then the mountain range of work to effectively support a 1000 player server is not worth it".
If anything, I often wish we could rely on many more cores being available than actually are! It is certainly far far easier to see real performance wins with multicore than by using multiple servers, which introduce very heavy coordination costs.
That's not quite what I said.
I'm happy you're using specs to work on an open source game. Specs and bevy and all the ECS work being done is super exciting and fun. I <3 Rust.
In the meantime there are no major games that effectively scale to, let's say, 64 cores. I don't know of anything shipped that can saturate a 12-core/24-thread Ryzen. ECS alone will not get us to that level of scaling.
Yes it's trivial to to throw audio, networking, and a few other subsystems onto separate threads. Modern games definitely leverage 4 cores. Although several of those cores will be severely under utilized.
Modern ECS designs are rapidly evolving and rapidly improving our ability to better leverage multiple cores. But we're not yet to a point where games can easily and efficiently saturate 10+ cores.
Personally I'd love to see a game like Eve Online that can effectively simulate a universe with tens of thousands of players either spread across the universe or all in one places during one giant battle.
> It is certainly far far easier to see real performance wins with multicore than by using multiple servers, which introduce very heavy coordination costs.
This is extremely true.
That's just for our server. Our clients can make use of cores in even more ways, although they have less work to do and generally have fewer cores available, and you can see Veloren taking advantage of 16 client threads with similar utilization here: https://twitter.com/sahajsarup/status/1431837669391142916.
The most important thing I want to note is that in both cases, you are not seeing tremendous imbalance between the cores most of the time. While there are definitely single-threaded bottlenecks in games, you have to be working pretty hard before they start bottlenecking the workload! Instead, we are just suffering from a combination of general inefficiency and lack of work to do.
So no, I'm going to push back against this notion that multicore scaling for games is some sort of crazy intractable problem. It's not. Like any other kind of parallel scaling, it's trivial in some places, more challenging in others, and depends a lot on your workload (including you actually having enough work to saturate the cores in the first place!). But there's nothing special about games here.
> Is there a specific reason why Minecraft hasn’t been “adapted” from the original game to a robust design that could scale to larger worlds and player populations before your project ?
The answer is because it's a lot of hard work.
I am happy that a bunch of smart and talented people are working really hard to optimize Veloren. Good for you. I hope you help push the state of the art.
Nor is scaling well on multicore the "point" of the game or even an explicit goal (though handling lots of players is)--taking advantage of multiple CPU cores is just one of many ways to improve performance, which we try to tackle on multiple fronts (including increased utilization of the GPU, explicit SIMD, smarter algorithms, allocation reduction, structure compression for improved cache locality and network utilization, etc. etc.). We haven't made any special effort to parallelize at the expense of single-threaded optimizations, and generally only do parallelization within a system, move things to the background, etc. where it is revealed as a bottleneck. And we've mostly done so by utilizing existing Rust libraries like crossbeam, specs, rayon, and wgpu, not rolling our own stuff. So again, there is nothing at all special about Veloren's design or focus here that makes it more amenable to parallelization than any other game would be, despite it being in a genre that is supposedly difficult to make scale.
And that's the thing I'm specifically trying to push back on--the idea that multicore scaling for games is only possible if you have some dedicated cabal of programming wizards who want to push the state of the art. We live in the age of libraries, and wizards are only needed deep in the guts of the implementations of those libraries (just as they always have been, and probably will be until the end of time). A game programmer does not need to understand how a lockfree work stealing queue is implemented (or even what it is!) in order to use "parallel for" to beat the snot out of a carefully optimized single-threaded version of the same task, and it's usually far easier to do the former than the latter.
I certainly understand why for a game like Minecraft, with a lot of legacy mechanics and mods that were never designed to be threadsafe, or a game engine like Unity, Unreal, or Roblox, that similarly have lots of plugins and customer code they would like to keep working, it would be very challenging to parallelize after the fact. And naturally, there are limits to what you can do on a single system, and your game design options become far more restricted once you're talking about 10k rather than 1k concurrent players. But for a brand new game without any legacy baggage, there's really no reason why it should scale poorly on multicore systems.
Try to to parallelize a simplified form of applied energistics.
Applied energistics is a mod that lets you create an item transportation network. There are storage containers and machines with an inventory (for the sake of simplicity make them hold exactly 1 item and let the machines just turn A into B, B into C, C into A). The network interacts with inventories through interfaces. A storage interface makes items in that inventory accessible to every machine. Machines receive inputs through exporter interfaces and send outputs through importer interfaces.
It effectively is a database for items and that is exactly what makes it difficult to parallelize. The vast majority of games have 1:1 interactions between entities. In this system interactions can be as bad as n:m. That's also why it lags so badly with large networks. A hundred machines periodically scan hundreds of inventories.
Similarly, fine-grained parallelism can be employed by storing each storage container behind a mutex or reader-writer lock, or even avoiding locking entirely and just using copy-on-write to update the item state when it is changed (we can either do this by executing all our state changes for each tick at once, in parallel, using Arc::make_mut, which is usually fastest, or if we need to do it asynchronously by using a crate like arcswap, which is slower). This is less efficient than a channel, but it has the advantage that the current inventory of a machine can be read without extracting the item (something you didn't specify as a requirement, but which I'm including for completeness).
Note from what I said previously that we don't actually need to continuously scan inventories for updates at all. The obvious optimization to perform is instead to have channel writes push changes directly to a change queue (this can be parallelized or sharded with some difficulty, but from experience a single channel usually suffices). The change queue can then be read or routed (in parallel or otherwise) to the appropriate storage devices to deliver its payload. If need be (since you haven't given a lot of details), we can also track which storage interfaces are being read by players, and each tick (in parallel) iterate through any players attached to the interface to notify them of new updates to that interface. There are other crates that automatically implement the incremental updates I mentioned, such as Frank McSherry's https://github.com/TimelyDataflow/timely-dataflow, for when you have something more complex to do; however, I have never had to reach for this because (which is why I wrote this post) it's actually uncommon to have something super complicated to parallelize!
From what I understand, this does not sound like it has nearly the complexity of a database :) The major thing that makes database performance harder to parallelize (though to be clear--they parallelize extremely well!) is not knowing what transactions are needed. In this case, though, we have perfect forward knowledge of what kinds of transactions there could be; the only things we would likely want to serialize would be attaching and detaching storage interfaces, and we can batch them up very easily on each tick due to the relatively "low" concurrent transaction count (keep in mind that some databases can process millions of transactions per second on a 16 core machine). And even if we did need to parallelize attaching and removing storage interfaces, it's not a strict requirement that we do that serially--crates like dashmap provide parallel reads, insertions, and deletions, and are basically an off-the-shelf, in-memory key-value database.
Finally, the kind of load you're talking about (hundreds of machines and hundreds of inventories) does not sound remotely sufficient to lag the game if it's optimized well, particularly since if we did do the naive scan strategy, it parallelizes easily (to see why: each scan tick, we first parallelize all imports into storage, then parallelize all scans from storage).
I suspect the problem here is not that the challenge you've provided is difficult to parallelize, or that it implements the functionality of a database or is M:N (by the way--something that is M:N in a hard to address way are entity-entity collisions!), but that the solution is designed in a very indirect way on top of existing Minecraft mechanics. As far as I can tell from what I've read about Redstone, it's completely possible to parallelize for most purposes to which it's put, since blocks can only update other blocks in very limited, local ways on each tick--it might even be amenable to GPU optimizations (in our own game, we would make sure that updates commuted on each tick to avoid needing to serialize merging operations on adjacent Redstone tiles). However, I could easily be misunderstanding both what you're asking for, and how Redstone works. If this is the case, please let me know!
Even more speculatively: I think a lot of game designers, when they think about parallelizing something, think about doing it in the background, or running things concurrently at different rates. While this can be done, this is primarily useful for performing a long-running background operation without blocking the game, not for improving the game's overall performance! In fact, running in the background in this way is often slower than just running single threaded, especially if it interacts with other world state. Many game developers therefore conclude that the task can't be profitably parallelized and move on. But the best (and simplest) solutions often involve keeping a sequential algorithm, but rewriting it so that each step of the algorithm can be performed massively in parallel, as in several of the possible solutions I outlined above. This is the bulk synchronous parallel model, which is the most commonly used parallelization strategy in HPC environments, and is also the primary parallelization strategy for GPU programming. It allows mixing fine-grained state updates with partitioning to maximize your utilization of all your cores, and because you're parallelizing a single workload and partitioning by write resources, it usually has far less contention with other threads than if you were trying to parallelize many workloads at once, each hitting the same stuff. This is the model we almost always turn to to parallelize things unless it's extremely obvious that we don't want them blocking a tick (like chunk generation, for example) and it reliably gives us great speedups without making the algorithm incomprehensible.
Of course, if you wanted to say that libraries like crossbeam or rayon are "miracles made by geniuses" then I'd be more inclined to agree :) But there are similar facilities in other languages too, e.g. folly and OpenMP for C++.
I'd like to hear more about that. nagle at animats.com, if you don't want to say much in public.
This ties into the "metaverse" business. Lots of metaverse talk about big seamless worlds full of user-created content, but not much is running. So far, nothing with user created content really scales. There are lots of little shared worlds, like Breakroom/Sinespace, Facebook Horizon, IMVU, etc. There are big-space voxel worlds, such as Dual Universe and Roblox. There are general purpose region-oriented big "seamless" worlds such as Second Life, which have trouble at the seams and can't handle crowds in one place, the same problem these Minecraft improvers hit in round 1.
Spatial OS was going to fix all this. Their system is basically objects which can be accessed remotely and which migrate to where the most accesses are coming from. It cost over $100 million to develop, they had to do the hosting, and the first four games all went broke due to the high cost of hosting. So, since they had too much venture capital, they set up an in-house game studio and created Scavengers, which is reportedly a so-so shooter.
Roblox has plans to solve this, somehow, by sheer money power. When you have a few billion dollars to spare, that might work.
If we're headed for the "metaverse", this has to be cracked. Somehow.
(I've been writing a multi-threaded Rust/Vulkan client for Second Life / Open Simulator so I'm painfully aware of these problems.)
Cynical to say that Microsoft are focused more on milking as much out of the brand and merchandising as possible rather than actually improving the core gameplay in meaningful ways.
"Can we have more performance?" "I can do you a glow squid?" "What about shaders and whatnot?" "Axolotl?"
This is super interesting and I'm excited that someone has finally made real progress on distributed game world software
Could you elaborate on this point? To me "need" seems to strong; it could also be addressed via replication or a distributed model.
Centralized is absolutely the most straight-forward (and may be the most suitable) but I'd love to see some reasoning as to why it's the only appropriate approach.
1. An optimistic execution strategy (the servers can run their own redstone)
2. A locking system allowing only one server to have redstone current in a given chunk at once.
3. A rollback-based system to repair race conditions caused by (1)'s optimism.
Hostile mobs are ALSO still WIP and:
1. Are only synced if two players are near each-other on two different servers.
2. Have their aggression entirely managed by a WorldQL script which sends messages to the appropriate servers instructing them to call LivingEntity.setTarget on the correct player.
Was hoping nail those both down before I shared it here, so please forgive me! Thanks for your interest and stay tuned.
Perhaps two chunks/servers could communicate fast and precise enough for redstone.
But Imagine the corner of a chunk receiving updates from two different servers.
Mojang have not yet added a "transfer packet"[0] which would allow for region-based switching.
For the most part, one would need to re-join the server under a different proxy pool located in their desired region, usually accessible via a subdomain (us.example.com, eu.example.com)
[0]: https://hypixel.net/threads/why-do-we-need-transfer-packets....
I made a custom minecraft proxy (similar to bungeecord but A LOT less resource intensive) that starts the real server when someone attempts to connect to it.
200 lines of go for the proxy and about as much for the http front end in django :).
Of course! One of the few games that is massively multiplayer all in the same persistent, connected world is Eve Online, and the dynamics that arise from its economy and faction warfare are fascinating!
Server boundaries could move based on where the population is via delaunay triangulation (instead of fixed boundaries), and servers could share high-importance information with their immediate bordering neighbours. (This would be recognized as ghost data on the neighbouring servers.) You could even go further and have neighbours share ghost data with their other neighbours at a lower fidelity/frequency.
A virtual distributed actor system could potentially be used to address any potential downtime or resource waste created by unpopulated zones.
I've been playing around with some of these ideas but haven't been able to turn them into an actual implementation, so kudos to you for actually doing it!
This new approach was primarily made because I thought it would be fun and cool :P
I play with a few friends on a modded 1.7.10 server and we’ve decided to restart with 1.16 due to the lag having become untenable.
We run it on an i7 machine at my house dedicated to it, so it’s not a hardware issue.
Like clockwork the server would freeze for about 5s about every 30s. Using opus our best guess is that it’s unloading chunks and the GC is happening.
We’re going to run 1.16 now which I hope has some performance enhancements and so that we can use more modern Java 16 runtime with its nicer GC systems.
I’m also hoping that Minecraft has at least since moved chunk generation off the main thread since there was no good reason for world exploration slowing down things like it did.
I also built a JavaScript redstone simulator website and I’m curious how you will handle that and other block updates at server boundaries.
There is something not quite right. 1.7.10
After a few years it attempted a re-launch using an approach similar to to the first one mentioned in the article. There were multiple worlds and as you approached the border of one world it would teleport you to another adjacent world. It was cool, but was jarring and suffered from its own complexity issues.
Anyway, this project is super cool. I would have loved to see something like it 10 years ago.
> a victim of its own success due to the large player count. It would often have 250+ players who would build these massive redstone machines
I always thought the redstone restrictions resulted in too many bots.
> Here's a demonstration showcasing 1000 cross-server players, this simulation is functionally identical to real player load. The server TPS never dips below 20 (perfect) and I'm running the whole thing on my laptop.
If it can run on one laptop, why does it need horizontal server scaling? :P
You don't really know where the bottle-necks are until you put 1000 actual players on the same "server".
What do you mean by holy grail? Is it not something that's already accomplished by several games/MMOs?
Unless "spatial MMO" means something specific here.
It's also about he level of trust you are willing to give the clients, you can for example offload all logic to the clients and just have the server broadcast all messages. But then you will have a problem with cheaters that use modified clients.
The same "holy grail" exist in database too, where you want low latency, high throughput/concurrency, and high availability. Where the solution is, just like in "MMO" games, to use "sharding".
1) Battle of B-R5RB in Eve online had 2,670 players on the same shard according to WikiPedia. Their solution to the problem was/is to lower the game physics tick-rate.
That demo is primarily meant to demonstrate the efficiency of the message broker and packet code as if there were 1000 players on different MC servers all forwarding their positions through WorldQL. I’ll make it more clear.
I assume that WorldQL is also used to store monsters besides the blocks, otherwise my understanding is that players cannot interact with monsters from other servers. Is it possible to create redstone circuits on different servers that then interact with each other?
> We're planning to introduce redstone, hostile mob, and weapon support ASAP.
>redstone, hostile mob, and weapon support
Aren't these the slow things? Is it just player position and map data that's synced through a central DB? I would assume the bulk of the work isn't done yet.
[1]: https://www.brailleinstitute.org/wp-content/uploads/2020/11/...
The zones are so large that by nature players will be spread out leading to less interactions. Raids are instanced to just your party. In areas with natural congestion (such as auction houses), things could lag at times.
While combat and movement is realtime, it's mostly waiting for timers, so the latency and bandwidth requirements are reduced compared to a first person shooter for example.
WoW later introduced a kind of dynamic overlay of the same open world area from two realms where it was not very populated in the area so that players would be more likely to actually see other players.
There was also a feature where you could party up with a friend on a different realm and you would transparently be playing on their server in the open world (with restrictions like not being able to trade with them).
I haven’t played WoW since Mysts, so I don’t know how it’s changed since, but WoW was definitely more complicated than just sharding and instancing.
In private servers using code you can look at (Mangos) they manage to run 4000+ players on a single gaming-spec machine without much trouble. Its more of a design problem compared to something like minecraft where the simulation is much more detailed. In some research projects for wow private servers people have reached 20.000 simulated players in a given machine.
Normally your requests would just be dropped if it had to process everything within a 1s tick. Now all reload times etc get slowed down by a huge factor depending on load.
I’m writing up some formal documentation on it now, I really wanted to have that done before this project was exposed to the scrutinizing Hacker News community, but I can’t control what people share! :)
Stay tuned.
Was this all on top of a 9 to 5?
However, once the documentation is complete, someone could hypothetically implement it using https://github.com/magmafoundation/Magma-1.16.x
It would be cute if Azure credits could be paying for a VM that's just running Compact Machines I never visit the inside of anyway, while the places my character actually goes are running near me. For example, once you've built it, who visits the inside of that first Compact Machine in Claustrophobia that's just a self-powering battery full of uranium, water and thermoelectric generators? You'd drown in there anyway if you visit for more than a few seconds. But it still needs simulating.
I could see this really changing how users interact with each other.