For monitoring on a meta-level, we used Prometheus and Grafana. Here's what that looked like near the very end: https://www.dropbox.com/s/ktlvvtmppq64u8o/trimmed%20grafana%... Since the ratelimit was heavily in place by then, the numbers are pretty low. For example, the headliner "blocks per second" was well over three thousand a few weeks before then.
One of the most common queries for actual usage though, was "what bases in the last 24 hours (or whatever) have the most chests currently". We'd then travel to these locations to borrow items. Here's that query https://pastebin.com/fjhHGSYz Roughly, it grabs chunks with chests in that timeframe, then uses a recursive Postgres CTE to traverse the https://en.wikipedia.org/wiki/Disjoint-set_data_structure that we used to implement DBSCAN, from the leaf to the root. Then it groups by the root, which in practice means grouping chunks together by which ones semantically have been determined to be part of the same base, then makes a report by quantity per base.
There was also a web UI that showed all active tracks. You could click any player and see their travel history, the route they had taken to get to that location. As well as clicking any cluster and seeing the usernames associated with it. We used this a lot to grab what members were present at a given item stash or base, to make sure we weren't stepping on anyone's toes.
In terms of what it took to make this happen, the stars of the show were probably the main Postgres, which was hand-tuned on top of a hand-tuned ZFS, on top of a NVME drive that was IOMMU passthrough'd to this VM for maximum performance. As well as https://github.com/OpenHFT/Chronicle-Map, which really helped out. We needed a mapping from block coordinate to a bit of data about that location, such as the last timestamp we checked it, what Minecraft block was there last, whether the block is interesting, uninteresting, different from seed, etc. High frequency trading software was perfect for this use case, because it greatly reduced GC pressure on the JVM, since the data was stored off-heap. It ended up at around (roughly) fifty million entries, served hundreds of thousands of requests for block data per second, and requested historical data (from the Postgres) at about thirty thousands SELECTs per second (from a table with over ten billion rows). The reason why this needed to be so fast, is that the use case is downloading bases. Sometimes someone would log in to 2b2t at their base, after having not played for weeks at a time (so the data is in Postgres, not the map in RAM), we wanted to be able to pick up right where we left off on downloading what they had built. This means it needs to get back up to speed with their entire base, as fast as possible, in the seconds before they move away to other chunks, or log off the server. We are downloading bases one click at a time, and bases are huge. The area that is loaded in by a player is 144 by 144 by 256. That’s 5 million blocks. And there are about 250 players online on 2b2t at any given moment in time. That’s well over a billion blocks that we could click on, and learn which Minecraft block they are. But, which locations should we use up our clicks per second budget on? It focused on areas that are being changed rapidly, and areas that are different from the base vanilla terrain. Those are certainly bases, large Minecraft builds. We do an expanding paintbucket pattern that expands in 3 dimensions to whatever has been built there. However, we have to do this quickly, and we have to pick up where we left off. Someone might be flying around their base with an elytra, at very fast speed. We would love to scoop up the terrain that they’re flying over. If the base is particularly interesting, someone flying over a certain area could have had about twenty Minecraft accounts spread all over 2b2t’s map suddenly refocus all of their left clicks on that one area, to grab all the interesting neighbors that it wasn’t able to last time. We peaked at about three thousand block clicks per second, averaged over weeks. So this was a massive data loading and management problem.
What do you mean by hand-tuned Postgres & hand-tuned ZFS. What did you change and why?
But for what I actually changed, off the top of my head, I set the recordsizes in ZFS large enough that Postgres could safely (because of ZFS CoW) have full_page_writes off, and combined with synchronous_commit off, that really sped up the overall system and made the WAL logs much smaller. After looking just now at postgresql.conf, various other things were tweaked, such as seq_page_cost, random_page_cost, effective_cache_size, effective_io_concurrency, max_worker_processes, default_statistics_target, dynamic_shared_memory_type, work_mem, maintenance_work_mem, shared_buffers, but those were not quite as important. (plus some uninteresting tweaks to WAL behavior, since we had replicas that got WAL logs shipped every few minutes with rsync)
Sort of like the NASA management in the Challenger report; as Feynman pointed out, the fact that the O-rings were eroding when they were not supposed to erode at all indicated that the models were simply wrong, and thus there was no guarantee of safety at all, as the "safety margins" were only margins if the models were right; but since no shuttle had visibly blown up yet... Cryptographic and anonymity and information leaks are particularly dangerous that way, because when they fail, they fail silently: a bad crypto scheme still encrypts and decrypts your emails, it just lets the enemy decrypt them too. I'm reminded of my work on darknet markets: DNMs often leaked their IP or other information because there are so many ways for Apache or Linux to do so by default, and when they do so, there's no visible error; the DNM will continue working exactly as intended, right up until the police kick in your door.
Aside from being proactive about investigating anomalies, information leaks are very hard to defend against in general. Hard to see how you could reasonably block this bug without some fancy capability security implementation which would let you do something like require the client to produce a proof that it is in sight of a block and permitted to query its current status. (Then requesting arbitrary blocks would fail by default, and a patch which mistakenly gave everyone a global capability to let requests succeed would be impossible to write without knowing one was introducing a grave vulnerability.)
The main expense was priority queue, which is a $20/account/month payment to 2b2t to be able to join the server faster. Probably nearing a thousand dollars on that. We lightened this by creating a proxy system where other builders could join through our headless, and it would add nocom exploit packets from their account to 2b2t, and peel off the responses on the way back. They got a more reliable experience with less disconnects, and we got some extra capacity for free (completely unbeknownst to them of course haha). By the end, all the overworld accounts were being paid for and used by other people, I only had to cover the 2 accounts in the nether.
About $300 on a new 2TB nvme drive for postgres. An unknown amount on electricity. Other than that, all the big processing ran on an existing homelab server (thanks to fr1kin). About $150 on digitalocean for a droplet to run the headless minecraft with minimal ping to 2b2t ($10/month). Backups were just WAL rsyncs to postgres replicas ran (again on homelabs) by myself and a few others.
How much time would you say went into this both from yourself and the others? Probably quite a lot over the years? (because again, it's very impressive)
Also, how many login sessions would it take to associate a track with an account?
Generally 2 or 3. 1 point of association was given whenever a track went cold, split up among however many players logged off within -6 seconds to +1 seconds of that timestamp (because of chunks unloading on a 5 second schedule). 1 point could be a false positive, but 2 or 3 would be reliably trustworthy.
One way to think about it, is that in "normal" Minecraft servers, it's you versus everyone, but in an anarchy server that allows hacks, it's you versus the server software. There have been many hacks to locate other players, like following the trails that they leave in the world, that you can only see if you inspect some specific bits in chunk packets from the server. Perhaps they posted a screenshot that shows bedrock, which you can brute force their location from in a few hours/days on a GPU with OpenCL that simulates Minecraft terrain generation. Nocom was similarly "fair game", in that it only used the Minecraft protocol, and stayed within the "bounds" of talking to 2b2t with Minecraft packets and extracting information from that.
All that being said, it's still plainly the case that people put a ton of time into their bases, and no matter how much moral sleight of hand and cope you apply, it's still kicking over someone's sandcastle. When it comes to actually traveling to these locations and raiding them, our aim in that campaign was pretty much solely to amass a large stash of items/blocks for ourselves for the future. I would pass on bases by-default, only handing them along to be raided if there was bad faith construction (e.g. swastikas), or if most of the builders were known to be naughty. The main campaign was against standalone item stashes. These do take some time to create (by having exploiting past item duplication glitches), but they don't really have any creative expression in the same way that a base does. I have no logical argument for why I say that the same amount of hours invested in artistically making a build is more morally valuable than duplicating building blocks, it's just what feels about right. During the time when I was involved, we passed up hundreds of bases just because there was something actively being built there, and essentially only hit places that were just rows of chests placed on the ground in the world.
Btw Mirai botnet also came out of minecraft.
For downloading bases, it is an essentially perfect recreation. Some things are missing, such as the color of banners, but for 99% of blocks, just setting the correct blockstates will recreate the build. Some disconnected sections of the build might not be found by the paintbucket floodfill algorithm though, so floating parts could be missing. We counteracted this by having a few random blocks in the chunk be checked on a schedule, so eventually it would get everything.
Connections from university though... I suspect that can be extremely useful, but didn't leverage that in my own time there, so can only guess as to how much (and also I didn't go to an ivy league school or anything where it could have made an even larger difference).