Moving Marginalia to a new server
marginalia.nu
marginalia.nu
I don't know what virtualisation the author has in mind, but static RAM allocation isn't a requirement for VMs. VirtIO's memory balloon can be used to dynamically grow and shrink available memory. It's not exactly something you can do willy-nilly, but when a VM has a large amount of free RAM (like 24GB) you can definitely use ballooning to temporarily reassign memory capacity.
Since the author is talking about using Debian, they could opt for Proxmox to handle ballooning for them. Without Proxmox, scripts similar to this (https://github.com/berkerogluu/auto-ballooning-kvm/blob/mast...) could also be used to control memory distribution, though I imagine a server like this needs something a little more sophisticated.
Then, when the VM runs out of memory, the virtio driver starts releasing pages for as long as the hypervisor lets it, slowly increasing the amount of RAM allocated to the virtual machine. This works until the hypervisor runs out of memory to allocate, after which the driver inside the VM will refuse to release pages and the VM's OOM will kick in.
The default is grow only (start with the minimum amount of RAM and only release more to the VM as time goes by), but it's possible to ask the driver to start "allocating" memory again with enhancements, like the ones proxmox do or by using scripting. With the guest tools installed, the hypervisor knows exactly how much free RAM each VM has.
This shouldn't cause problems for page tables on the client side as far as I know. However, certain workloads may not work well; if the VM allocates memory faster than the balloon driver canrrelease it, you'll get out of memory errors. Similarly, quick allocate-free cycles will prevent any auto shrink mechanism from working right.
You could try and run an experiment if you wish; the balloon driver works just as well regardless of whether you're allocating 16GB of RAM or 512MB. You could run a bunch of VMs with a model of the memory allocation you expect on your desktop and see if you end up page trashing like you fear.
IIRC it locks these pages so they are unavailable until the driver releases them back.
> A potential way around this is virtualization, to run multiple operating systems on the same machine.
What’s the actual issue? Linux can have trouble servicing many page faults in parallel in a single process (technically mm) due to lock contention. Multiple processes should reduce this contention.
Do you have more details?
They’re not, though. There’s a tree of “page tables” (that is, a tree with branching factor 512 or so, in the format used by the CPU) per process. Also, per process, there’s a tree of VMAs (the logical maps from contiguous virtual address ranges to whatever logically backs them) — these are created by mmap and friends. And, regrettably, a lock, also per process (although this lock is a read-write lock, and page faults are reads).
If you have a whole bunch of processes mmapping the same file and thrashing against it, you could end up with contention for that files’s data structures that track mappings, but that seems unlikely. Mappings of different pages of ordinary memory should scale well.
And a VM, for this purpose, is more or less like a process. QEMU (or whatever other userspace host you use) literally maps everything that the VM logically maps, and VM faults are handled as though QEMU triggered a page fault.
As you wrote in the article, you still have the old server that can support your current load. That likely won’t be an option in the future as your load continues to grow.
I'm also feeling out what's a way of working with this machine that isn't a huge pain in the ass. When you've got one instance running on one machine, manual deployments is fine, but I think something more CI-driven is probably going to be necessary to keep sane with 8 index shards and a test environment as well.
Kubernetes would be an option but I have bad experiences with that too. A bit too much spooky action at a distance for my taste. Whole ecosystem feels very fragile and churny in a way I'm not very happy with, and the abstractions designed for hiding away the complexities of dealing with a cluster make running it on a single machine where those abstractions aren't necessary just pointlessly awkward.
Congrats on the recent success by the way ;)
How unprofessional! It's like you're building something efficient that isn't going to give massive amounts of money to cloud providers.
I'm not sure why the change <https://github.com/MarginaliaSearch/MarginaliaSearch/commit/...> wouldn't just point to GitHub itself; is there analytics value in having people click on your domain first?
I think it works now though. It's a pain to migrate several dozen nginx sites, especially with browsers helpfully remembering stale redirects for weeks.
between mandatory 2FA "for your safety", the copilot scandal, and the way M$ treated paying mojang customers who refused to make microsoft accounts - github's future looks bleak.
IMHO it's just a matter of time before the inevitable email: "all github accounts are being migrated to microsoft accounts. your github credentials will no longer work after mm/dd/2026. please migrate to continue using github."
codeberg is excellent.
-“John”
Loved this, didn’t find it chaotic at all.
Not sure if I missed it, but how are you planning on moving data from the old to the new storage? Do you have any concerns with corruption at that stage (validation)?
As described in the post, a lot of the data is also heavily redundant so even if something goes wrong in one or a few places, the missing parts can be reconstructed from the rest.
It's really hard to say how much faster it is going to be, but it's at definitely much faster than the old server. I was not really having performance problems before either, though. The main obstacle was just dealing with insane volumes of data with limited RAM and disk.
Do you know roughly how many pages you have in your index?