Scaling Mastodon with systemd template units
eigenmagic.com
eigenmagic.com
Environment="NODE_ENV=production"
Environment="PORT=4000"
Environment="BIND=::"
Environment="STREAMING_CLUSTER_NUM=2"
You probably want to use EnvironmentFile= and shove that file with the rest of app's configsAnd, uh, shove the queue size parameter there instead of
sudo systemctl enable mastodon-sidekiq-default@25.service
sudo systemctl enable mastodon-sidekiq-mailers@2.service
sudo systemctl enable mastodon-sidekiq-scheduler@2.service
using feature designed to run multiple instances as a poor man's way to configure stuff. Let's just say nothing will stop you from also doing systemctl enable mastodon-sidekiq-default@15.service
and systemd will just try to start the sidekiq twiceThe article's proposed way of using templates is TERRIBLE. I had assumed from the title it will show off how to run multiple instances using same template units (say to vary a run dir or whatever), but no, it is using them as glorified config file for number of threads/db conns. Just shove that into env file, you can use env parameters in rest of the file too...
The directives we got most use in org are probably Restart* family and RequiresMountsFor, althought the last one is more and more common in distro packages so it is not even needed that often. The usecase being "don't start DB if mount fails" usually.
Overrides are probably the most used feature overall, we often find that the standard distro unit just needs a little bit of tweaking for our purpose (like adding restart), so we just create file /etc/systemd/system/servicename.service.d/01-restart.conf with
[Service]
Restart=always
RestartSec=30
and voila, service now auto-restarts and that behaviour won't break when package updates, as we're not changing any package file. The few other modifications of existing ones included adding new read-write dirs or adding env variables for proxy.Technically you can even go as far as making services into little containers with all the options for limiting permissions but I haven't got that far. Most apps don't exactly need it either, have readonly/readwrite dir + some limits and it is pretty safe out of the box, I've also noticed many distro packages going that way.
* [1] https://devrandom.eu/blog/post/2022-02-27_increasing_system_...
* [2] https://www.freedesktop.org/software/systemd/man/systemd.exe...
* [3] https://www.freedesktop.org/software/systemd/man/systemd.uni...
* [4] https://www.freedesktop.org/software/systemd/man/systemd.ser...
Another feature of systemd I really like are the timers [1][2] it introduces. It offers a much more sane approach to handling scheduled jobs, and allows for an easy overview of which jobs are running and when the next job will fire via systemctl --list-timers.
[1] - https://www.freedesktop.org/software/systemd/man/systemd.tim...
...but we did remove few thousands lines of fixed init scripts (anyone that tells you sysv scripts are easy and simple is lying to you, multiple projects fail there) *and* an bunch of Monit instances thanks to the features the base service management has. And simplified a bunch of other cases.
So yeah, overall even with issues it is a huge benefit.
It takes 4 whole seconds [1] for systemctl/journald to tell me it has no logs for the daemon on my NAS and it opens 985 files while it is doing it
If it just used sqlite as a backend it might've been useful for analysis (on top of way faster...) but Lennart wanted to have a go at implementing binary DB format badly so we're stuck with it.
If it at least kept a pointer to last file the app's logs were written to the lack of proper indexing also wouldn't be a problem.
But nope, it's just entirely worse than text format. And I do mean that in entirety, acking thru text files is faster than systemctl trying to find log files for the app...
That’s not true, Wants=/Requires= do not specify any order between units, only that without them the unit cannot be marked as succesful. This way they can still be started in parallel.
If you really want it to happen after network is up, you should add a dependency of After= under the Requires=.
Also network-online.target might be a better choice but YMMV.
...that being said
make template unit `myapp@` then just
systemctl start myapp@1
systemctl start myapp@2
systemctl start myapp@3
can also tie them to start/stop together via dependenciesor make wrapper script that starts all and use man systemd.kill modes to send signal to all of them.
In the past I did actually use the templated solution, e.g.
systemctl enable myapp-worker@{1..4}
systemctl start/restart/reload/stop myapp-worker@*
However, in order to scale the number of processes down, you have to know how many are running already. Not that difficult to write a wrapper script around of course, but why isn't it just built-in?You increase it, it spawns more, you decrease it, what next ?
How systemd knows which process to kill ? You want to kill the idle one, not the one currently doing the work.
The working process MIGHT BE IDLE coz it is waiting on network. Systemd have no way of knowing. There is no way to find out which process should be killed without processes itself cooperating.
And frankly because it's rare use case. Majority of apps that use have workers follow pattern of having a process that accepts requests and spawns or sends them to the workers and just accept number of workers as a parameter. Precisely because of the above, you do want to have logic on stopping the worker. And it's cheaper on memory to spawn threads instead of fully fledged processes anyway.
Systemd doesn't know when you need more processes so you have to script it anyway even if it had a feature to modify number of spawned daemons within single service. If you have that already, might as well write the the worker spawner there.
Personally the feature I miss is ability to have watchdog. Auto-restart is nice but if there was an option to "run this command, if it fails then restart main ExecStart app" it would've been perfect, easy way to add healthcheck that checks whether app actually works vs being unresponsive.
I found that the sd-notify[1] package on NPM works well. If your app happens to use NodeJS then it's pretty simple to set up. If not, looking at the way that package is implemented is probably a good way to figure out how the feature works so that you can do the same thing in your own application. Maybe the documentation for the feature has improved, however, the last time I checked (only a few months ago) everything I found was bad. I did stop looking once I found that the NPM package works well.
It seems a bit niche to be in a situation where you need the flexibility of dynamic scheduling whilst only having a single machine. Once you have more than 1 machine, systemd isn’t designed for it anyway, so you’d be looking at a cluster scheduler like Nomad or Kubernetes anyway.
In the single machine case, would it not be better to either:
a) Just spin up a static number n of workers (where n is the number of machine cores) and let them idle if not in use.
b) Use a thread pool (or green threads) inside a single process with shared memory (if e.g. idle memory use by workers is a concern).
I can’t immediately see the benefit of being able to have dynamic control over a number of workers AND keep them as separate processes that need to be adjusted manually.
I guess there may not be a choice if the software is not designed to be run inside threads AND memory use is a concern.
We have the following files in /lib/systemd/system/ to define our sidekiq
unit templates, which are all just copies of the same content as above.
[email protected]
[email protected]
[email protected]
[email protected]
[email protected]
[email protected]
Watsigh