The curse of scalable technology
lukeplant.me.uk
lukeplant.me.uk
1.) We have a very large, powerful server we are using for webapps/database. I wanted a simple way to "orchestrate" containers on it (from a friendly web interface or something). I've come to the conclusion that such a solution does not exist - it's either CLI (docker/docker compose) or run kubernetes (which I tried, and got it working, but it's too complicated for me right now, even on one server. Maybe someday)
2.) I want to aggregate various logs from the apps running on the server, and be able to visualize them in something like Grafana. Some currently log to their own postgres database, others to a file. The answer to this is to install half a dozen (or more) services, each with their own config language, quirks, and 200 page manuals, and hook it all together. The good news is I can use a "simple" config since I have less than 100GB of logs a day (wtf?, I have more like 50MB a day))
There seems to be a completely missing middle class of software/devops/sysadmin information. It's either toy programs or "web-scale" 1000 node clusters.
But coming back to the article, it's really frustrating to try to talk to others or find answers. "I want to aggregate logs" causes people to think my logging needs are bigger than they really are. Same with "container orchestration". And then I get told I'm doing it wrong (which believe me, under my current constraints, is the best we can do). I guess overall I wish people would respect my current constraints.
k8s is a non-starter for this use case. So incredibly overcomplicated, and what did you think you were getting for that complexity? It's for large-scale deployments. You have one server. Resume driven development is getting out of control...
If you don't have a dedicated devops team, or you're not using a managed k8s service like AKS, just don't. Please don't. Stop spreading the madness to these beautifully simple environments. KISS: keep it simple, stupid. Running k8s because Docker CLI is too hard, is like learning to fly a F-16 because riding a bike is too hard.
Now I could (and did) add them to the 'docker' group so they can run docker(-compose), but the issue ends up being permissions on any bind mounts. They end up not being able to read logs, view some config, etc.
It's not necessarily a single-server monolith, but a server that is meant to handle a multitude of (possibly unrelated) web-apps at once. I might have to think about VMs more. I really do like docker, and traefik is a god-send (making routing and certificates a heck of a lot easier).
https://docs.docker.com/engine/security/#docker-daemon-attac...
"Running containers (and applications) with Docker implies running the Docker daemon. This daemon requires root privileges unless you opt-in to Rootless mode..."
If they can log into the server and run docker containers themselves, you've probably inadvertently granted permissions that could be escalated to gain root access, if they were a sufficiently skilled adversary (or their user account is compromised by one).
You typically avoid this particular problem by having a CI/CD pipeline build, deploy, and run the containers. No one has permission to do this except the pipeline. Grant devs ssh access to the containers, using appropriate network rules to limit access to trusted IPs, and ideally behind a jump box in a VPC.
VMs give you a lot of OS-level protections for free. For one thing, getting root access to a VM shouldn't compromise the host or other VMs (in the absence of a 0-day, anyway).
Unfortunately, it sounds like you're pretty far down the Docker rabbit hole already. It's probably much more trouble to switch now than it's worth.
Logging - look at Loki, also by Grafana Corp, will run in a container, Nomad has a docker driver extension to send logs to a Loki endpoint (podman has some sharper edges, but probably less so on a single-host setup).
EDIT: I haven't updated this in a while, but https://github.com/jedd/nomad-recipes/blob/master/loki.nomad gives a taste for a very basic Loki job under Nomad. In terms of prereqs, you'd need docker & nomad running, and have a persistent volume (to local disk) configured in Nomad.
On logs, I agree and have looked for the same. A simple way to aggregate logs in one machine, heck it could even be running SQLite, and query via a web UI. Doesn't seem to exist for this scale.
[1] https://github.com/mrsked/mrsk [2] https://k3s.io/ [3] https://dokku.com/
I looked at k3s some, and it is promising. I think overall anything kubernetes related is probably in my future, but I am not familiar enough with the ecosystem to run it in production right now.
I did install rancher w/ rke2 and even managed to deploy some of my apps. It's really nice, and runs okay-ish on a single server once you increase the pod limit, but again, not comfortable running it for production (yet).
Had never heard of mrsk, but bookmarked for the future. Thanks!
Have you tried https://www.portainer.io/ ?
I recently started to use Portainer for this. It seems pretty serviceable.
Companies that size I've worked at often had 2–5 technical staff total. That's where I've spent most of my career.
It contrast, some people feel a project with 25 developers is best described as having "only" 25. And I know that's not the top of the scale.
It's easy to imagine how many choices I'd make differently if I had too many people to reach organic consensus, or could count on a half-dozen team members leaving and getting replaced every year, or couldn't count on a certain baseline of skill or technical taste.
This stuff doesn't really come up when we talk online. We just hand each other assertions without context - "K8s is overkill, you should use systemd on bare metal", or "your app will collapse if you use the DB as a queue", etc.
I think the way we discuss our choices is worse for not clarifying our assumptions about the environment.
It's a knowledge gap problem, which is precisely why it's comforting to seek out and lean on the consensus opinion. The alternative takes work: filling that knowledge gap with data from empirical observation and logical analysis. Well-designed experiments? Well-specified requirements? Design? Nobody has time for that /s. It's much easier to say "k8s bad 'cause I read it on HN"
Ah yes.
My favourite is when I look at a system that is slow as molasses even for one user, and the predictable refrain is: “we can scale up!”
Adding more lanes to a road with a speed limit of 5 doesn’t fix the problem of each car going slowly.
Let’s go on an adventure instead of having a sober conversation about what should be table stakes skills.
- spend time to understand the problem being solved.
- research applicable technologies for the problem which align closest with what the team knows.
- keep an open mind for other applicable technologies, but do not incorporate them without cause.
- accept that there are innumerable alternate ways to solve any given problem, but the time to solve is finite.
- do not emotionally attach oneself to the current solution so that other approaches can be objectively considered and employed when beneficial.
- alter technology used when benefit is identified and risk to success is minimized.