Running servers and services well is not trivial (2018)
utcc.utoronto.ca
utcc.utoronto.ca
I've noticed not everyone shares this bias, and I'm wondering if I'm unnecessarily conservative or other people are underestimating maintenance costs.
IMO folks managing services for personal use vastly underestimate how much harder everything can get in an actual business environment.
It's not just a technical problem, either. Bigger companies tend to have more expertise-oriented teams (security, compliance, developer tooling, operations, internal infrastructure) which tends to make decisions more difficult than when a single person or team can do it themselves.
A week-long project initially. Now you have to install updates, set up and maintain secure access, reboot or troubleshoot when it dies, etc. Installing things is the easy part.
Our team of 6 sysadmins manages:
- DNS appliances, storage appliances, NTP appliances,
- hypervisors, Dev/stage/prod k8s clusters, some other k8s clusters
- dev/prod Elasticsearch/Logstash/Kibana clusters
- internal GitLab, Jira, Confluence, nautobot, OpenDCIM, a deprecated Twiki
- several internal custom apps
- Probably more I am forgetting.
Nothing gets patched consistently. Everything is neglected to a certain degree.
Every month we had a recipe to spin up a new instance of the Github appliance, and import the most recent backup into it.
A lot of systems do tend to get a little neglected over time, and of course different organizations have different priorities, but I insisted on this because I figured for that particular company the Github instance was one of the most critical components - if it is down people can't work.
One company I was at heavily used SaaS and AWS for everything so there weren't very many systems to manage. We'd just jam everything in Docker containers and use AWS ECS. Patching hosts was just rotating to a new AMI on AWS; patching applications was updating dependencies.
Even if they have one, that sysadmin could be incompetent or out of their depth.
Just started a new job where the last guy was decent enough, but what ended up being built up needs to basically be redone, e.g., a giant /21 network: no server VLAN/subnet, or separate network for network management interfaces or server IPMI.
Or just overworked. In theory you can self-host everything, but there's a always a time tradeoff.
If a business does not have someone knowledge about IT, how can they know that they have 'deficiencies' with regards to IT?
Things are the way they are and they think that's 'normal' because they don't know any better.
Small business: Please IT help my 5 year old laptop can't keep up and I swear it's not a virus. Yes, I updated everything
It's amazing watching friends/family outside of tech work. Within 2 minutes of watching someone's machine slow to a crawl, I open Task Manager on Windows and see they're out of RAM and their machine is heavily paging: Yeah, if you had another $40 8GB stick of RAM your machine would be significantly more responsive.
If you keep it simple, it stays simple.
> If you keep it simple, it stays simple.
Things that work well for one person in isolation don't work at scale or for teams. How do you handle authentication for your git server? What about backups? Manage disk space? Updates? That's before you get to the point of dealing with workflows and integrations, or "it's slow when 10 people clone at the same time"
Most of the things you're struggling to solve are effectively preventing it from actually working as intended.
Is it? If I checksum the .git folder on my workstation and my co-workers workstation they're going to come back different. There's no guarantees that I haven't rebased main, or that I have all of the branches that were stored on the remote. If something catastrophic happens to our main remote, which one of our versions do we restore to?
> It's decentralized version control.
Just because git is decentralised, doesn't mean that it can only be used in a decentralised way. How many teams are pushing/pulling like a p2p network, and deploying to servers/clients from their workstations and verifying that the commit hash of their local repository matches what's deployed? A vanishingly small number of people.
> Most of the things you're struggling to solve are effectively preventing it from actually working as intended
If everyone is using it wrong, the tool is wrong. There are billion dollar companies out there that are based on a centralised git service, which proves that people can (and do) use tools in the way that makes sense, not necessarily as they were designed. Personally I'm glad I don't have to share patches over mailing lists with my coworkers, but you do you.
You may have rebased your local main branch, but that doesn't affect your origin/main reference.
> or that I have all of the branches that were stored on the remote.
Everytime you pull or fetch, you get all the branches stored on the remote. Of course, you're not going to have any branches that were added after the last time you communicated with the remote.
> If something catastrophic happens to our main remote, which one of our versions do we restore to?
The origin/main that's the most recent.
> that doesn't affect your origin/main reference.
Who is to say that my origin/main is the same as your origin/main, or that my origin is the same as what our running application is using as it's origin?
> Of course, you're not going to have any branches that were added after the last time you communicated with the remote.
Exactly, so you're relying on the fact that _someone_ has the latest version without actually verifying it.
> The origin/main that's the most recent.
Assuming all our origins are the same.
However, this core argument is obscured by a very emotional rejection of what the parent is saying - that you don't always need these additional things, and that you can (sometimes? often?) keep things simple. I think that's an interesting point to discuss.
> If something catastrophic happens to our main remote, which one of our versions do we restore to?
Dunno, talk it through? I hope you have a good enough relationship with your coworker that you can discuss your work with them.
> Just because git is decentralised, doesn't mean that it can only be used in a decentralised way
The OP not only did not say git can only be used in a decentralised way, they actually mentioned a git server - ie. a central point.
> There are billion dollar companies out there that are based on a centralised git service, which proves that people can (and do) use tools in the way that makes sense, not necessarily as they were designed.
Nobody argued otherwise. But, it is also true that there are billion-dollar companies out there that use an internal git service. How do I know that? Both GitHub and GitLab sell on-premises versions to those types of companies :)
> Personally I'm glad I don't have to share patches over mailing lists with my coworkers, but you do you.
Rationally, this argument is so off it can only be result of an emotional outburst. OP never mentioned sharing patches over mailing lists, and has in fact stated that it's easy to host git server.
I understand and respect your argument and agree GitHub, GitLab and others provide valuable service. But gees, chill out, man. https://xkcd.com/386/
Agreed, however in my experience advocates for this kind of simplicity "don't need" these things, except they ad-hoc rely on piecemeal solutions.
> Dunno, talk it through? I hope you have a good enough relationship with your coworker that you can discuss your work with them.
This works on a team of 2. On a team of 10/20/50/100, pausing everything for everyone to figure out seems like a terrible idea. And teams of 2 quickly become teams of 10.
> they actually mentioned a git server - ie. a central point.
But they ignore all of the overhead of running a server, and fall back on it being decentralized as a solution to the "problems" of runnign a server.
> Rationally, this argument is so off it can only be result of an emotional outburst.
I'd really rather you didn't stoop to personal attacks on me, especially as I've done nothing of the sort.
> OP never mentioned sharing patches over mailing lists, and has in fact stated that it's easy to host git server.
OP said in his comment "Most of the things you're struggling to solve are effectively preventing it from actually working as intended. ". Given that git was designed for the linux kernel [0], a reasonable criticism of "working as intended" is criticising the workflow it was designed around. Running a git server doesn't give you _any_ way to collaborate or work with people, you need to build all of that tooling on top of it. The linux kernel uses patches distributed by email, despite hosting a git server.
> But gees, chill out, man. https://xkcd.com/386/
I've downvoted you specifically for this part of your comment, it's an unnecessary personal attack.
[0] https://git-scm.com/book/en/v2/Getting-Started-A-Short-Histo...
Instead, I'll just apologize: I'm sorry if my comment came across as an attack. I didn't mean it as such. I've seen too many flamewars over trivialities.
Edit: To expand on my comment a bit.
1. You will have to check with every user when they last pulled their repos and/or made any local change and wanted to push it.
2. While your git is offline and you're figuring out which version is the most up to date your users can't do any work with git
3. You just lost all your issues, pull requests, wiki articles and more that isn't stored in git
4. Making backups is your job as a systems administrator and you just failed spectacularly
You’re responsible for 1000 Git repositories used by developers all over your department. Some (or even all) of those have just been wiped in a ransomware attack.
Whom are you going to tell to `git clone` what from where?
DVCS is designed to propagate code to many people and allow them to easily modify and share with others who also have a copy. If people don't have a copy they can get a copy of a copy (which at some point may have been modified). In a large company you need a Single source of centralized truth. You cannot build a company on the concept of "it works on my dev workstation" or the worse suggestion here "all the company's IP source code is on only on my dev workstation".
Answering your other questions:
Authentication: ssh keys
Backups: the same way you back up the rest of the machine (s3 snapshots of ebs?), or run a second server at a different site with a cron job that runs "git fetch --all" or whatever.
Manage disk space: It's not the 90's anymore. How are you running a 1TB machine out of disk space with a git repo?
Updates: Enable unattended updates in whatever distro you are running. If you are running a separate backup server, pick more than one upstream operating system (redhat, Debian, arch, BSD), so a botched update won't break both.
Git hooks work fine for workflows and integrations.
Is it really slow when 10 people clone at once? How is that even possible on modern hardware with 100's GB of RAM and dozens of cores?
The reason people use GitHub and Gitlab is usually not because they want a git server. For that there are much better tools like gitolite.
SSH and Git.
While most of those tchotchke apps are still still functioning as designed, the PDF generator (that uses pdf.js) in Life-in-Weeks app [2] has somehow broken apart. It doesn't generate the PDFs like it used to.
Despite so much of care and effort put into making these decisions, to not have to maintain/upkeep the software, the utopia remains elusive.
[1]: http://q.ht (served via Github Pages, see https://github.com/gurjeet/q.ht)
That's not to say that a experienced team can run an infinite number of arbitrary services without cost or anything like that, though. There may be a select few situations where the cost of deployment and maintenance is negligible, but that's going to be the exception rather than the rule.
Are you comparing server vs local desktop software?
The author's article is actually comparing server-SaaS vs server-on-premise (or server-self-managed-cloud-vm-container).
Check out some of the HN posts where people talk about how much companies spend on AWS and calculate how many sysadmin/devop salaries those bills could pay for (it is commonly >1). Probably even easier would be to find a company that's about 10-15 years old and see how much their tech spend declined when they switched to the cloud. ;)
Once it’s off site then it becomes a hassle cause security and per usage billing and all the other fun surprises that cloud comes with
Even cloud providers, who have APIs for everything, will fuck up at making services that can be deployed as code in a sane way.
You could always make an Ansible module but then there's the overhead to managing/installing that
Data wrangling can also make idempotent playbooks a bit clunky, too. You get into this 2-4 play "run check, reshape results, conditionally run play" pattern
If I have to check or change something or recreate it even 6 months later I often have no clue what the heck I did.
Github costs $4/user/month, sentry costs $25/month, and 2 digitalocean k8s clusters costs $20/month. For $175/month I can have a decelopment environment with basically 0 maintenance for a team of 25 people, including monitoring and alerting for my app.
Compared to running a local gitlab instance, deploying OpenTelemetry and running my own k8s cluster in aws, it's a complete no brainer to buy SAAS.
These potential vendors have zero or near zero additional cost to activate service (eg: the OptiTap is only a few tens of feet from where an ONT would be install), yet the sales reps call me every few months asking when I want to light up service and aren't ready to even talk price or SLA, despite knowing what we pay their competitor.
The article touches on it, but there's also compliance concerns (access control is mentioned, also retention policies, DR, ability to redact/remove improperly added data)
They briefly mentioned monitoring but, like auth, can be non trivial for a well built system. Something like email is easy until the entire system is down and the alerts don't go out anymore (so you need an external monitoring system)
I also don't see labor cost mentioned very often. How many hours will it take to support?
"Fossil does not require a central server. Data sharing and synchronization can be entirely peer-to-peer. Fossil uses conflict-free replicated data types to ensure that (in the limit) all participating peers see the same content. "
...but still it is easy to set up on a central server :
https://fossil-scm.org/home/doc/trunk/www/server/whyuseaserv... https://fossil-scm.org/home/doc/trunk/www/webui.wiki https://fossil-scm.org/home/doc/trunk/www/server/
Not even in the least - they treat them like cattle. If data center scale tools were easier to use for the average engineer - the gap between managed and self-hosted would start to close. Obviously, software, tooling, process, scale: there are plenty of huge challenges to making self-hosting viable. Personally, I see it as far more doable and less radical than trying to distribute the internet in any other fashion (blockchain or otherwise).
No consumer refines their own petroleum or makes their own steel these days, except maybe as a hobby. Similarly I don't think the market share of consumers who host their own email, git servers or payment infrastructure will ever rise again.
Every IT shop I've ever interacted with has a sizable on-prem footprint. And an even larger self-managed footprint if we include things that are run in the cloud but are managed by the IT shop itself (i.e. Red Hat OpenShift on AWS, HashiCorp Vault in Azure, GitLab in Alibaba, etc). So I think we've yet to see the initial demise of that paradigm.
Even if the trend is that most Enterprises are moving towards SaaS and PaaS products, I think we still have a long way to go until the majority of IT infrastructure is managed by a third-party.
If you're running 10,000 machines then you divide your management costs across them and treat them as cattle and end up spending (not real numbers, obviously) 0.02 person-days-per-month-per-machine or whatever on managing them. But that doesn't mean that with the same tooling you could run just one machine with just 0.02 person-days-per-month, because a lot of the benefit you're getting from scale is the ability to make one decision and do something to all 10,000 machines at once, and it's the decision that takes time and effort.
Does that about sum it up?
Why are there even hackers? People should stop doing tricky things because others might find those things difficult. We're making them uncomfortable and should stop.
One clear thing that needs fixing: Linux desperately needs definitive how tos for every common thing an admin/user might need to do posted and maintained on a distro specific/owned site.
As it is now, when I want to learn how to do {thing1}, I have to sift through a complex maze of stack questions, blog posts, youtube videos, etc. Many are ancient, don't really apply to the distro I'm on, fail to mention that there are other ways of doing it (some of which might be better for a given situation), etc.
Then, when I finally settle on a howto, it fails to work and then I burn hours troubleshooting and tweaking. Eventually, I get it to work but when I reflect on what it took to get there, I couldn't really follow the breadcrumbs of my frustrated efforts well enough to document it for posterity.
Eventually, I run into problems with {thing2} and I realize that the fastest way to troubleshoot is to wipe the box and start clean, but I can't because recreating {thing1} is a multi-hour task.
I think most people don't consider checking those wikis because they aren't using those distros, but right now they are the largest and most comprehensive source of information that we have and most of the content is generic and works with any distro or even non-linux OSes.
Plus the use of Packer and Ansible with CI/CD, you'll see where your builds and deploys break quickly and see those failures before you need a server replacement now.