Running servers and services well is not trivial (2018)
utcc.utoronto.ca
utcc.utoronto.ca
I am therefore regularly shocked by how much money is poured into or otherwise spent on tech and tech workers and how laughable the result is. I've really never seen a server infrastructure at any company where I haven't shuttered in horror within a few hours of looking at design, practices, or systems themselves. When I was much more junior I was perplexed by the sudden appeal of i.e. AWS EC2 but then I've seen from the inside multinational conglomerates try to acquire and deploy servers (i.e. just the logistics that a non-technical person could handle, no software) and understand. The irony to me is, if you fail at logistics, you have near zero chance of doing anything else right, so cloud doesn't reduce your liability much unless it is fully managed like GApps or Salesforce. This should be a huge warning sign for a variety of stakeholders, but money's cheap and the livin's easy right now.
Yet, all these companies remain in existence.
Maybe what you consider to be "laughable" is actually "good enough" when it comes down to dollars and cents?
The issue is largely orthogonal to my own skillset because one needs to only view quality regressions in order to assert my point, not be the cause or hero of them (ain't philosophy generous?!). The times a 2 year experience system operator could go in and wreck havoc on a mainframe operation were not particularly common and the standards and practices (i.e. airgapped backups on tape, geographically dispersed sysplex etc) ensured business continuity were generally planned right in from the initial purchase order with the vendors' help once upon a time. In this decade people can and did launch massive financial exchanges on mongodb when it had serve and well known issues.
Running a mail server is easy. Running a reliable mail server with good deliverability is not. Not being spammed into the next millennium is hard. Yes, it's possible (I do it on some inexpensive OVH VPSs) but it takes time to come up with a solution that works without being too restrictive. Plus, spam is a moving target.
That and there are so many nosey devices on the home LAN these days I had to create 4 subnets with isolation so the nosey IoT crap cant poke around my personal stuff
I admin a small, low traffic business site for a friend. In theory, running a small PHP site should be easy.
In practice, I cannot think of a week that has gone by where everything just "worked." From script kiddies to hackers, to web-components suddenly disappearing. To YouTube audits, to random bugs that aren't reproducible.
Large providers will always have an advantadge in that, they can identify adverasrial activity before it reaches my site, and then apply what they learn to every site they operate.
- Email (SES)
- SSL certificates + renewal (ACM + lets encrypt)
- Backup + Restore (RDS snapshots + AWS Backup for EFS)
- ALB (if desired)
- CDN (Cloudfront)
- Firewall
- Auto scaling database (RDS pgsql serverless) with automatic pausing
- Auto scaling storage (EFS)
You can adjust capacity by just dialing up or down the ASG, and our Elixir app auto-clusters using the Habitat ring for service discovery. Packages are upgraded when new versions are pushed to our package repository. All binaries run in jailed process environments scoped to Habitat packages, with configuration management and supervision handled by Habitat.
This is a non-container approach, focusing on VMs. However, the theory is that this will be pretty much turnkey for a scalable self-hosted product on AWS, including software updates. It's hard to say how well the theory will fall out in practice, but I'm optimistic. Avoiding a mess of microservices was fairly important in making this kind of thing possible, we have a few services, with a dominant monolith in Elixir.
It was a bit of a bumpy ride but the product has stabilized enough we have confidence in it and it turns out to be quite a good fit for the self hosting use-case. Basically, it's an all-in-one solution for packaging, isolation, service discovery, configuration management, and deploys. So it was relatively easy to transition our production deployment bits to a packaged up deployed solution to 3rd party servers, with all of the bells and whistles.
Our setup is a fairly minimal custom AMI that just installs the basic package dependencies for Habitat and configures the supervisor. Everything else is bootstrapped within Habitat, including our own proprietary configuration management service which allows a nice web-based GUI for configuring the ring state. The user data script just does some basic UNIX setup and DNS initialization and then does a bunch of Habitat service loads. By abstracting over all configuration via Habitat, a unified interface exists for configuring all the services across all the machines, regardless of their underlying configuration management approach, programming language, conventions, etc.
Something like kafka requires figuring out the configuration you want, putting that into Salt, adding the correct configuration options, deploying it and making sure zookeeper works, and then generating certificates and what not. It's not a simple process.
Setting up monitoring, and other things like floating IPs is a pain. Custom wrappers for Terraform scripts and other components required to deploy the systems you need to run an app. It's a lot
Kafka requirements adding salt to keep your zookeeper happy.
JavaScript went into this stage not too many years back. "I test my react code with mocha chai cucumber in gulp with puppeteer."
Given that there are better and worse abstractions, most are not suitable for most use-cases, a lot of times simpler and more centralized is actually more stable. But when you start needing high-availability-coordination you're going to be damn happy that there are things such as Consul, etcd and Zookeeper.
As for Salt; by all means, start off by having bash scripts in source control, which at some points becomes a pain-point and you'll realize you need some kind of configuration management to not go insane.
Several of these tools are just putting a framework and common language for tasks we have to do either way.
Edit: As a developer, I would rather sketch out the infrastructure that I need, then hand it off to someone whose entire job is to set up infrastructure. I work at a small startup where the devs are still doing a lot of infrastructure, and all it does is create technical debt. Without a solid team that owns infrastructural concerns, developers just end up digging themselves a grave.
A certain amount of configuration complexity will always be there, but there’s still a lot of incidental complexity that could be cut away if we just generated these config a with a general purpose language.
Some days I just want a load balancer that can route traffic to whatever servers are up now and then I'd just have Docker, a service registry, a load balancer and I'd be done for a while. The easy bits should be easy, and they can't even manage that.
Isn't that sort of like someone in Sales complaining about the "state of development" and asking why they can't just press a button and have a custom app to sell, like, tomorrow?
You can, however install the banzaicloud kafka operator with only a few commands and very little config and yield a production-grade kafka cluster pretty quickly.
There will always be a meaningful amount of effort involved in supporting your own systems, and that cost (be it in ongoing maintenance time, or in lost time because your own system failed) _can_ always be higher than paying to outsource the work to a trusted third party.
I don't think this is accidental complexity, it's essential complexity. To the extent that it seemed easier in the past, it's because we ignored some of that complexity and paid the price. Putting up a truly production-grade service is fundamentally hard.
Now, I think it probably will get easier over time as we grapple with these problems. Integrating with auth, for instance, should be easier, and that can be solved with some more code, and formal and informal standards. It's not all essential complexity. But I think a good deal of it is, or at least it is from anything remotely resembling our current perspective.
I think this is the kind of essential complexity which we faced when developing operating systems. OSes are fundamentally helpers, which don't solve the application problem, but make the solving easier (a good OS makes it easier by much).
So we can and should use the results which we got from OS development.
I've been working on building something like this for a year now [1] and it's certainly quite different from how we do backend development normally,but that's also the point :). The tricky part is to find abstractions that are general and non-leaky so they can be leveraged to build a wide variety of software. Feedback appreciated!
Static typing would also support editors that let you choose from allowed values.
Maybe such a language exists? If not, why not?
So you should ask, can you use the common languages - and if not, why?
Good question. I do not know the answer to that. But I think there must be some answer since there are all these different configuration languages, Ant, Gradle etc. They must have been created for a reason.
Just started using it and I think it's probably one of the best software ensembles I've ever used in my career. Completely knocks Docker, K8S etc. out of the water.
I describe the deployment I run with it here: https://news.ycombinator.com/item?id=21468506
Nix is not typed (it would be slightly better if it were), but everything is evaluated before it hits your servers which allows for lots of static checks.
I started writing a tutorial on NixOps here if you want to learn it: https://github.com/nh2/nixops-tutorial (only has 1 part so far, I'd like to show how to bootstrap a Consul cluster and distributed file system next).
https://people.cs.umass.edu/~arjun/main/papers/2016-rehearsa...
No, it’s not Go. Go is a general purpose programming language. What parent poster wants is more like Terraform, but better (I assume). With semantics that reflect devops actions.
You might write that tool in Go :) but the users wouldn’t necessarily know. Just like a gamer doesn’t need to know about C++.
Ansible is not as 'data-structure' oriented, so testing that is more difficult, but each element (modules, libraries, roles, playbooks) can be written to be fairly environment-neutral, which you could then use a test a fixture like BATS to easily check for regressions.
To answer your 'type-checking' expectation, its generally on the developer to implement those tests. In my experience, its rare that those are implemented in publicly available devops libraries, even though we have most of the tools to do so. It's definitely an area where the practice can improve.
https://nickcraver.com/blog/2016/02/03/stack-overflow-a-tech...
That said, thanks to VPS and Raspberry Pi, it's fun and extremely inexpensive to privately deploy and maintain semi-reliable personal services. Anything beyond that requires immense planning, expense, and dedication.
Similarly free CI services are very attractive but when they stop being free, or die, and you have a lot of investment in tests specific to their infrastructure spread over all your apps and again no agency to keep it going yourself, it's less attractive.
"Why not pay someone else to do it" glosses the privacy and security results are not the same when you pass all your email or IP to a large foreign company who may compete in some of your markets, compared to doing it in-house.
Yes it can be unexpectedly difficult to do even a small thing securely and well. But it doesn't mean that it's not the right approach.
Will skip reading comments because I know some k8 fan boy or some other type of religious enthusiast will insist all of it is solved.
To decide of course you need enough data about the situation at hand:
- initial costs
- keep system running costs
- other pro/con items like "availability" (e.g. you can use a local server even if there is issues with your companies internet access) that can be prioritized and rationalized with methods like the cost-utility analysis or the cost-benefit analysis
I think it is hard to give general advice on what to do in such cases.
None of this stuff is easy when you're starting out. Even when you've become used to whatever environment you've created, there are a lot of moving parts to monitor and manage.
I've been running mine on a $5 Digital Ocean droplet for several years without any fuss. Maintenance boils down to logging in once a year or so and installing whatever updates are required (which I'm a few years behind on...). It's survived a few HN front page posts even with this unoptimized, non-scalable, sqlite backed ramshackle setup all without breaking a sweat.
If I'm a 1000 person company, I've got resources to spend on my corporate 2FA infrastructure.
What do I do when it's 10 people? 20 people? 50 people?
We self-hosted git though (using gitolite for access control), running servers was a core competency for the team, so having a little baby server on the side that just dealt with text files for 50-100 people wasn't a big deal. It was running on a mac mini at the CEOs house until he forgot to pay his cable bill once and we couldn't push code for a day.
edit to add costs: a small company can do this for 3-5 bucks/user/month. that kind of cost is doable for a small shop, and worth it.
But how do I now take over their repositories and email, for example?
> edit to add costs: a small company can do this for 3-5 bucks/user/month. that kind of cost is doable for a small shop, and worth it.
Do you have a concrete reference? I really don't want to use a Google-based system. Microsoft could be an answer even if I reflexively cringe at that--they at least seem to be able to deal with businesses properly.
I'm really not averse to paying money for this, but it needs to be seamless enough that we can use it from the CEO to the receptionist.