You Don’t Need All That Complex/Expensive/Distracting Infrastructure
medium.com
medium.com
Sure, Kubernetes, fully automated CI/CD, autoscaling, lambda functions are all cool and very useful if you have a problem and paying users in that scale.
Majority of the applications even today in 2019 would be served well by the reliable Nginx + (Python, C++, Rust, Go or other language of choice), PostgreSQL and backups. If you keep things simple, you can even have a standby mirror of the same infrastructure to take over if needed, and it would still cost you less than spinning up nodes after nodes for the complexity of running your own Kubernetes.
When you set up such an infrastructure like you proposed, there are still too many wheels in motion(complexity), and even worse they are abstracted away from you further and further. And, I'm sure the AWS bill is not going to be pretty considering they do all those bells and whistles.
Elastic Beanstalk itself is free.
1. It is clearly a vendor lock in. You build your business upon all AWS services, using an AWS orchestrator - good luck ever getting out.
2. Abstraction. It is a controversial opinion. I'd rather have less magic when possible and know what goes on underneath, and possibly have control over. Now, a startup may choose to offload the duties of operating a database(RDS), storing files(S3), performing backups(S3/Glacier) in the interest of time and trade-off control, optimisation and $$. However, considering the generalisation of simpler startups, IMHO operating a small scale Postgres on a standard EC2 instance is not really much of a hassle...
3. Cost - however small it may be in terms of actual dollars, AWS managed services(S3, RDS, SES, SQS etc..) come with added cost over simply running their open source equivalent yourselves.
Again, Keeping it simple - not in terms of what the user needs to do to deploy their application, but in terms of what components do you need to get your application running in production.
If you are starting a business, the least of your business risks is lock in to AWS.
Besides that, you're handing a regular old zip file with your code to Elastic Beanstalk, it creates a VM with the standard web server package installed depending on your language. If later on, you need to go to another environment (and in the grand scheme of things, few companies ever switch out their infrastructure no matter how many Repository classes they write to abstract their database), you create your same webserver set up in the new environment.
Every cloud vendor has a hand-me-your-code-in-a-zip-file-and-I-setup-everything-for-you service.
However, considering the generalisation of simpler startups, IMHO operating a small scale Postgres on a standard EC2 instance is not really much of a hassle...
Why would I want to manage my own Postgres database and worry about backups, patching, etc. when I can pay someone else to do it?
Cost - however small it may be in terms of actual dollars, AWS managed services(S3, RDS, SES, SQS etc..) come with added cost over simply running their open source equivalent yourselves.
And that's more for your developers to manage, I'm sure if you can find some gullible young developers to work 60+ hours per week pulling double duty as developers and infrastructure operators it will be cheaper. I for one am not willing to do that. Every hour that your developers are spending maintaining infrastructure and doing the "undifferentiated heavy lifting" they are spending time not creating features adding business value.
You're talking about adding more servers (and no redundancy) to save a few pennies. I would avoid any company managed like that like the plague.
I actually have walked away from offers where the company was "on AWS" but all they were doing was running a bunch of EC2 instances and hosting everything themselves and expected their developers to be on call for infrastructure failures.
On the other hand, I did take a job where the new manager's marching orders were to move as far away from managing any infrastructure that could be done by AWS. We have to have a serious justification to put anything new on EC2 instances instead of using Lambda or Fargate.
In addition to wasting engineer time and sanity, this is bad for the bottom line. Hardware fails, and if new hardware doesn’t cause systems to auto failover, you lose revenue, customers and trust. If you don’t autoscale, your service degrades when you need it most and you lose new customers and business when you get spiky load.
With all the advances in things like fargate and Aurora Serverless, it doesn’t make sense not to use these. The baseline cost is ~$200/mo for a setup that doesn’t go down when hardware fails or get slow when you get an influx of users. And it’s straightforward to set up.
You are locking yourself into a cloud provider that is a part of a really hostile country in terms of digital privacy. Your customers may not care right now, but they really should. If you can sell that kind of behaviour to yourself because you're "starting", then your customers deserve the kind of business they are getting from you. To me it sounds incredibly unethical.
Seeing that all of the most valuable tech companies in the world besides the ones in China are based in the US, do you suggest they host with Alibaba Cloud?
> If you are starting a business, the least of your business risks is lock in to AWS.
True and I agree, for the largest part.
> Every cloud vendor has a hand-me-your-code-in-a-zip-file-and-I-setup-everything-for-you service.
This is where I'd disagree (perhaps because of my infrastructure management background). The magical abstraction of "give me your code and don't worry about what happens underneath" IMO is not a good engineering practice. It may be for some. But, you're spending time anyway reading through documentation of some AWS service setting up your business only to walk away from it all as soon as you need to scale. I'd rather spend the same time reading the docs of the FOSS software that the AWS service in question is built on. Why? - Now you know what runs your code, and how. There is an inflection point to this idea. I'm not advocating for running your own DNS, DHCP, personal firewall etc. Just the core of what your business depends on. In my example - nginx and Postgresql.
Again, if you have an idea which you want to see if it works - perhaps setting up nginx and PostgreSQL by yourself just to see feasibility is silly. But, if you are starting up a real business and expect it run seriously, it'd be a good idea understanding what runs your business.
> You're talking about adding more servers (and no redundancy) to save a few pennies. I would avoid any company managed like that like the plague.
I intended to bring up cost only as a tangential benefit for a startup in its 0th phase of life. I did not mean to suggest that a company should make infrastructure reliability decision based on saving a few pennies.
By the time you need the complexity of dependable infrastructure redundancy, failover handling, and auto scaling, your company is big enough to hire infrastructure engineers whose job it is to do those and not have developers half-assing them under stress. My point is - for a young enough startup, avoid the need for complexity in the first place.
> I actually have walked away from offers where the company was "on AWS" but all they were doing was running a bunch of EC2 instances and hosting everything themselves and expected their developers to be on call for infrastructure failures.
If a company expects their developers to do infrastructure operations as "extra load", and be on call for some $AWS_SERVICE/$INFRASTRUCTURE, then I'd avoid them like plague as well. They are setting themselves up for failure by skimping on hiring SRE and infrastructure engineers.
By the time you need on call service for infrastructure, you're simply not in the nginx+PostgreSQL realm that I was talking about anymore.
The context of the discussion is not needing “complex” infrastructure starting out. If you are trying to ramp up quickly, there is a way with a click of a few buttons you can have a scalable, fault tolerant solution with Blue/Green or rolling deployments etc.
It may be for some. But, you're spending time anyway reading through documentation of some AWS service setting up your business only to walk away from it all as soon as you need to scale.
Scaling your web servers with EB is just going into the console and increasing the maximum number of servers in your web farm.
I'd rather spend the same time reading the docs of the FOSS software that the AWS service in question is built on. Why? - Now you know what runs your code, and how.
Developers don’t develop on AWS. They still need to know how to set up their local machine to run their software during development.
There is an inflection point to this idea. I'm not advocating for running your own DNS, DHCP, personal firewall etc. Just the core of what your business depends on. In my example - nginx and Postgresql.
They would set up Nginx on their computer as part of their development cycle. They would need to know that anyway.
Again, if you have an idea which you want to see if it works - perhaps setting up nginx and PostgreSQL by yourself just to see feasibility is silly. But, if you are starting up a real business and expect it run seriously, it'd be a good idea understanding what runs your business.
Is setting up Postgres on my computer locally going to help me understand how Aurora/Postgres stores data redundantly across multiple availability zones with Multi-AZ redundancy, automatic failover, read replicas, automated backups, point in time recovery etc?
By the time you need the complexity of dependable infrastructure redundancy, failover handling, and auto scaling, your company is big enough to hire infrastructure engineers whose job it is to do those and not have developers half-assing them under stress. My point is - for a young enough startup, avoid the need for complexity in the first place.
Why do you assume that reliability and scalability isn’t a requirement from day one? Both small companies I’ve worked for were B2B companies that had a few large “whales” as clients. They wouldn’t have been able to keep their first client if they weren’t reliable on scalable.
If a company expects their developers to do infrastructure operations as "extra load", and be on call for some $AWS_SERVICE/$INFRASTRUCTURE, then I'd avoid them like plague as well. They are setting themselves up for failure by skimping on hiring SRE and infrastructure engineers. By the time you need on call service for infrastructure, you're simply not in the nginx+PostgreSQL realm that I was talking about anymore.
If an AWS infrastructure service goes down, there is really nothing an infrastructure engineer can do at that time. It’s dead simple to build out on AWS in a fault tolerant, scalable manner between autoscaling groups of EC2 instances, Aurora (MySQL/Postgres) with read replicas and multi AZ deployments, lambda, etc.
But you still don’t need to hire a full time infrastructure person. You can farm that stuff out to a managed service provider cheaper.
Don’t get me wrong, I personally hate Elastic Beanstalk. If I want to spin something up fast, I’ll use another service that AWS has called Codestar. It creates a code template in your chosen language, your CI/CD pipeline, an EC2 instance (or lambda pipeline for serverless) and your buildspec/appspec files (containing a bunch of shell commands) and your CloudFormation template. But to make the skeleton usable, you still have to know how everything on AWS works.
But the whole point of the submission is that you shouldn’t have to know all that from day one. At the end of the day, you still have a standard Windows or Linux based VM with EB using the same stack you would use anywhere else.
I don't like arguments like this. This is true for any number of tasks that a layman wouldn't care to understand. You still will do them anyway, regardless of whether or not your user cares. It doesn't take a lot of experience to see the correlation between poor engineering practices and poor user experience.
Now, I don't think the author was using this argument to pitch poor engineering, but there are good reasons for setting up resilient and consistent build deploys, as he addresses. I think the only real takeaway is "dont over-engineer your infrastructure" which is just an extension of generally "don't over-engineer."
Every time infrastructure gets talked about on HN the conversation gets so sidetracked. The real problem is that people don't think about infrastructure in the context of a holistic view of a project. The hours you spend setting up your project with resilient and consistent CI builds are many times more valuable down the road.
I think people need to not be so afraid of spending time on infrastructure. People already err on the side of spending less than more time on infrastructure. I don't think it's wrong at all that the author spent hours building an automatic build process for his project, especially because now that he's set it up he can refer to it in the future. Maybe with his newfound knowledge and experience on the problem means that he can leverage it with little effort in the future. None of this is a bad thing especially if you think about your engineering career on a macro scale instead of week to week.
One obvious benefit of the whole single/simple server idea is that you can make changes to the system so quickly and easily, which is very important at the beginning.
And yeah, this may be layered on later when it seems like the idea has legs and you'd benefit from this. It most likely doesn't need to be the first step.
It's called a dedicated server. These days you can get a monster of a machine. Even on AWS you can get a machine with 2TB of RAM and 128 cores.
Once you have a working pipeline in place, you can focus on the features, that's the whole point of the exercise: enabling developers to ship features for users. If I have to SSH into some machine and restart a set of services manually in the correct order, it will keep me from releasing features. I am afraid / lazy / whatever. Spending time on this is not something that will never ever reach the user in any way.
The user does not care if something is using Heroku, K8s, or something else, yes, but the user definitely appreciates a product that can reliably iterate on things without major downtime or similar problems.
It's kind of like saying, "well unit tests are nice and stuff, but does the user ever notice?". I would strongly argue the user will notice unit tests in the long run as they ensure the software keeps being maintainable.
But you do notice when you food is slow to arrive, the food doesn't taste fresh, or the chicken gives you food poisoning.
Email web interfaces got worse, login forms got worse, documentation got worse, search got worse (peaked around 2004).
If developers don't want to work for you because they feel held back in creativity or career progression, support staff have to stay up late for downtime deployments, and managers and/or ownership feel blind to the process, then you can still have a miserable company with happy users.
Quality of the work is an important part of a compensation plan. Eventually this hits an organization's bottom line in the form of turnover, employee productivity, etc. In software development with high salary, difficult to recruit employees, this can be significant.
Bullshit, its just another layer of complexity to be managed and that may fail. My guess is that it will take more time maintaining it than it will save you in the majority of cases.
From my experience, this has almost never been the case. Exceptions were when there was infrastructure for the sake of infrastructure.
I have two types of pipelines both based on AWS:
1. Push code to Github -> AWS CodeBuild spins up a prebuilt or custom Docker container (on a server that I don’t have to manage) to build and run unit tests -> Code Deploy agent runs on EC2 instance that runs a script to deploy website.
2. Same as above but Code Deploy runs a CloudFormation template to deploy/update lambdas and other related resources.
The only part of either process that requires me to maintain a server at all is the web server.
It takes me about an hour to set up a new pipeline at the most. It’s basically modified a bunch of yaml/JSON files.
It’s a one time setup that pays dividends through the whole process.
I’ve also used hosted builds with Azure Devops before.
A "working pipeline" is an abstraction. Your actual pipeline will still sometimes break or misbehave in unpredictable ways, except now it takes a lot more context to debug and resolve those issues. Which is fine if you have dedicated infrastructure teams working behind the scenes to keep all that stuff stable (Heroku is an entire company that does this) but it's no silver bullet.
If you have to SSH into some machine and restart services whenever you release a feature, you will learn how to do that. And in that process, you will have a better idea how to fix it if it breaks and how to design features in an operationally stable way. If you just push to master and let the Machine take it from there, you're living in willful ignorance and are helpless when the Machine inevitably breaks down. And, in the long run, that will have a much greater impact on your feature velocity.
Instead on working on the core products, engineers would have "fun" adding more microservices, redundant Kubernetes clusters, istio controllers etc etc. (all of that with less than 10 users).
As an engineer, you need growth in headcount or complexity if you want a raise.
It's a people management problem that flows from the CEO and form the HR department, not a technical one.
"Because it scales" "But we don't have a scaling problem, and if we did it could easily be solved with some caching".
Unfortunately there is also the bullshit which is hiring in this industry, where solving things the simple way doesn't add anything interesting to your CV.
Especially when you start putting caches on top of caches. You can end up with a hunt which cache has the incorrect value problem.
What it does though when organically implemented, is it allows you to scale dev teams more easilly and it will also let you build functionallity and features more rapidly for an ever evolving business.
This is my take-away from two massive undertakings at two separate medium-to-large scale enterprises that I’ve been involved in.
One was a complete failure and one an enormous success that both have had business altering implications.
You deploy tech to solve a problem. Don't go solving problems that you don't have yet.
people using something just to put it on their resume is a huuge red flag to me.
Is it his fault that the incentives are borked?
Also FWIW configuring Linux machines manually over SSH and building and deploying your code too it, isn't always easier. Make sure you write everything you did down, because in 6 months you will completely forget how you built this thing, where it was deployed, etc...
As other people here have said, agreeing with you, paying for multiple availability zone RDS or paying Heroku is also a good plan.
I am very anti-spending money on software that isn't directly related to a project but a $10 confluence wiki lifetime license is one of the best purchases I have ever made.
I believe it is true about infrastructure, and about features and code as well.
When my team needs to release a functionality for the "product"[1] that we maintain, we have a very simple strategy : we go step by step. The users tell us what they need, and we start to deliver as soon as we can. Our first release only solves a portion of the user's problem, but the second release solves more, and so does the next, and so on. I devised this strategy with the users to make sure they are fine with it: they expect a perfect product... eventually. But in the mean time, they having only some portions of their needs met, and we avoid the feeling of being overwhelmed by a great apparent complexity.
[1] The context is a bit complicated... and irrelevant. We don't exactly maintain a product, but it's close enough for the argument here.
(It is a bit more than a typo)
The real question is why business allows them to. I find it entertaining to try to substitute another profession for engineer in the sentence above and see if it still works. Lawyers, teachers, doctors, civil architects...
This problem usually indicates a weak team, or lack of leadership within engineering. But granted, selecting a weak CTO (who picks snowflake & diva engineers) should be attributed to the business.
Tooling breaks down more often when it's more complicated, and you need more complexity when your tooling abstracts over a large amount of complexity, so the more that your tooling is supposed to reduce your cognitive load, the more likely it will break in such a way that it actually increases your cognitive load.
Still works, mostly. Certainly true of architects, partially true of lawyers and doctors, less true of teachers.
The key difference is the amount of f2f time with users/customers/clients. If you add a dumb feature to your site and you have to talk to users about it f2f - not over Twitter or Slack or Medium, but in person - you're going to give yourself a much more realistic idea of what matters to them.
They seem particularly enarmoured with re-arranging the org-charts.
Ahhhhh! You got me!
Before, when I built the same app for two years, the nifty pipeline was overwrought.
I think the context switching inherent in what I do now makes the nifty pipeline useful.
Kubernetes, multi region, canary deployment, immutable infrastructure, even infrastructure as code is overengineering for a brand new indie project, for sure.
BUT, there is a certain minimum amount of quality, availability and scalability any product needs, especially a new product that might show up on hacker news or product hunt and get 100-300 concurrent requests a second for multiple hours at unpredictable times.
To not use fargate because you have to dockerize your app and build it during a codepipeline stage is silly. That takes what, a week, max, the first time you ever do it? Once you've done it, you can do it faster for future products. To not use an autoscaling and high availability database when it costs the same and is so simple to set up makes no sense either. To not use ci/cd to test your db migrations and new code on staging before pushing to prod is just dumb and asking for pain and suffering. Why put yourself through that when products have been built to make these things _so_ easy and inexpensive now?
For deployment: if you can tolerate rare service downtimes then it is difficult to beat Linode or Digital Ocean. I am probably in a minority, but I also really like GCP’s AppEngine: turn on billing, set the minimum instance count to one and the maximum to two or three, and you have a resilient system. That said, I run my primary web site on a single tiny GCP VPS instance with no fault tolerance.
For heavier compute requirements, I also really like a physical server from Hetzner- for about $100/month you can get a GPU, i7, and lots of RAM and bandwidth.
> I spent hours setting up a simple and free continuous deployment pipeline. Commit to git, Semaphore CI builds the Docker image, uploads the image to a container registry, then pulls & restarts the container on my server. It’s really cool. I felt like a wizard once I got it working … and it was free! I actually sat down here to write all about how to do it yourself, until I realised how long I spent on it and how much value it delivered to my users
If you have a team of some size, the time you gain over the lifecycle of the project can be very significant when you have a smoothly running automated dev/test/UA/live pipeline.
> Not build the fanciest deployment pipelines, or multi-zone, multi-region, multi-cloud Nuclear Winter proof high availability setup.
If your downtime can be measured in appreciable lost revenue or reputational cost, then yes, you need all of that, as well as a solid SRE or 5 to keep it all running.
I typically advise my clients what uptime will cost them extra, per "9" (as in the 5-nines), and let them make their own decisions.
I’ve been in enough meetings working for small startups where our clients grilled us about our resiliency, availability, redundancy etc. and where they forced us to put our code in escrow as part of our contract.
You will need a remote backup and a secondary located service, just in case. Potentially a third development one to not push breaking changes directly into production.
Skipping any of this will make your service brittle in addition to not being scalable.
And it already hits the rule of three... (of "don't repeat yourself") that makes a lightweight Kubernetes deployment worthwhile. (Perhaps without all the bells and whistles.)
Remember, downtime is usually lost money. A day's worth can be thousands of dollars at least, and that's two nines already. (99.99%)
It makes no sense to start with a more complicated setup, until there is clear evidence of a need.
For a start, Heroku or AWS instance will do and they scale. (As the other commenter said)
If you're doing your own infra, you need at least basic fallbacks and safety.
It is by no means constrained to infrastructure, it is architecture, technology, management the whole lot.
[1]http://jesperreiche.com/the-push-for-new-and-shiny-solutions...
I'd make it more complicated and resilient but why?
I think that there is a lot of worries around this, it conditions technology choices, sometimes pushing early optimisations, and it is hard because you don't have the data on how many concurrent users your product will have. The best you can do is to extrapolate from you current data or other products data.
So many people end up following a recipe, it could be Firebase recipe, AWS Fargate, Serverless... you name it. Recipes that promise limitless autoscale.
I think indeed is distracting and premature, but again is hard to avoid it, if the premise is that there will be a massive growth overnight, which is imposible to predict that will potentially kill us, and we have to be ready.
Next idea I have, I swear I'm going to throw all my effort into the frontend and go with as minimal a backend as I can get away with (heck, depending on the application, human-in-the-loop might be fine early on) until I get some validation. It'll feel "wrong", but whatever.
Like, no shit don't build more infrastructure than you need. If you're a one or two man team you don't need anything more complicated than 'push container here, pull down here' or 'push code here' et. al. Once you get above that having a repeatable process (repeatable by someone besides the person who built it) becomes valuable.
The 'every second spent on infrastructure isn't spent on features' line makes me think I'm going to have to blackmail into working on security updates or paying down technical debt. And the picture you picked is cheeky.
I will agree with this statement, "You need the most simple, least time consuming infrastructure that gets you to that point." But that WILL become more than just 'throw it on a linode instance' after you get some customers and more than 2 people working on whatever uber for dogs project you have going on.