Do not do in code what can be done in infrastructure (2017)
rogerjohansson.blog
rogerjohansson.blog
Infra is hard to get right, to change, and standardize.
That's why people use cloud providers, but they are very expensive.
With code, you can solve problems using cheap fleets of heterogenous vps from different providers.
And when you have a small budget, that makes a x10 difference in hosting cost. Not to mention you don't need to hire a super dupper 300k <cloud brand> expert to manage your system, a regular dev will do.
Most products don't need unlimited scaling, zero downtime and so on. It can even crash in prod once in a while, espacially in the early years.
You can alway migrate from code to infra later when you need to scale up and standardize. The opposite is way harder, if you even reach the point at all.
E.G: on a video streaming plateform with a decent audience (800k users a day), we had to transcode user uploaded videos then ensure we had several copies of them load balanced accross more or less servers depending of the media popularity. Today you would put some web scale key value store to manage that, maybe a cloud virtual FS or at least some CDN. At minimum, you'd spawn docker instances, kubernatis if you feel fancy. We solved the problem with redis, nginx and the cheapest naked vps possible. The site ran using only 7 servers, including encoding, streaming, database and the django app itself and was super fast.
To this day it is still maintained by one single dev working part time. It's 10 years old and still uses python 2.7, jquery and a hellish handcrafted css mess.
So sure, infra is the clean solution. But Pareto has something to say about it.
k3s is similar, but older
k8s is kubernetes
k9s is a canine joke and terminal client for k8s
Just so I'm not just nitpicking, k0s and k3s are both good solutions to hosting kubernetes on bare metal or cheap VPSs. Hetzner in particular is starting to get most of the features you typically lack on VPS providers in that price range (software-defined networks, load balancers, storage) packaged in a kubernetes cloud-controller.
What do the cloud providers use? That's what you're doing whether you know it or not. I'm gonna bet it's a mix.
Oh, the classic "There is one administrator" falacy.
> kubernatis if you feel fancy
I hear this mindset about kubernetes very often from people who can do some stuff with, but don't get infrastructure at all.
EDIT: thinking that a regular dev can do what an admin does is simply naive.
But I would think that as my only experience with infrastructure is when it fails and I have to fix it.
But again, it depends upon whom you ask. An infrastructure engineer or ops guys are more comfortable with infrastructure scaling your app whereas a developer is more comfortable with application scaling itself.
Global state is global state.
The enemy is your own mind.
And with cloud instances and VMs providing abstractions that don't map 1:1 to the hardware they're running on, all your infrastructure becomes code to create reproducible deployments, or serverless execution.
You don't build your roads immediately before you start driving, and destroy them after you're done, and build only as many miles of road as you need each time you go for a ride. They're not a great analogy to what we do with computers.
While I agree, I just wanted to point out that Terraform allows defining infrastructure as code.
You need to do some processing, like for example create a summary report of several gigs of really big json docs every few seconds. This is a problem that can really be sped up by using more cores.
Which option will you choose:
1.) Install configure and maintain a stream processing tool/framework and add dependencies to your code? Oh and don't forget to add service discovery, special filesystems and install 7 different runtimes and all the dependencies.
2.) Run Kubernetes + Kafka and several python or node.js microservices (each with many instances, because node.js is practically single-threaded) - one to chop up the data and put it on a queue, another that reads it from the queue and actually processes it then placing the results on the queue and another one reading it from there and responding to the ui.
3.) Create a function that does all the processing per document. Run it using pmap (parallel map, starts off threads in the background) or reducers in Clojure or your favourite framework. It uses all the cores on your machine.
I know I'll chose the more Boring technology in most cases, as most cases are pretty boring and don't need all that advanced services, orchestrators and architectures.
Then, you get pulled into regular meetings to explain the time line for migrating to a proper structure.
Now, there will come a time where you'll need to automate this. If/when that time comes, if you already have decent cloud competency, then it is far easier to tie everything using S3, SNS, SQS, Lambda (or other equivalent) than going the Kubernetes/Kafka route.
Unless you buy hiring-from-a-bootcamp-to-be-told-your-practices-are-archaic-and-pulled-into-regular-meetings as a service, of course.
In AWS for instance, you are one API call away from running a Spark job in Glue or EMR, that effectively does all that for you. Plus you get logging, monitoring, concurrency limiting, retries.
Or, even easier, you could just dump the big json docs in S3 and use Athena to directly query them with SQL.
Zero infrastructure needed in that case.
Be wary of doing in infrastructure what could be done with a little bit of code.
Maintaining a dozen bits of pieces written consistently in one programming language may be cheaper than maintaining complicated infrastructure with various components each requiring separate competencies.
As with everything, tradeoffs are usually involved. Be suspicious of people claiming otherwise.
Pick 5 tools and use them religiously. They're rarely the optimal tool for anything, but unless you're operating at large scale, you'll always get better results sooner.
Nowadays I seem to need 20 different complex technologies to be able to do even a simple service.
I can do that because I had the luxury of learning them slowly over the years as they were taking over the market.
But what about those new guys? They can't learn everything well all at the same time. So they must do shoddy work at least somewhere, the whole system is stacked against them.
Or they need to Ctrl+c/Ctrl+v most of their solution without understanding where it came from or what it does exactly.
And compromise on understanding the important parts like computer architecture, OS internals or networking protocols that they just don't have bandwidth to learn after learning all those frameworks, libraries, DSLs, tools and so on so that they can see something running.
I think it is responsibility of the senior people in the project to manage cognitive load of the rest of the team and keep it at a reasonable level. Sometimes a new tool looks shiny and fun. But is the added benefit worth the disruption to everybody? I think of ability to learn and keep things in memory as a kind of budget.
I try to make good use of it.
Yes, managing the team's cognitive load is one of the primary jobs of a senior technical leader. I feel lucky that I've been able to pick up as many tools as I have, but it was generally only two or three at a time, and the more I look around, the more great tools I see that I still don't know how to use. That doesn't mean they aren't great, but I know that even with my decades of experience, I wouldn't be able to build something with those tools without a big upfront investment in learning them.
Anyone who is trying to sell you on spreading your product value proposition across some cloud product portfolio is either incompetent, or simply deprecating the engineering of your product for a profit motive. In very few cases does it make engineering sense to move product complexity from the most manageable domain (codebase with line-by-line scrutiny) to some bloated web interface with no documentation. Certainly, there are declarative configuration techniques for infrastructure, but I think we can all agree these abstractions leak like a sieve and are prone to frequent breaking changes.
Once you experience the degree of control that you get with all in-house development, you will never ever want to give it up. We don't have total vertical integration today, but it is still really fucking nice compared to what I typically see being bandied about on HN every day. Our foundation is bare metal hosting, SQLite & .NET Core (on windows but we could move to linux with minimum pain). It is really hard to get more independent than this in software without rolling your own OS/DB/compiler. Certainly, we could have selected a more "independent" language & framework, but productivity/security/stability is also a huge consideration for our customers.
We have some other 3rd parties we work with, but at no point do we consider our core product value to exist in their spaces. It is just a 2-way integration with clean separation of duties between businesses.
At large scale, and for application running on end-user devices, I try to keep as much as possible complexity in the software because you don't pay for its execution.
But again, that's a compromise that depends on scale and infrastructure costs.
As with anything software, the more there's code, the more the potential for bugs. Things should always be engineered to optimize for that metric (lower bugs). Of course, if it is cheaper/safer/secure to do something at the client, then its benefits must be factored against the complexity it may introduce.
I think what the author is trying to say, albeit does a poor job of explaining it, is that we could leverage well matured Infra level tools like Envoy, Consul etc instead of executing their features at application layer. Which I completely agree. You dont have to rewrite consensus protocol or traffic routing at application layer, instead leverage the existing infra level tools to achieve it. But you still cannot get away without understanding how these tools function at core.
It may be good, or it may have features that are unsuitable, however you have to remember that adapting it can be harder than writing redundancy handling from scratch on a good VM.
Especially if you have special requirements for data consistency that a heavy database can others handle.
A particular trap is getting this heavy service infrastructure and then trying to scale 100x, which not even Docker (the lightest) can pull off without a lot of hardware, and nanokernels also have problems. And then you get additional infrastructure problems which require even more infrastructure and staff. Which you cannot handle then outsource, which then gets done partly unsuitable or expensive, and you no longer own your stuff.
As an operations person scaling up large-scale distributed infrastructure and tools like Kafka, RabbitMQ & OpenSearch being mission-critical, it's becoming increasingly clear that the consumer applications themselves need to manage the failover/maintenance process of this infrastructure for things to run smoothly and people to stay sane.
As someone also with a fair bit of experience writing Erlang, there's a temptation to not have any of these external dependencies and have everything in software anyway. Erlang can handle it.
The issue I see is that most of the problems they're trying to solve this way could be solved with good application architecture.
However, if the component is key to your value-add, the less likely such solutions will fit your needs exactly, and so it might quickly become a bottleneck. You might then find yourself rolling your own solution anyway. Just use good judgment.
This is why there are so many different database types now, ie. key-value, time series, etc. Relational databases aren't perfect for every problem, they're just good at most problems. But you should probably stick with relational databases unless you know for sure you're better off without one.
> We can make sure that we always have X instances of a given service or process running in our infrastructure.
Kubernetes will make sure to try to start new instances if existing ones fail. It cannot make sure that you always have those X instances running by any means!
This may sound nitpicky - and of course, any code-level solution to redundancy or failure issues won't work if you don't have anywhere available to run on - but it's important to understand both your requirements and what specific guarantees your infrastructure can make. There are cases where "let it fail fast, Kubernetes will restart it" may not be your best choice.
I'd say that all of the examples on the article are bad. Rate limiting, circuit breaking and system configuration always happen on the infrastructure. If you do something on code, it will be redundant and not solve the entire problem (and create many more). The opposite applies to call retrial, it always happens on code, and if you do it on the infrastructure, it won't solve the entire problem (and create many more).
> Pick the right tool for the job
This is such an abused phrase. How about pick the right methodology for the use case? Tooling quickly becomes a cargo cult, be it languages, frameworks, or "infra tools". The deficiencies in said "tools" can often force bad designs or patterns that people apologize for rather than fix. You don't have to pick one tool for the job if you can pick a method that allows an array of tools used in a standard way.
For example, S3 bucket changes. Everyone first thinks "Terraform". But there are certain S3 changes that are either difficult or impossible with Terraform. What if it's easier to allow a developer IAM role to make specific changes to specific objects and leverage said role outside Terraform? You can craft your Terraform to ignore those changes, and use any AWS SDK to make ad-hoc object changes. Maybe later you find that not using Terraform at all for that bucket, and instead just managing roles that can manipulate that bucket, works better for a larger array of use cases.
I don't think I agree with this article though. Infrastructure solutions are good if you require robustness or the same solution is needed by many components that can share the infrastructure.
Robustness in this case would be comparing a layer 8 load balancer and a thread that load balances workers. You could seemingly start up many threads on the same machine, or even segregate them by virtual machines (for example Tomcat does this) but if failure of a single host is of concern (or catastrophic enough concern) then dedicating the load balancer to its own machine is desirable. My point being, if you step back and look at what you're actually trying to solve for it will tell you where the infrastructure should go, and has the added benefit of influencing the rest of your surrounding ecosystem.
If you need a distributed hashmap or distributed lock, (or really any distributed dedicated purpose code), something premade is a good option, especially if that's not your company's specialty.
These articles are always a bit interesting that they seem to ignore the debate is "3rd party/outsource" vs "insource". EC2, S3, etc are all just internal "Amazon applications" to support horizontal scaling that Amazon realized there was a huge market for. Somewhere there is code for these someone is writing, maintaining, and operating just like any other "business app". It's especially apparent at large companies for shared services like email, authentication, storage, compute.
To the end user, anything across the internet is "infrastructure" for whatever app they're using
If you have a "cattle" setup, how often do your server instances die and get recreated? Is this something you measure and try to minimise? Do you work out why any particular instance died?
Vs "yeah everyone's wire is the same length"
Which is fairly easily provable
The gimmick that IEX brought to the table is that they are delaying orders a fixed amount. They own the matching engine so can trivially add that fixed amount in code.
But that’s not nearly as fun to show off as a big box of spooled fiber.
That what is the case?
Based on the 2 previous sentences, he's either saying something like:
"are you aware that this is the [right use] case [for programming language framework distributed orchestration]?"
or
>are you aware that this [overlap between Kubernetes-style container orchestration vs programming language frameworks] is the case?
It appears Roger Johansson is author of Akka.NET from Sweden so English may not be his best language.
Years ago, I would have said the same. But living in the AWS world for a while, I now see that infrastructure-as-code is so damn powerful.
The reason I posted this link was because I think I've turned around and have embrased this next evolution in industry (but its taken a while for me), and thought others here would have agreed with me. I guess not!
VMs and containers probably won't cut it.
Because for your code you already figured out how to test it and how to deploy it. Probably those processes are already set up and running.
Really? Isn't that one of the primary use cases for something like Apache Kafka? (And Kafka seems to be very popular even today.)
note, the term "message bus" is similar, but not quite the same as "service bus"