I've been saying this repeatedly (and downvoted for it repeatedly): if you want truly reliable systems, use simple, boring technology, and don't fuck with it after it's set up, and run it yourself. 99.99% of all these outages are due to screwing up something that already works, something that if it was in your own rack you could just leave alone and not touch at all.
You can't create "technical debt" if you don't change anything in the first place.
> You can’t create “technical debt” if you don’t change anything in the first place.
Rubbish. The bits really do rot, and if you don’t do _something_ on occasion you end up with an entire data center no one wants to touch because the dust in the servers might be structural at this point.
I’m not saying go rewrite your apps against the Kafka instance your junior devs are fucking with, but you have to do something to fight the entropy.
The counter-story to yours is running that database on MongoDB in the cloud on a cluster. Instead you'd be having crazy MongoDB issues, data inconsistencies, connectivity issues when the cloud is down, etc etc.
The solution is somewhere in the middle. You can have modern, supported hardware running a LTS Linux and that counts as boring.
But over time boring IT turns into legacy, and without some tension to the system pushing it forward your standards end up locking you into legacy forever.
Important things go onto clusters, or at least have a (hot or cold) standby server.
Stuff breaks. So you fix it. Boring old stuff needs fixing too sometimes. Problem is, old stuff gets obsolete, can't get replacement parts, because of progress. (or something). It's the same story since the first looms were made centuries ago.
What you can't fix you can't really depend on. Our time scales are just compressed to ridiculousness because the pace of change is off the charts these days. So basically, you can't really depend on anything working more than a few months before falling over. Sucks.
Or it may simply not meet the needs of users anymore.
I would hardly hold the air traffic control system up as a model to aspire to, for example. The only reason we run the old one is that the upgrade attempts all failed.
Tell that to your security team.
Don't screw with it and it will have security issues after a few months?
Fiber optic cables are a great technology, but they don't react well to being cut in half by a backhoe. Is the solution you are recommending that we stop using fiber optic cables, or that we stop using backhoes?
Do you:
1) Don't fuck with it?
2) Make a mitigating code change. Patch / fix it (fuck with it)?
If you must fix it, the correct solution is to replace the affected software with the same (or almost the same) version of the software with the fix. No API changes, no other fixes.
Once an attacker is in your organization he will look for exactly that kind of internal-only backend were exploits are already available and the attack vector is known.
There is no such thing as a internal-only backend regarding security.
Let's assume the attacker used social engineering to get credentials from an unprivileged user and uses these to log in to a remote desktop. (I know there are ways to prevent that but I think there are many examples shown that public facing remote desktop is not two unrealistic) Once he is inside your company he can reach the "internal-only" backend and uses the privilege escalation bug you thought is not worth fixing to get root.
But it's true, it's much cheaper if you can find a way to replicate those or do without.
Are they really showing that? None of the major cloud providers, even constrained to a single region (or even AZ) seems on average less reliable than the on prem datacenters I've seen, and there's
> not allowing customers to do so shows how little you care about them.
While the solutions may not be as complete for all use cases as public-cloud-only ones, are any of the major cloud providers not working to enable and selling their capacity to support hybrid-cloud deployments?
I've definitely seen this where I work - the "old guard" setup the system that put the company in a prime market position, the newer people are just doing API calls and scratching their heads if it doesn't work.
Here's a reddit link because YouTube is blocked here.
https://www.reddit.com/r/programming/comments/bq1dt6/jonatha...
For small hobby projects I simply use a 3rd party 2ndary DNS service.
The more "the cloud" replaces many, many servers at lots of different places, the more the outages (which once happened all the time, but to many different organizations at different times) will become big enough to notice.
So, yeah, not just your imagination.
This is just for the last few months...?