Hod do you handle multi-DC, multi-AZ resilience? What are you using instead of IAM policies that cover every resource?
(Asking unironically, would just like to know. I worked at cloud-centric companies for last 15 years or so.)
Hod do you handle multi-DC, multi-AZ resilience? What are you using instead of IAM policies that cover every resource?
(Asking unironically, would just like to know. I worked at cloud-centric companies for last 15 years or so.)
The IAM etc. will probably done by a combination of technical on organizational measures. You will have certain people doing certain things at least before the solution is ready for IaC. People will create roles, accounts, accesses and such. With networking gear that can be still tricky to implement, with virtualization solutions that is easier today. For databases etc. you can create accounts in them too. Of course, K8s and similar make these things more formalized/ transferable. However there is a lot of stuff before you can deploy that.
People forget however that even if you have hundreds of servers you are tiny compared to the cloud providers. You don't have to have the same breadth and depth of offering. So while you need more baby sitting of hardware, probably will not get nearly as good deals on hardware as the big providers do, you will save their considerable margins. Also, they actually have some of the same expenses too - if a harddrive goes bad they will still swap them basically the same as you do. Big cloud providers will not get substantially different energy pricing than what e.g. a steel foundry would get.
By hosting things on premise or in a nearby datacenter(s) you can shave off a lot of latency too. Some machinery likes to store a lot of data and shaving off latency will decrease your need for thick router buffers because you will not have such a big Bandwidth Delay Product and will achieve the same speeds with much smaller buffers. Building stuff on premise just for you makes some things easier too. Even if you loose some credentials usually you can just hard-reset the equipment as the last resort. There will be no credit card blocking that would affect the operations. If you are less strict with security it will usually matter much less - you are not sharing the hardware with unknown parties and all people that touch it have a contract with the company. So usually everybody want the company to succeed to get the paycheck. You build a deeper know how and can do some optimizations the cloud providers cannot do because you don't have to be general.
I mostly thought about the software infrastructure side. With thousands and even mere hundreds of servers over several locations you already want some uniformity. Would you run k8s? Nomad + Consul? MinIO or Ceph? MySQL + Galera? Would OwnCloud scale to many hundreds of users? How would you unify or integrate access control to all that?
Nothing unsurmountable here, just interesting how it's done in real big on-prem installations.
What I would run depends on the particular customer I would have. I probably wouldn't try to unify or integrate access control much. You are not trying to build another public cloud, you want to develop and deploy applications with reasonable robustness in comparison to the cost and benefit of the solution.
Real on-premise installations usually are a mix of open and proprietary technologies. Many companies probably have a few Windows Server file servers on top of VMware vSAN or some kind of EMC2/ NetApp/ HPE 3PAR or whatever storage with HA capability. There is no S3 compatible storage and all the data is stored on the network drive and referenced in some MSSQL database or stored directly in it. If MSSQL or Oracle is used, they probably run on local storage and have are part of a cluster that are marked such that they are never on the same physical node. You can do similar things with Proxmox, GNU/ Linux and PostgreSQL too with a little more effort. You can run MinIO or Garage for S3 if you want to be cool or buy a supported Ceph installation from e.g. SUSE. Everything will be more rudimentary and half automatic but still a few people will be able to manage that even on a rather large scale without many issues. Of course, if you have bigger needs, there are companies you can by computers from by the rack. Some will offer you a complete cloud management stack with it. It can be Oxide Computers (their own stack), or Cloud&Heat in Germany for instance (that is built on OpenStack). There are so many options.
Given that my experience with AWS is to use the same application level cross region resilience techniques im used to on prem(i have worked with high end unix boxen most of my carear) im genuinely baffled when people start talking about cloud resilience as something magical, and nearly all our traffic happens inside of an private network(MPLS/VPN) that stretches across the different sites.
I really haven't seen any magic multi-az resiliiance in aws that dont have an onprem counter part.
None of this is in house whitebox hardware but relatively standard solutions from established vendors(VMware recently started abusing their near monopoly so everyone is looking for/at alternatives like proxmox, xen and nutanix but arent ready to move just yet).
One thing that's more "magical" in AWS is S3. Their architecture is impressive, see e.g. https://www.allthingsdistributed.com/2021/04/s3-strong-consi...