There is a reason to pick a good strong vendor and stick with it. Big names end up building their own data center (they can probably capx it for tax purpose). Small to medium usually don't have time to deal with multiple vendors. Try to be vendor agnostic is really great but again, even in the world of open stack, you either manage all of your open stack installation yourself (which is expensive), or you end up one vendor running one version behind, or not offering 100% feature.
I think it is more worthwhile to first complete automation and self-healing in one vendor, before thinking leaping to another one. It took so many engineers at Netflix to build a resilient AWS environments. VMs go down all the time, more often than AWS pushing some bad commits and cause half of their native services go down. There is so much more to engineer in other areas and those are more likely to put you out of service first, so unless you want to all your team dealing with fires every day.... we were putting band-aids together most of the time.
I'm having trouble imagining what problems have to be constantly dealt with such that multiple millions of dollars have to go into abstracting the multiple vendors.
If you're spending millions of dollars on making systems redundant, I completely understand. But that is not the same as spending millions of dollars to allow datacenters 4-7 to be on other vendors.
Maybe if you name some problems specific to multi-vendor support that took several man-months to fix I might comprehend. Just a sentence or two description. Thanks if you do, sorry for being confused if you don't.
Anyway, the problem started with bad management to be honest, which I don't think my case is too rare to hear. When the first project rolled in, we ran PaaS on a single provider, but more projects joined, it was time to choose another vendor because of better equipment and better deals. Yet, none of them really provide good resiliency.
With the third vendor probably around 12-18 physical servers were purchased and managed by the vendor. It was openstack. That version of OpenStack at the time was not compatible with AWS, which later becomes our 4th vendor, and runs our dev environment. You want to run additional performance testing, well, we can't get the same elasticity as AWS because we only have so many physical servers to provision VMs, but that's where our real data lives. So we had to do our QA on AWS and doing the data copy can take a whole business day. The script we wrote for Amazon doesn't work on OpenStack. Security, network, I/O metrics are not consistent across vendors. That adds complexity to code. I know exactly how to abstract things, but just trust me :-) it really makes code hard to maintain, and really painful to integrate with multiple vendors. As a matter of fact, I don't like working with AWS API (with boto) myself because API response formats are inconsistent!
Finally, we got rid of one of them, then two and finally we are on the final stage of consolidating everything on AWS and focus on infrastructure automation and lower the number of chores.
Like many projects out there, things usually start out real nice, but then once you get too busy fighting fire here and there, you will accumulate some debts. If one is not careful, the debt can backfire and we had our lessons. There is no resiliency in most of the other vendors because they require you to purchase more servers and they themselves have hard time to go true elasticity. AWS, at least, for the most part, doesn't run out of instance availability that often (it happened a few times to our EMR processing). I do have some issues with AWS myself, but so far, AWS seems to be the only true cloud provider you can hang on to for several years.
Before building a grand multi-vendor infrastructure, build on a single-vendor well, then decide on the next step.
Also, steer clear of "cloud neutral" services and products that will magically move data and services across cloud providers. Interop is the last thing on any proprietary vendor's mind. An example. You have a pair of border edge routers. Do you buy 2 Junipers, 2 Ciscos, or 1 Juniper and 1 Cisco for fear of a bad vendor bug taking out all the routers? I'll tell you which one I would not choose. The .com TLD nameserver requirements used to mandate dual vendor setups. They sure did learn their lesson.
Is there any relevant reading material? I would have assumed diversification to be a good idea for such critical infra.
Vendor diversification is a bad thing for critical infrastructure when interoperation is required. "It's a Cisco problem!" says the Juniper rep. "It's a Juniper problem!" says the Cisco rep. You're stuck in the middle. It's terrible. You can only hold one vendor's feet to the fire and they won't care at all if you're in a heterogeneous environment.
Remember all the middleware products and companies from the late 90's? Neither do I.
Dual-sourcing makes sense in a lot of cases.
It does make sense to standardize on one open-source database package (MySQL in the case of my employer). But that's software that I can take with me anywhere. And it's open-source. So the risks of vendor lock-in don't apply.
> Otherwise, why are you running stuff in the cloud anyway?
How does it follow that if I don't lock myself into one cloud provider all the way, it's not worthwhile to use the cloud at all? Maybe using cloud providers is worthwhile simply because, at a certain scale, they're less expensive than leased dedicated servers, never mind the up-front cost of buying and colocating hardware. Also, it's easy to provision cloud VMs on demand, then throw them away when you're done with them. Those are good reasons to use cloud providers without locking into just one.
It seems to me that the best approach is to use only the subset of features that are common to DigitalOcean, Vultr, and maybe Linode, and abstract over those multiple providers with software like Ansible that can access multiple provider APIs.
I am indeed suspicious of proprietary solutions for deploying and migrating across cloud providers, such as Cloud66. But that's only because using one of those solutions would itself be an instance of vendor lock-in. If there were an open-source package with similar functionality to Cloud66, I would probably use it.
We currently have ~50 servers in 8 cities, across Linode, Digital Ocean, and Vultr. It took me two weeks to craft a ~400 line script that abstracted the server creation APIs for each. Once spun up, they're each bootstrapped with a script that builds each server from scratch identically regardless of the provider (with a couple one-offs for Vultr), because they're all running the same distro.
A whole data center can go down, and there's no reason for me to get out of bed.
Step 1: An if/else-heavy script that will take in a few parameters (for us it's city, a server type, and a numeral for naming) and build a clean server with all of the needed keys populated.
Step 2: A "yum install"-heavy script passed into the clean server, that builds everything needed from scratch, sending status emails throughout the build process.
There is so much to with than just be able to spin up an VM and then run Ansible/Chef/Puppet on it. Heck I can write all of that in Fabric. There is no direct connect on Digital Ocean. I am not sure how you set up VPN with Digital Ocean or Linode. We use cloudformation on AWS, and I am pretty sure there is no such thing on Linode or Digital Ocean. Exception and response codes different across providers. Able to reproduce an environment from scratch is important to us, and of course, we try to do that in stages. I own a DO box myself, and that box turns out to be really slow in the NY region (where I live), maybe I am just an lucky bastard.
But to be honest, did you really build your entire infrastructure in three vendors to begin with? What are your reasons to really build on Linode, Digital Ocean and Vultr? How do you copy your data across environments? Are you splitting dev/qa/ci/sandbox/stage/prod?
It takes a bit of work and determination, but it comes together in the end.
One problem with building on AWS is their lack of network diversity. If there is a network cut, causing congested links, they will do little to try and alleviate congestion to save on cost. Issues in Oregon have caused network degradation between WEST-1 and WEST-2 that last for days with no improvement.
(I haven't used it myself.)
We use it for automatic configuration of everything from colocated hardware to $5/mo VMs on DigitalOcean and other low-cost virtual server providers and it works great.
most of the time they are spending someone else's money, also. you should see some of the deals that have come across my desk. the 'big names' can literally charge 2-5x more than a competitive quote and get away with it, oftentimes with worse deliverables (i.e. long stretches of downtime that somehow get a pass from their customers).
at the end of the day it's a server sitting in a rack in a datacenter, connected to ethernet. beyond a certain level of quality (tier 3 dc, server class hardware, an enterprise quality network) it's really all the same. people should choose their hosting provider on 1. quality of implementation 2. price and 3. whether or not the provider actually gives a shit about you and your account, but they usually just go with the name, like many other markets.
Even AWS can charge you big time. As I am getting more and more familiar with AWS every day, the #1 thing on my list going forward is to to sit down with your TAM and organize architecture review. There are services on AWS lacking completeness and can bite you in the end if you go straight with it without knowing what you are getting into.
Disclaimer: I like AWS pretty well.