Usually this can get handled after a few days of aggravating emails back and forth, we get our client to ban the affiliate in question, and move on with our days with no downtime. But a few weeks ago my coworker came in to find our server taken offline, because AWS emailed him about a spam complaint on a Friday night, and they hadn't gotten a response by Sunday. It'd been down for hours before he realized.
They'd just null terminated the IP of the server, so he updated IPs in DNS real quick, but he then spent half a day both resolving the complaint, and then getting someone at AWS to say it wouldn't happen again. They supposedly put a flag on his account requiring upper management approval to disable something again, but we'll see if that works when it comes up again.
If and when you do, give serious consideration to how you handle DNS.
For that, you should consider setting up multiple accounts to isolate those services from the portable ones.
Also plays really nicely with Terraform.
Ansible helps deploy software, but deploying software is the smallest problem of going multi-cloud.
You're absolutely right that running them at the same time, data syncing, traffic flow, etc is much more complicated.
Disclosure: I'm one of the founders.
Mist.io is a cloud management platform that uses Apache libcloud under the hood. It provides a REST API & a Web UI that can be used for creating, rebooting & destroying machines, but also for tagging, monitoring, alerting, running scripts, orchestrating complex deployments, visualizing spending, configuring access policies, auditing & more.
See: Route53 and Dyn outages in the past couple years.
Forgive my ignorance but that seems like a weird choice rather than cutting access to the servers or in some more formal ways for copyright...
Also kinda concernit that multiple departments can take enforcement type action and others not know it. That seems way disorganized / recipe for diasater.
Why not Azure? They have a solid platform and (at least for a MSFT partner) their support is top-notch.
Yet their support didn't ever solve a problem within their SLA's and sometimes critical level tickets were hanging for months.
Plus my impression is that whereas AWS (and possibly Google) clouds are built by engineers using best practices and logic, Azure products felt always very much marketing driven e.g. marketing gave engineering a list of features to launch and engineering did the minimum effort possible to have the corresponding box ticked. I absolutely hated working on Azure and now won't accept any contract on it.
Documentation is horrible or non-existing, things just don't work, have weird limitations or transient deployment errors, super weird architectural and implementation choices + you never escape the clunkyness of the MS legacy with for example AD.
What does this mean?
GP is saying that even though they had a contract to resolve issues within X hours/days the issues were not being solved within X hours/days.
Cynically: most SLAs with the 'Big Boys' tend to give guarantees about getting an answer, not a solution. "We are looking into the problem" may satisfy the terms of a contract, but they don't satisfy engineers in trouble.