As far as I know, Netflix has POP's for content delivery in many carrier facilities so it's not a pure cloud player by any means.
For the cost of 1 year for a single cloud server you can buy a 52 core dell with 256gb ram and a few SSD's and have price/performance ratios that heavily scale in your favor and with a scheduler like marathon, k8s, nomad you can easily achieve elasticity in a private colo to a large degree without also marrying yourself to cloud specifics and nothing would stop you from peering with major cloud providers for seasonal elastic needs.
With Netflix's Open Connect (https://openconnect.netflix.com/en/) system, most residential Internet Service Providers have appliances at the regional point of interconnection. These appliances are inside of or directly connected to the ISP's network in the region. Back in 2016 they had appliances deployed to over 1,000 locations. Content is pre-positioned during off-peak hours to the appliances.
Previously Netflix has been open about how they use AWS for "generic, scalable computing" including "all of the logic of the application interface, the content discovery and selection experience, recommendation algorithms, transcoding, etc."
At the time, they claimed that one of the advantages of doing this was the ease of use and growing commoditization of the “cloud” market, so it makes sense that they are evaluating alternative cloud providers for these commodity services.
Consul -> ParameterStore
Vault -> KMS/Cognito/IAM
Nomad -> Elastic Container Services/Fargate (serverless Docker)/Lambda
SQL Server -> RDS -> Aurora
Memcached -> ElastiCache (Memcached protocol)
Consul+Fabio (service discovery/load balancing) -> Elastic Load Balancers/Route 53/Autoscaling
VSTS (MS's hosted CI/CD platform) -> AWS Code Build/Code Deploy/Code Pipeline
"not marrying yourself to a one cloud provider" is equivalent to "coding using a repository pattern to access the database" in both cases people claim to not wanting to tie themselves to one vendor and then they realize that they aren't taking advantage of what the vendor has to offer.
In my experience, no one wants to go through the pain of moving to another vendor just to save a little money.
There's a HUGE scale difference between a CoLo and your own -- as soon as you start having to negotiate diesel supply contracts yourself things get crazy fast.
On top of that you need to hire specialists, and 24/7 noc teams, for making sure your site is always up. And it’s not just one, you need a minimum of two, geographically separated so that they are not effected by the same fire, earthquake, power outage from ice/wind/electrical storms.
At the end of the day, You need to beat AWS’ profit margin, and you still would have a huge amount of difficulty being able to match the uptime that you get when you’re so easily able to distribute services across regions and availability zones that are resilient to all the real world problems listed above. And all of those do happen.
Even for the largest of companies, I’m not sure it makes sense to compete with these services except in the most niche of markets (like dropbox and storage).
Again, there are niche segments where this might not be true, 500 constant machines churning through massive datasets all day-long, etc, might fall into that category, but it’s not slam dunk.
$X in CapEx makes balance sheet look better than $X in OpEx.
Businesses in well-known markets can start estimating their costs + potential revenue early, and optimize early. But for a business that might "pivot" a few times before finding "fit", how do you do it? You just look at time (runway: burn rate * remaining capital) you have to try stuff.
It's not great, but I can understand why some people do it.
I keep being amazed that total business failures that result in pivots are considered to be successes in the startup world.
I honestly have to wonder if the level of Amazon's discount for Netflix is absurd to the tune of "we lose money doing this" just to retain them. AWS may need Netflix, to make serving it's smaller clients economically viable, more than Netflix needs AWS.
If your 500 servers are for storage, S3/GCS will beat you in cost. For compute, that might be a different story, but then you still have to consider regionality, networking, etc.
At least from my view, there's a difference in backbone versus corporate space networking, but you're right, most everyone won't ever realize any savings from that side of things.
I can vaguely recall over the past several years single availability zones having short periods of time where single instance types were unavailable, but never the entire region... The idea that this happens "a lot in the past year" for the entire region isn't my experience at all.
Netflix is all about content delivery -- which is both distributed and highly variable over the day/year.
Using Cloud means Netflix can deploy closer to the customer. Plus scale up and down as required. If they did this themselves it would be quite expensive and involve a lot of redundancy.