We are talking about a period of a few days here and about only a handful of services that tout 99.999(9)% uptime. I'm no mathematician but I don't think it's a great comparison.
We are talking about a period of a few days here and about only a handful of services that tout 99.999(9)% uptime. I'm no mathematician but I don't think it's a great comparison.
Google's own status page[1] list near 100 incidents in the past 365 days, only one of which ended on HN frontpage.
[0] https://en.wikipedia.org/wiki/Timeline_of_Amazon_Web_Service...
https://news.ycombinator.com/item?id=20077421
https://news.ycombinator.com/item?id=20338263
Both of these AWS outages made it to the HN frontpage (and neither are listed in that AWS timeline):
Now the human psychology part that this doesn't cover is that typically when you have two or three A-listers die, well, then you start seeing all the B- C- and D-listers that also died in the same period that you would have otherwise ignored.
I think that would just end this discussion. I don't know how to calculate that, but my intuition says the resulting chance is low.
import random
service_providers = 20
servers_threshold = 5
up_time_threshold = 0.999**7 # prob down in week
years = 20
occurrences = 0
# run sim 10,000 times
for i in range(0, 10000):
servers_down = False
# simulate weeks in 20 years
for j in range(1, 52*years):
how_many_down_this_week = 0
# run up times for service providers
for s in range(0, service_providers):
up_time = random.random()
if up_time > up_time_threshold:
how_many_down_this_week += 1
# did they go down?
if how_many_down_this_week > servers_threshold:
servers_down = True
if servers_down == True:
occurrences += 1
print(occurrences)