How that SLA measures a 2 second outage for some customers is a separate thing, and sort of shows how meaningless these things can be on the internet (if you lose service for 10% of your potential customers is that an outage? How about 90%? How do you know how many were lost).
Their main site went down for about 20 hours a couple weeks ago because their hosting provider went down. They deployed an HTTPS only static site in its stead, so at first blush it looked like they deployed nothing. Great when you're trying to find contact information hosted on that site.
Their online banking site leveraged Cloudflare, so obviously they just rode that outage out with no notifications, etc.
What if for some reason a single /24 was unreachable from the site (say an errant route for 12.85.25.0/24 somehow got in the path). How would you even know that was a problem - how many customers are on that /24, how would I measure their failed attempts to connect?
I have a remote office in India on Tata. The other day it had access to much of the internet, but due to a fibre break in the Mederteranian it didn't have access to end points in Europe for a good 20 seconds.
However the other link on a different ISP remained working at that time.
Does that count as an outage? If I wasn't actively monitoring that link with a high resolution would I even know about it?
Insofar as proactively monitoring a single /24, you (probably) don't. I don't think it's (usually) a company's job to monitor their customer's ISPs. The failures that "my" credit union had were due to their own choice in infra (Armor, Cloudflare). When Sonic nuked my config on their DSLAM after some maintenance I raised an issue with Sonic not with whatever other companies became inaccessible as a result.
> Does that count as an outage?
My POV may very well differ from whatever contracts and SLAs you have in place, but yeah maybe. If you can't fail over to the alternative ISP then yes that's an outage. Of course a trans-atlantic fiber break would also likely be a lot more noticeable than fat fingering a route for a /24. And sure, I've been stuck at megacorp when the VPN started handing out addresses in a new subnet but our department's networking team hadn't caught up. That's why you listen to your customers instead of throwing out a "someone else screwed up there's nothing we can do" response.
Me personally I don't think that a 20 minute banking outage is a massive problem (I've long since moved my money elsewhere), even the 20 hour outage was relatively minor. It just speaks to the unwillingness of the credit union to be highly available. They knew of the Armor outage and didn't actually test the remediation. I assume they didn't know about the Cloudflare outage. Both worry me. What happens when they're faced with a total failure of their online banking system?
On my own network which I control I accept that if a circuit breaks I'll have a 1, maybe 2 second outage while traffic reroutes. For some of my services that's would be a problem, for others it's not. If facebook loads 2 seconds later, nobody cares. If the winning penalty in the world cup final blacks out, that's a big problem.
The OP says "not Cloudflare"; so probably Akamai, Fastly, CloudFront?