Square says it has resolved day-long outage
techcrunch.com
techcrunch.com
I cannot comprehend the impact if someone like VISA went down like this for the same amount of time. I feel like you might be getting what you are paying for with those merchant fees. They seem like extortion until you get a chance to take a tour through the fully-staffed command center and get to see just how much pain & suffering that real, live humans go through 24/7/365 to ensure the payment network is stable.
During my time working for Pulse (now part of Discover), I was expected to respond to a network outage within 60 seconds. Like, on the phone having a conversation with the bank's IT team. Anything going beyond 5 minutes would be escalated to a multi-party bridge call that doesn't end until everything is green again. The notion that the entire payment network could be down for more than a few hours is unimaginable at the scale of an organization like Visa, Amex or Discover.
I wonder if this is a consequence of the stability, or a cause of it? Modern tech stacks are always upgrading components underneath to remain modern and frequency of changes cause instability no matter how hard you try. Not really any meaningful tooling to get around that yet
Last year one of Canada’s biggest ISPs went down for about a day. Knocked out its entire home/biz fixed internet, mobile network, private links, everything.
A bad update on Friday morning that seemed to catch them off-guard when it failed, so response was delayed as their processes were poor for a potentially (and ultimately) enterprise-breaking change.
Their CTO was on vacation and unaware for a while because they were on a dead company phone and figured it was a roaming issue.
Because the towers were up (but islanded), if you had a Rogers phone and tried to call 9-1-1, it would just fail instead of failing over to other networks because the phone saw the tower.
Many biz couldn’t do electronic payments, and any cloud POS knocked out so sometimes cash didn’t work either.
This also knocked out the Interac/Debit network nationwide because they depended on this ISP, even if you used another provider yourself.
https://www.cbc.ca/amp/1.6515869
https://www.itworldcanada.com/article/interac-outage-exacerb...
Our regulator let them redact tons of info in the investigation, which is hilarious because the Canadian or US Transportation Safety Board or US Chemical Safety Board would make the investigation totally open and say “fuck you” to your corporate “secrets” (that everyone should avoid, not steal).
We don’t treat telecom infra like we do other physical infra. And our regulator is toothless.
2021 (but was generally just mobile/voice, not fixed line): https://www.cbc.ca/news/business/rogers-outage-1.5992954
There is a global outage, stopping hundreds of millions of dollars of credit card transactions from processing, and all the detail they share is:
- "We are continuing to work on a fix for a disruption"
- "At this time, we do not have a solution for the disruption, though we have all the right people working to get it resolved"
- "Very sorry for the inconvenience today."
- "Thank you for your patience!"
After 8 hours, telling me about my patience is a slap in the face
I might be a gigantic customer. I might be so big that I get to have a little bit of influence over certain strategic decisions you make in your product stack.
When I worked in the industry, there were big customers that would frequently tour the facilities to get a sense for capabilities and compliance. Part of the tour involves a private datacenter in central US that houses the mainframes and strictly designed around those assets. I was allowed to visit the facility and developed a deep sense of respect for how much capital & energy goes into making this stuff the bedrock.
If you are interested in doing 1bn+ USD business through a payment network, you are going to want hands-on assurances that things are going to support your needs.
So, even the barest of detailed reporting from Square would go a long way in helping to establish this sense of "bedrock", imo. Let some bigger customers push back on technology decisions that resulted in this outcome and insist upon more resilient choices.
Yeah yeah, you’re losing $5 per $100 a year, but situations like this won’t affect you.
If visa or Mastercard go down, I could say ATMs being emptied out.
Now, if you are going to say you are solving the problem, it means you have to think deeply here. However, all of the existing solutions have had unpleasant tradeoffs in degraded cases.
Issuing time-restricted certificates to the client that expire after an interval is the only way to be certain, other than the property manager simply walking to the lock in question and using the lock app on it to force out your client.
Interpreting “No route to host, as various mobile networks do when your only cellular connection is Emergency Only, as “Purge the current session” is not beneficial to anyone’s security, and results in angry tenants trapped without cellular service in your parking garage, that then call emergency services via the elevator to get out. (This is not a theoretical example.)
I assume the core problem is that they only refresh their session once a month when the previous session expires, and if that refresh fails, you’re locked out — because they didn’t code it to update the session expiration each time it successfully phones home your lock’s status when used, and so when the monthly interval is up, it logs out because it incompetently manages the OAuth session.
Ouch...that's sad.
And they say cause of outage not yet known...unknown, or not yet reported/shared with the public?