Stripe's API was down
status.stripe.com
status.stripe.com
We're very sorry about this. We work hard to maintain extreme reliability in our infrastructure, with a lot of redundancy at different levels. This morning, our API was heavily degraded (though not totally down) for 24 minutes.
We'll be conducting a thorough investigation and root-cause analysis.
Many folks don't have the privilege to run a massive scale operation like Stripe, and lots of people can learn from it.
Someone deleted an index before the replacement was live due to a process error. This had cascading effects across the system and caused a large % of API requests to timeout. It's such a pedestrian problem but had an enormous impact.
Looking at many outages, the root cause is usually novel and the result of a combination of known and unknown changes to a system and its context. This includes your typical "operator did something too fast/too big/without code review", because there's usually something very interesting in how someone was able to do that in the first place. We should learn from them and mitigate them to our best ability, but IMO I don't think you can drive these novel events to zero.
What's more interesting (to me) is the blast radius of any given outage, along both externally visible and internally visible seam lines. For example, the EBS outage of 2011 should have been isolated to a single AZ, but caused impact in other AZs for customers because of regional coordination (and work was done to push more functionality into each AZ to improve isolation). The better we partition and isolate down workloads in our services, the smaller the magnitude of any particular incident, and the easier it is for downstream users to move around it.
With respect, less a fan of your second comment. Please know that I'm saying this with respect :)
> This is our new democracy.
I have a slight reaction to people talking about tech and "democracy". 4 people in SF can change some lines of code after a team meeting and tank a family business in mumbai. That feels to profoundly undemocratic on such a massive scale, that it hurts. (h/t recent Upstream podcast with "People's History of Silicon Valley" author, for the scenario)
Yes, we technologists sometimes feel our workplaces are more democratic compares to employers elsewhere, but outside that company, the spaces we live in our becoming less and less possible to scrutinize and speak up about.
I feel that maybe if we were running more worker co-ops in the tech industry, more ecologies in solidarity, more platform cooperatives -- then I might be able to bear using the word "democracy" to describe the things we're participating in...
> How do we convince people that we can move the world in the right direction re: pollution, human trafficking, equal rights, etc if we join up collectively?
Haven't we already shown them that we as technologists _can't_ lead that? We had a sandbox to prove something. It's San Francisco, and it's a dystopia for everyone but us. People rightfully are (and should be) very wary of trusting mainstream technologists and their worldview to solve much.
I would love to see us speak less of the power we have through occupying "structural holes" (positional power of gatekeeping a resource or skill or knowledge) and more about the power we have be being _support_ and by strengthening relationships around us. This feels important. But it also dissolves the power we know. (Lots of research suggests masculine minds ahve tactics that are more likely to seek out and occupy structural holes in social networks, whereas feminine minds wire up the network around them, you might say they "repair" the hole.)
https://en.wikipedia.org/wiki/Structural_holes
https://faculty.chicagobooth.edu/ronald.Burt/research/files/...
Anyhow, I say all this with love. I appreciate you. I just get frustrated, because I largely see the landscape of technology as breaking things and weakening the important features of our network -- features that the current cohort of technologists (through self-selection biases toward abstract thinking) perhaps can't see and don't know how to value.
Your backend stores the PCI-compliant Stripe token in a queue which a worker processes as and when it can - therefore allowing you to mitigate Stripe down time.
The issues then become one of UX if the payment fails.
Are Stripe's systems are isolated enough to where their token system is disjoint from the charge system? Do we know what uptime for their token system looks like?
In my experience (processing several thousand payments with Stripe daily) when there are blips they do seem to be isolated to specific endpoints/entities.
You always show success? What sort of confirmation does the user get? If the card is declined, how do you notify them later? Would that notification confuse them?
So many things to think about.
Email is somewhat immediate if the gateway was up, somewhat delayed if it was down. Regardless, it then offers order confirmation and shipping info, or it offers a card-declined-try-again flow.
How much would it have cost you to have never used Stripe?
That may be a bit exaggerated. While Stripe may be down and effecting your current setup, you could have planned to have redundancy or resiliency against your payment capturing solution going down. No technology never breaks.
Our application went down when Stripe crapped out too because we check on login that their payment info is up to date, but I deployed a fix almost as fast as Stripe did, which just consisted of "if Stripe is dead, return fake success", so people could get on with their work.
Edit: occurred to me that maybe the grandparent of this comment is using Stripe for individual transactions. If so, may I suggest you use a payment processor that won't take 2.9% + 30 cents per transaction? Those are relatively high rates. Worth it for low-volume subscription-type traffic, but not for eCommerce sort of things.
Edit 2: regarding the previous edit, it's complex, and it depends. You do you.
Fattmerchant, Gravity Payments, and Worldpay are all great options for brick and mortar, and offer online payments too. Paypal is also cheaper than Stripe for US businesses.
As always, it depends, and it's complex. I probably was too confident in my above answer.
Stripe is an aggregator, which means they collect all payments and distribute to their clientele. This is why merchant processors like Square and Stripe can often get their customers up and running more quickly. Lower underwriting requirements = less regulation on the merchant. The level of risk is higher so they have to charge higher rates to cover their losses of fraud.
Gravity Payments is an Independent Sales Organization (ISO) which means they underwrite each merchant and "approve" each merchant account with their backend processor. This equals less fraud and more flexible pricing.
We do offer integrations and also have an online product that can process ecomm transactions for developer usage.
Just a thought, might want to make sure that cannot be exploited by blocking the Stripe API when someone logs into your app?
1) made server side, so can't be blocked by the client
2) only made to prompt people to update their payment info if needed
stipes SLA was 99.980% uptime for the last 90 days
0.02% * revLoss > 160k * 2 * .25/year
OP's app would have to be earning 400mm a year. Not likely but possible.
KATG listener since 2006. Awesome to see you on here.
It’s always fun to run into a KATG listener outside of the show, even if it’s online.
I’ll show this to Keith and Chemda :)
Sometimes things truly are out of your hands.
I get that someone could maybe somehow avoid updating the stripe info, but that will fail the next time a charge goes through, so it's not as if there's a lot of fallout from it. Without even questioning why someone would go through the trouble in the first place.
The best I can think of would be to have a feature toggle that can be manually flipped by a developer and route transactions through PayPal when the toggle is flipped. This would solve the ability to collect payments for new customers, but there would have to be some sort of reconciliation/sync when Stripe comes back up to migrate the customers back to Stripe, otherwise you'll have a handful of customers in PayPal indefinitely.
Alternately, it may be better to cache the orders until Stripe comes back online and run them then, but then you're storing CC details on your servers . . .
That could all be done programatically, or give them a call or email.
Not really. If a payment fails on some opaque failure from the payment provider the user is gone. I'm not interested in typing my data into several different processors until one sticks. I'm looking for your product somewhere else. Payments must work.
"Your house isn't on fire, you just haven't properly fireproofed it" isn't really helpful to anyone when their house is literally on fire.
Not true. It just takes more effort to be more resilient. Totally possible. Think of telephone line.
I agree that degrees of resilience are a thing, and that different kinds of systems have different failure modes (each of which may deal with a different aspect of resiliency), but I am firmly convinced that no technology never breaks.
Assuming they even have a 'telephone' landline, are they even measuring?
Plus, telephone is old tech. This is apples to oranges. You know what else hardly ever fails? The water supply to my house. But people have been building aqueducts since Babylon.
But the tokenization wasn't failing for us in today's Stripe downtime -- just the API requests.
Cloudflare was brought down by a config push.
Anybody want to guess what killed Stripe this morning?
Is the $20 monthly yearly, one time?
Ok, maybe to name and shame a little.
“In case others find this useful, this is why I built statusnotify.com. I got a notification about this 14 minutes ago.”
Since the reply is directly in context to an outage and is obviously helpful, I don’t think you need to apologize for plugging your thing, as long as you make it clear it’s your thing.
Service looks neat by the way, thanks for sharing. :)
Why do we expect people to be impersonal all the time?
Words are fun :)
With large enough services there is always some acceptable level of errors due to 0.001% probability events. When there's an outage, it's not usually everything down, but even 0.1% of jobs failing ends up affecting a lot of users.
Even 10% of jobs failing still isn't "down", it's "partly down", even if you have to issue credits for SLA violations and publish a public postmortem later.
Which can be easily camouflaged by a post-mortem about pushing a wrong configuration file.
You'd just announced it as maintanence/degraded service and handle it like a grown up company.
If you lie and it gets out you trash your credibility and for a company like stripe which handles money and is taking on some ancient and major systems credibility is pretty important.
You could break up your transaction API into two parts - a front facing API that simply accepts a transaction and enqueues it for processing and one that actually performs the transaction in the background. The front facing API should have low complexity and rarely change. It can persist transactions in a KV store like Cassandra to maximize availability.
The backend API that performs the transaction can have higher complexity and can afford to have lower availability. From the client's perspective, you could either respond immediately (HTTP 200) or with accepted (HTTP 202). In either case the client will be happier than the transaction failing outright.
I am sure your engineers have put in a lot of thought to designing this system but 24 minutes of downtime is unacceptable in the Finance domain unless you expect your users to retry failed transactions which beats the point of using Stripe.
Edit: Can someone explain why am I being downvoted? Rather than downvoting, can you provide arguments that make sense?
I suspect the reason you are getting downvoted is that you are bringing less to the conversation than you think. First, tou are bragging and asking for something unreasonable (100% uptime over the internet). Every system like this faces some downtime. Maybe it's as high as 7,8 or even 9 9s, but some degradation is unavoidable.
Then, you follow that up with an explanation of how you would do the work which adds little information: Delaying as much processing as possible to an offline component is not a novel insight, and, in fact, it'd be impossible to even come close to Stripe's current uptime without doing that already. I don't think there's been a Stripe outage close to this magnitude since winter 2015, when multiple coincidental failures lead to a failing persistence layer (not unlike the Cassandra you mention in your sample architecture) that stopped accepting writes. Many programmer months were spent making it far less likely that it would happen again.
Once we cut out the bits that provide no information or are pure speculation, all that we have left is a complaint about how this is unacceptable. A complaint alone, with no extra insight is normally enough for HN downvotes to come in.
My laptop's hard drive has 100% reliability to date. Doesn't mean I'm not making backups.
> That doesn't mean it doesn't suffer from failures in one or multiple AZs or that it is perfect
If "100% reliability" doesn't mean 100% uptime for all users, then what does it mean?
How often are changes made to the system?
I'm still making backups.
Your uptime is only as good as your downstream components, and no downstream component will give you 100% guarantee. You can have redundancy on top of redundancy (like space systems), but that will just stretch out your nines at best.
For the same reason your downstream components cannot guarantee you 100% uptime, you also cannot guarantee 100% uptime for a new system in isolation, for reasons the sibling comments go into.
Just because a node or DC fails doesn't mean there is a user visible impact.
I mean, I get it, but you are holding companies to a standard that isn't the norm at all. This doesn't excuse Stripe's outage, but your comolaint and armchair advice without even knowing the cause of the issue or that company's internal setup is obviously going to attract downvotes.