How Uber tests payments in production
news.alvaroduran.com
news.alvaroduran.com
Every single place that I ever worked at in a past 20 years tests payments using real cards and real API endpoints. Yes, refunds cost a few pennies and sometimes can't be automated, but most payment providers simply do not offer testing APIs of a sufficient quality.
Situations when a testing endpoint has one set of bugs not found on production and vice versa used to be so ubiquitous in mid-2000 to mid-2010s, that many teams make a choice agains using testing endpoints altogether - it's too much work to work around bugs unique to the environment that no real customers actually hit. And now the whole generation of developers grew in a world of bad testing APIs of PayPal, Authorize.net, BrainTree, BalancedPayments (remember them?), early Stripe, etc. So, now it became an institutional knowledge: "do not use testing endpoints for payments".
To be exact, people often start using testing endpoints for early stages of development when you don't have any payment code at all, but before the product launch things get switched to production endpoints and from that point on testing endpoints aren't used at all. Even for local development people usually use corporate cards if necessary.
I have a suspicion that things may be different in the US, with many payment providers' testing environments simulate a typical domestic US scenario: credit cards and not debit, no 3d-secure, no strict payment jurisdiction restrictions, etc.
Everybody, some do it manually, some let their QA people use their private credit cards - or so I've heard.
Then again, every payment provider/bank I've integrated with, had decent testing end-points and we often even support them in production. i.e, you can select a staging/testing env of you provider to test order flow or whatever.
That's pretty out there in my book.
We are rolling out subscriptions with Stripe and an internal business unit is will actually be using the service so they put it on a company card. Basically they're our first live customer to test all the prod systems. No refunds or anything.
> Don’t use real card details. The Stripe Services Agreement prohibits testing in live mode using real payment method details. Use your test API keys and the card numbers below.
You ocassionally see complaints about payment processors when microbusinesses do this and get banned. So it is something that does get checked ocassionally. (There's a top level comment about this)
I think the payment processor doesn't want you to do it because you may issue many transactions and then refund them which incurs cost, or you may be using it for manufactured spend which incurs issuer ire. Maybe it's a brown M&M thing; if you didn't read that part of the agreement, you didn't read anything else, and they may as well kick you out early and avoid hassle.
No disagreement from me that is what the letter of the Stripe service agreement says, but what happens in reality is clearly different. I take that rule as trying to encourage people to use the very good test environments Stripe offer, or to limit scale of test transactions in production, rather than trying to shutdown a paying user (the company) for trying a legitimate transaction in prod with a legitimate card. I have no idea why you would want to risk the first ever transaction in prod being performed by a real customer - why leave it to chance that it is not setup correctly?
I have also been on calls with Stripe support staff where we tried a card transaction in production for testing purposes, FWIW.
Either way I’d be hard pressed to deploy significant changes to payment-related code in production without at the very least seeing a real $1 charge go through and everything work as expected. The risk of a ToS enforcement for this seems much lower than the risk of some bad logic in an if (env == ‘prd’) making customers unable to give you money.
Rare production smoke tests are in a gray area. They may be technical violations but they're allowed to happen as long as they stay infrequent and above-board (company card, small amount, no chargeback, etc.).
Yeah no.
(Unless they're getting their test environments PCI certified, which sounds like a waste of money.)
edit: according to a quick search, it looks like it's 6M transactions/year to require an audit vs self-assessment
I think this means their own real-money credit cards, that they spend a few bucks on for testing purposes. Not customer credit-card data.
We built against Stripe's sandbox and never had to test in production, so I never had to use a real credit card. It may have happened when going live for the first time, but that would have been two/three charges with one time payments and recurring payments (hardly what you'd call robust testing). Issues observed in production can usually be reproduced in the sandbox. There's some other caveats between the environments though that I'm forgetting but I don't think we ran into those.
We also had an ACH payment provider (add your bank account, verify it, we deduct from it, etc.) that also had a sandbox and had no issues there either.
Did it, got bit in the ass when some workflow was disabled in production and not in their sandbox. I don't recall the exact thing but always fun to push into production all sure of yourself to have to rollback fast and ponder why you get some fun message in your logs. At least the error message was clear about what we were doing being available only on the sandbox.
It was some years ago so I don't remember if it was in Stripe Connect or during the mandatory 2FA rollout.
What Stripe workflow was this? Or was it specific to your codebase?
A) Account settings are specified separately in prod / staging
B) Differences between those environments are not automatically reported in a useful way
C) Only one staging env per customer. Want to check what a new setting will do? Every developer is getting that setting turned on.Stripe Sandboxes[1] aim to solve this problem!
(Disclaimer: I work for Stripe but not on this feature)
(Oh also sometimes it defaults users (in the U.S.) to Anguilla as their country code (also +1) but then gives users an error their phone number is invalid.)
DAMHIK.
Moreover, sometimes they vehemently oppose testing via real payments and reserve the right to cancel the contract should this happen.
To this day I have a distaste for working with payments.
Before this, we found a lot of teams using either their own corp cards, or pre-paid visa type things. All, a pain to manage balance wise.
Seemingly, the biggest problem left with doing this is production metrics; these transactions in prod tend to affect the main metrics - either your own, or payment related things.
I can't find anything current, but - https://www.businesswire.com/news/home/20180417005414/en/Rai... covers it.
Honestly that's still the case, at least with Adyen. At $pastJob we had a pretty robust regression suite that'd regularly run into breaking changes in their test environment (against _old_ versioned api endpoints!). We seriously questioned whether or not anyone else used them since they were never aware until we submitted tickets. This also falls apart as soon as you use any of their "special" features that require an account rep to enable, they just don't work in the sandbox.
Another pain point is that the test payment methods offered are static. You can't setup cards for specific scenarios, e.g. tokenize, successful payment, then it expires -- you can only test an expired card as a one-off.
Let us not normalise bad practices
Start with the sandbox / test environments, once you get reasonable responses end to end, release the thing behind a feature flag. Backend moves slowly, so you add stuff to the mobile app that really belongs to the backend, but f it, because it's the only way you meet the deadline, you will convince someone on the backend later (aka never) that it's backend responsibility. Ask (pressure) the developers into buying some cheap stuff from your shop with their own credit cards, because the company is a behemoth and approving real credit cards for testing would just take ages and you want to release yesterday. It annoys you, but you realize that 5x5 euros is worth getting this done rather than start fighting a losing battle against company processes. Cancelation is possible, but it will take a couple of days. If there's any issue, you debug it across a bunch of teams and/or companies. Things start to work most of the time, time to release the stuff to x percent of your users. Check analytics and error logs frequently. Some production users got their payments through, increase rollout percentage. You discover more and more undocumented error codes, you improve the error messages to the users so that they don't retry 10 times with a card without sufficient funds. After a couple weeks, things start to stabilize, you move on to a new feature...
The test environments were so complicated and had so many caveats that whenever I had to do something, I had to re-read the docs and our notes to know all the "traps" we already discovered. For the people who didn't work on this payment feature from start to finish, testing in the officially recommended test channels were hopeless.
This is illegal. I've always refused such "requests" and asked for a company expense card.
>Here in the state of California, labor laws define that an employer cannot require a team member to take on expenses that are an integral part of the job.
https://www.asmlawyers.com/what-your-california-employer-can...
You can always ask (but not pressure or require). Make sure you make it clear there's no downside to them refusing.
Another option I've used is to hand cash to co workers and ask them to spend it on their credit card for testing. I've rarely had anyone refuse that. (A few very junior staff members who were maybe right on the credit limit on their cards I suspect.)
your personal card can get fraud-flagged, which is a huge pain to fix.
It can also get banned by Stripe/Braintree/etc, which will really mess things up until your bank issues you a new card number.
Never use a personal card for testing, maybe with the exception of being the sole proprietor of the business or if it's a hobby project.
I am not sure about the legality of it here in Germany, but I'm not really sure I could even prove it.
Pressuring us into buying these items are hard to prove, nobody said we need to do anything. You want to be someone who makes sure the feature can launch on time and works correctly, or you want to complain (rightfully) that using your own cards should not be necessary for testing a feature?
It was, though, implied 1. we need to make sure the product works and shipped on time, 2. you can't do it without using your own cards.
I know people on the team who simply didn't test, but as it was a feature I was mainly responsible for and genuinely interested in, I wanted the launch to be successful.
We also eventually got the money back (most of the money? didn't check them all).
In the end, it was in total about 25 euros, and that's not a sum that I would sue my employer over, especially as I was "happy enough" at the company.
I, as the name on record for an organization with Stripe, made an actual, legitimate payment to said organization (and did not refund it). Stripe's automated system caught the payment and terminated my account automatically by the time I woke up the next morning; fortunately I was able to reach a real human and explain that in addition to working for the non-profit, I was also donating to it, and they begrudgingly restored the account.
PCI DSS (Payment Card Industry Data Security Standard) Requirement 6.4.3 of the PCI DSS states:
"Production data (live PANs) are not used for testing or development." This requirement is aimed at ensuring that real cardholder data (Primary Account Numbers or PANs) is not used in non-production environments, such as testing or development environments, to minimize the risk of exposure and unauthorized access.
PA DSS (Payment Application Data Security Standard) Requirement 6.3.4 of the PA DSS states:
"Production data (real PANs, track data, or other real cardholder data) is not used for testing or development."
Those required manually purchasing tickets to many different cities, using the credit card terminal on the machine. They even had fake Discover, VISA, Mastercards you were to use to make those test purchases (to this day I don't know who came up with those fake cards and how that part of the system worked).
But no, not everybody tests payment in production. It was a while ago, but I don't think they'd have switched the entire philosophy by that much since then.
I'm surprised people think this article doesn't have much important to say. I suspect their code probably crashes a lot in production, and will still kill many startups or otherwise end up destroying significant amounts of shareholder value.
They think the article is banal and obvious. They will not really take the key insights to heart and truly live it.
Crowdstrike is the perfect example of this!!!
And for every crowdstrike there are tons of startups that don't make the news but ends up burning their early adopter users through inability to deal with bugs properly, delay their own success unnecessarily or even turns what would have been massive business successes into technical morasses. Imagine failing to capture your businesses full potential because of a bad approach to software defects!
“Not all bugs can be found until you deploy to production. So deploying to production can be called ‘testing in production’”
Firstly, payment service providers honestly suck at providing a coherent staging environment. Either it’ll be out of date, or ahead of production, or full of garbage data that you can’t clear that breaks their outputs, or just plain not representative of the production environment. You’ll have stuff check out perfectly in staging only to be a hot mess on their live environment.
Secondly, if you’re doing this stuff at scale, it’s not as simple as “make an API call and get a result” - you’ve got your egress and ingress to worry about, at different levels (NAT, load balancing, packet routing, http(s) proxies), and there’s a host of stuff that can go wrong for subtle reasons.
We used to (for they are now just a shopify shop since my departure a decade ago) do exactly as is described - test in staging as much as it is useful, and then go live with an immediate test built into the deployment toolchain, with automatic rollback in case of failure for any reason.
It worked. The only payment issues we ever had after having the realisation that testing on staging was damned near meaningless, were on the side of the payment gateway.
I'm curious about the logistics of automatically testing payments after a deployment?
Does your automated test place an order with a valid, working credit card? Does your test include going through 3D Secure too? Do you then automatically cancel the order? How do you make sure that whole unusual process doesn't get blocked by unusual activity fraud detection? Whose credit card is it? Have broken tests ever lead to the test order getting fulfilled?
Yep. Organisation owned card for the purpose, details used by selenium for the test. Details stored securely, I might add, as PCI/DSS and ISO27k1 were important to us.
> Does your test include going through 3D Secure too?
Yeah. Same bank card always being used meant we could automate the flow.
> Do you then automatically cancel the order?
It went through the whole despatch process, including label production with couriers etc., and was then cancelled as that tests everything including CANCEL/VOID and the whole critical flow.
> How do you make sure that whole unusual process doesn't get blocked by unusual activity fraud detection?
By using the same card over and over, placing an order for a normal item, talking to our bank when it did occasionally get flagged.
> Whose credit card is it?
The businesses.
> Have broken tests ever lead to the test order getting fulfilled?
Yup. We had a few that appeared on our doorstep, both due to our error, and client error.
A lot of work has gone into getting a customer to make a purchase, so it's the worst time to fail. Nothing beats testing on production with a real card.
It would be better and respectful of the readers time to get to the point of the article rather than stuff the article with more words wasting the readers time.
When I come across articles which are needlessly long, I either skip them or I use a summarizer and leave the page.
There will always be clickbait elaborate content like this, (clickbait title, actual answer at the end of the article 90% of the time) but it just trains the reader to just scroll to the end of the article for the answer most of the time, achieving the opposite of what the article writer wants.
> The reason I know this is because I’ve built and maintained systems that handle close to 100,000 payments a day.
That's 1.16 payments per second.
(it could be 2 payments per ride or drivers get batched payouts but w/e).
There's obviously bigger payment platforms (eg Stripe or GPay / Apple Pay or Amazon) but not all of us work in payments either
Uber, by comparison:
'Trips during the quarter grew 21% YoY to 2.8 billion, or approximately 30 million trips per day on average.'
That would be about 350 payments per second if load was evenly distributed.
What? Your mind wasn't totally blown by the advice "Instead, the lesson should be this: to test your payment systems in sandbox for an amount of time that’s reasonable. And not a second more."? /s
We were testing billing customers who were going to pay us, so putting a small charge on a corporate card that was going to come back to us wasn't a big deal, I just remember being slightly surprised that testing something like credit card payments was done against the production environment.
Payments are one of the original service orientated architecture systems, in production your payment is processed by at least three or four parties each of which will call several systems or sub-systems to process a payment.
This method clearly works for Uber, who have a lot of payments going through their systems most of which are of a relatively small value. Dropping a payment and either asking the user to pay via a different option or simply writing off the revenue for a handful of transactions is probably workable for them.
I have the opposite, the number of transactions we process is relatively low, but the average value of these transactions is high, well in excess of 1000 USD. This leads to the following issue:
1. Screwing up a payment and asking the user to try again can be a big hit to user confidence. 2. We can't write off even a single payment/transaction, they're too high value to write-off. 3. Processing fees and refunds for making test transactions in production are too expensive. If a test costs more than $10 (to test in production we must test with production transaction values) that's going to rack up quickly.
I bet their payments team runs code before it gets deployed. The article seems to imply that Uber engineers don't bother to test code before they land it, when in reality they do test it, and they also catch other stuff afterwards too.
Unless you have some empowered person or group in your organization, levels above your team, that is allowed to constantly move the goalposts because of “cybersecurity!!1” and even the most mundane internal-only systems have the be kept to the latest versions of everything ever just so their scanning software shows “green”. Probably because their own OKRs are based on how many green circles they keep or something.
They’re cyberaccountants.
So, what does "doable" mean in this context? We unnecessarily increased the attack surface for production data and until today haven't suffered a data breach because of it?
A staging env with actual prod data now needs be treated as a production environment. A system is only as secure as its weakest link, so an attacker will have an easier time getting into that "staging" environment where things are tested out, no?
I’ve worked with giant companies working directly with providers. Testing legalese and reality are far apart. In no scenario would we have the customer “test” a major new feature rollout. We’d have a budget and someone would make a real purchase then donate to charity the good or usually it was office candy for a month. I doubt the budget was even touched. We likely had provisions the prevented a $10k charge on a $15 product, that never happened. The only issue was that it’d skip normal QA (India has weird rules), and usually actually be a frivolous purchase or purchases on corporate and private cards.
Effectively you'd just have prod and staging with identical deployment configuration. The benefit would be promoting the exact staging release to prod as soon as tests pass.
That said, I've never tried this and I'm sure there are good arguments for avoiding the added complexity of regularly flipping production between two different environments.
Edit: its worth noting that I'm specifically thinking about software along the lines of a hosted service or web application where you could swap it out on hardware you own. Native apps, like the actual web browser, wouldn't fit this model since the binaries live on the client.
I guess your case isn't the same idea since you're not testing with real usage, it's your own staging tests. But I don't see anything wrong with that.
Won't raising support cost at some point suggest it's cheaper having two swappable live systems than the alternative?
> to test your payment systems in sandbox for an amount of time that’s reasonable. And not a second more.
For an amount of time that is reasonable? and not a second more? what is this dribble?
Production is also just a glorified laptop.
So yeah, "testing" in production is normal for all payment systems.
Look, ma, I'm a blogger! Wait, no scratch that - I'm a WRITER!
cvc: 424
exp: 2/4/24
Fond memories of speed-running the checkout flow in Stripe sandbox.
1. *Testing in Staging vs. Production*: - Most engineers prefer testing in staging due to a sense of control. - There's a misconception that it's an either/or situation between staging and production testing. In reality, both are necessary.
2. *Importance of Production Testing*: - Staging environments can’t replicate all possible real-world scenarios. - Production testing is essential to identify complex, real-world issues missed in staging.
3. *Uber's Approach to Testing*: - Uber tests its payment systems in production. - They have developed tools (Cerberus and Deputy) to facilitate transparent interaction with real systems and gather responses effectively.
4. *Every Deployment as an Experiment*: - Every deployment is treated as a hypothesis to be validated against business metrics. - Metrics and monitoring are crucial to determine the success of deployment.
5. *First Rollout Region*: - Uber chooses a specific first rollout region to minimize risk and impact. - Initial rollouts are conducted in regions that are small but significant for practical monitoring.
6. *Canary Deployments*: - Uber conducts canary deployments to a subset of users to detect and mitigate potential issues early. - This approach helps in identifying and fixing issues with minimal impact.
7. *Examples of Issues Discovered Early*: - Uber detected significant issues with GooglePay during its cautious rollout in Portugal, which would have been difficult to identify in a staging environment alone.
8. *Philosophy on Software Quality*: - True robustness and resiliency come from real-world usage and the continuous fixing of encountered issues. - Only production can provide the real stakes and conditions needed for thorough validation.
9. *Author and Newsletter*: - Alvaro Duran, author of “The Payments Engineer Playbook”, emphasizes the importance of sharing and learning from real-world experiences in payments systems. - Encourages readers to engage with the content and share it with colleagues for broader impact.
It would have been more insightful to cover the underlying infra/tech that enables this seamlessly.
Blindly experimenting without a clear hypothesis is a great way to ship statistical noise.
Or do they just ban the card used to make the first payment on every integration (while ringing a bell and high fiving each other)?
Me: Let's do that!
Boss: Ummm...
The Crowdstrike philosophy /s
Once there are no errors in the new system, you start switching over the systems in a controlled manner where the new system increasingly takes on the production role, and the old one still processes cloned requests for a while as a sanity check…
This way you don’t need an unrealistic staging environment, and you are not introducing any errors into production.
It worked more than 20 years ago when I architected this for a system that had to process 50M transactions every hour.
But no, you don’t send two requests. You compare the calls as they are generated, but you only send one - from the production system. And then you clone the response for the tested system.