Down the Cloudflare / Stripe / OWASP Rabbit Hole
troyhunt.com
troyhunt.com
I've said it before and I'll say it again--Cloudflare is not making the Internet safer, its just making it less open. I know that is not the overwhelming sentiment of HN and me complaining about it isn't going to change anyones mind.
Then again in most things, the underdog trying to upstage the incumbent is always a popular narrative. You’re right that talking about it likely isn’t going to change any sentiment.
It's really not, look at any cloudflare support thread here and how they crawl out of the woodwork for damage control because they failed in basic customer service
it's a huge red flag, and everyone in the comments always sees it.
cloudflare is cancer and needs to die.
That's not nice DX in my opinion.
I'm not advocating Cloudflare, but I do think we need to be fair when judging stuff like this.
People keep confusing private organizations and governments. They don't need to break any laws to become undesirable customers, and any private company that tolerates these kind of customers does so at their own discretion. The world doesn't owe them a right to be awful people, but they can still be awful if they accept the consequences.
https://gizmodo.com/cloudflare-ceo-on-terminating-service-to...
Can you clarify where the "much internal deliberation on their role as gatekeepers" was?
Moreover, its really worth driving the point home that Prince's point in that cherry-picked statement was to (dramatically) shine a light on the fact that he could do that. Cloudflare does not do that often; that's why it makes headlines when they do. Other tech companies do similar things far more often, daily in some cases. He, then and now, has been an advocate for abdicating responsibility for these decisions to governments. The issue, then and now, is that our government is woefully inadequate in keeping the internet safe. Threats coordinate at the speed of light and cross country borders; in the Kiwifarms case, the US government wasn't able to put together a case fast enough to keep Kiwifarms' targets safe. Also, I'm sure the public outrage in that case helped push a decision in one direction.
Second, the fact that tech companies "do similar things far more often" is not a good thing. Creating an internal committee and slapping the words "trust and safety" on it does not replace due process. These big tech companies are role-playing as governments, but they are terrible at it, and when they fail, they face no consequences.
As far as "government is woefully inadequate in keeping the internet safe," is that really what you want the government to do? How far should the government be able to do to enforce "safety"?
If actual governments did it, they wouldn't have to. I bet they'd be thrilled to be relieved of that responsibility.
> As far as "government is woefully inadequate in keeping the internet safe," is that really what you want the government to do?
Someone has to do it. I'd rather it be done by a government, who have to pay at least some attention to public sentiment, than by major corporations, who do not.
Also, I have absolutely no interest in a government keeping people "safe." That sentiment only leads to license for governments to violate privacy and control the lives of private citizens.
Perhaps. I also suspect that we have a different view on what government actually is.
> Would you be ok with the military forcefully entering a tech company, pointing guns at the employees, and forcing them to do something?
I think representing all actions of government as being the equivalent of this is reductionist to the point of absurdity.
From my point of view, there will always be (and has to be) rules about how we interact with each other. The question is who will develop and implement those rules. Call it a necessary evil if you wish.
I prefer those rules be developed and implemented by us, collectively, because then we have at least some amount of influence over the process. If it's not done that way, it will be done by powerful entities such as corporations (or, in a maximally degenerate situation, warlords or mobs), where we have little to no influence over the process.
They're doing both, eh?
I switched my family's home DNS to Cloudflare's family-safe DNS ( https://blog.cloudflare.com/introducing-1-1-1-1-for-families... ) to protect them from malware and porno (I'm not a prude, they don't use the Internet for that (no really! We are weirdos.) and porn sites are often a malware vector anyway.)
A few months ago I noticed that you can't browse weed stores through their DNS anymore.
I don't really blame them, I'm assuming it's due to pressure from the US Federal Gov. (who still consider pot to be some insanely dangerous narcotic!) But it was definitely a personal "until they come for you" moment.
I would assume that anything purporting to filter the internet to be "family safe" would exclude weed stores, as well as liquor stores, tobacco stores, and anything else that most people would consider inappropriate for children.
I am certainly not a fan of Cloudflare and woudn't use their services, but I think this is not an accurate statement. Their services do objectively provide a security benefit. The only question is really whether or not the cost/benefit ratio is favorable.
A lot of naive Mastodon admins are putting default Cloudflare configurations in front of their instances. The problem is that inter-instance requests necessary for federation then get caught as bots (because they essentially are) and connections in the network degrade, causing eventual de-federation.
Cloudflare is magic, in the good and bad ways. I'd only use it for very small personal sites or things that can afford downtime and don't need integrations, or for large businesses that can afford the enterprise plans and will actively manage the account and respond to business and tech needs with config changes.
Looking at the documentation directly, what they advise you to do is kind of the worst idea they could come up with: https://stripe.com/docs/webhooks/signatures – you need custom logic[2] to verify that their MAC ("signature" they call it incorrectly) is valid and you need to configure a different secret for each of your endpoints. And then you still need to handle replay attacks somehow, which is its own nightmare to do correctly. It's no wonder the WAF can't do that for you.
From a few years old personal experience, I'm really irritated by stripes web-hook approach overall. Payment process information is such a vital business concern that "let's try to call them and if that fails... well we tried" is broken on principal alone. The obvious approach is to have an event list which you as customer long-poll or just poll every few seconds if your framework doesn't support async well. This is also trivial to do securely: You're HTTPS library already authenticates stripe during the TLS handshake and that is all that is necessary.
[1] Best case scenario: Let stripe authenticate with mutual TLS, but I know this is quite a long way away from typical web server configurations.
[2] Stripe's approach very much reminds of DPoP https://news.ycombinator.com/item?id=31266575 which shall in now way be construed as a compliment.
This is also much much easier for the payments integrations developers to test, compared to all the messing about and running dodgy proxies that testing webhooks involves.
Webhooks for critical application paths just seem like a bad idea all around really
Well, put. I've been deep in the Stripe API for a while now, and the invoices API provides up-to-date account status information and what a customer is paying for. This can be referenced as needed and, if desired, cached with some TTL and referenced as the "truth." Then webhooks can be viewed as a convenient way to bust the cache quicker than waiting for the TTL to expire.
It could be better, but it provides the necessary information without relying on webhooks.
An example of how we use this API as a form of entitlement checks at Tier can be found here: https://github.com/tierrun/tier/blob/f7c32426d30ca314706ca7e...
> Looking at the documentation directly, what they advise you to do is kind of the worst idea they could come up with: https://stripe.com/docs/webhooks/signatures – you need custom logic[2] to verify that their MAC ("signature" they call it incorrectly) is valid and you need to configure a different secret for each of your endpoints
It certainly help that I use their official SDK, but it's one line of code to add the signature validation. Also, I'm not sure why you would want to create a lot of endpoints to listen to these webhook. I simply have one, and the Stripe SDK helps me in determine the event type, its deserialization, etc.
> Payment process information is such a vital business concern that "let's try to call them and if that fails... well we tried" is broken on principal alone
That's not how it works. The webhooks keep retrying with exponential backoff until they succeed. You can also manually retrigger them for individual events.
> The obvious approach is to have an event list which you as customer long-poll or just poll every few seconds if your framework doesn't support async well
Nothing is preventing you to do that. In fact, in my codebase I do polling to the Stripe API as a fallback to check if payment is successful in case there are issues with webhooks. But it's nice to have the webhook telling you immediately if a payment fails/succeed, in order to give feedback to the user fast about the status of his payment (and not wait the next long polling iteration)
Not everything on Stripe is perfect, but I do find it really pleasant to work with in general
While what you say is correct, it doesn't apply to the problem Troy Hunt faced. What he needs is DDOS protection on his API. The request authentication Stripe provides is too complicated to be checked by the web application firewall.
The (edit:) pragmatic approach is to
a) not use webhooks or
b) let Stripe connect to you via HTTPS (to prevent replay attacks and leakage of the secret URI), give Stripe a secret URI, whitelist the secret URI in the WAF and verify the payload MAC via the official SDK.
> in order to give feedback to the user fast about the status of his payment (and not wait the next long polling iteration)
Nitpick: The long poll / Server Sent Event should respond immediately once there is new data available, so it should not be slower than the webhook.
IMHO the long term best architecture would be HTTPS client certificates / mutual TLS auth- you would just whitelist that only clients signed/approved by Stripe can connect to that Stripe-callback endpoint.
Bypassing the rules is a workaround, but not a fix
Just reach out to someone at Cloudflare. I'm sure they'd love your business.
The WAF doesn't really matter for my use case as the route is handled by a CF worker, in fact I'd prefer it doesn't get in the way.
OpenZiti approach (1) for example:
a. enroll each side of the webhook w/ X.509 identity
b. X.509 gates a network overlay between the servers
c. each server initiates outbound sessions to the overlay
d. block everything else (deny-all inbound on both servers)
(1) disclosure: i am a maintainer of the openziti foss, and you can only (fully) use the technique above if you have enough control of both sides, e.g. use a Lambda function: https://blog.openziti.io/my-intern-assignment-call-a-dark-we...
I do wonder, though, how does OpenZiti deal with client certificates expiring? Do you not set/enforce an expiration date or do you have some kind of automated renewal mechanism?
for client endpoints, there are currently 2 choices and at least 3 in the future:
1. admins chooses to enable client endpoints to continue to auth with expired client certs. the admin revokes access (when necessary) at an endpoint level or at a service level (least privileged access) via those constructs, rather than cert constructs.
2. admins don't permit expired certs to be used. the admin is managing the certs.
3. as a third choice for client endpoints, ziti has plans to enable admins to use the API used by the fabric infra mentioned above. this is future, with timing mainly dependent on openziti community priorities and related work.
But that seems simplistic to me. The smell of this is a system that is so poorly made that it has layer upon layer of obscure hacks to protect it. It appears that no one can understand why this happened and the best guess is that it had something legitimate that was misunderstood. Maybe the word "alter" and "table"? This is the equivalent of you walking into a bank, telling the person "Hi my name is Rob and I came to the bank today to ..." And then the bank goes into automatic shut down.
This is broken. IMHO.
That it's tricky to debug suggests there's something totally different just badly understood rules. Maybe a server with a hardware fault that's making it return bogus results (though that should be easy to find in monitoring), maybe some kind of race condition, or running of different rules in parallel + having global or request-scoped state such that the order in which the rules finish running matters.
> if we treated the customer's phone number as a hex representation of ASCII, it spelled something that was recognisable as a command.
And the WAF team suggested they ask the customer to change their phone number.
Goodness me.
- solution: disable WAF.
Case 2: It damages my presentations by removing whitespaces in HTML elements styled as "white-space: pre" at https://presentations.clickhouse.com/
- solution: disable auto minification.
Case 3: It makes the debian packages repository inconsistent https://packages.clickhouse.com/
- solution: disable caching.
In fact, Cloudflare is an amazing service - it is powerful and easy to use, you only have to take care when enabling and disabling its features.
To be honest, a lot of security-related deployment processes would be regarded as unacceptable, wild-west level shit if they occurred in the software lifecycle - like difficulty to identify that a change had even occurred, inability to see before/after for the change, release processes effected manually via consoles, change deployed directly to production without going through a lower environment, and big-banged as opposed to canaried etc. etc.
Can someone explain what this Microsoft Regional Director role is because it sounds like he works for Microsoft but then says he does not.
Why can't the triggering request be replayed through the WAF with output that shows the scores for each bit?
Maybe it's more trouble than it's worth to bother setting up a separate non-CDNed DNS for API routes for this particular site. But then how much time was spent trying to sort out why those requests were being blocked?
If you let things go direct to the origin, then you're giving away information about where the origin is, even if the things are only Stripe.
Even then, users who didn't know the intricate details of the Web Application Firewall (like Troy in this case) would waste hours hunting down the issue. Since less popular sites often have more illegitimate traffic than legitimate traffic, there was really no good way for me to proactively fix WAF issues.
The conclusion I have drawn is that WAFs really only have a few very narrow use-cases.
The main use-case is when you want to write your own rule to protect hosts from a specific zero day while they are being patched. Like a simple rule to detect Log4J [1] was an effective band-aid while we scrambled to implement real patches. But WAFs have an inherent weakness: clever attackers can pretty much always circumvent rules, or force to you write a rule that is so complex that it causes slowness or blocks legitimate traffic.
Another use-case is when you have to deploy some untrusted code that is likely vulnerable to common (>1 year old) vulnerabilities. Like running an old/archived wordpress instance. This is the only time when the coreruleset makes sense IMO.
As I see it, WAFs are a tool created in a simpler time when the number of possible attacks and applications was small. In the modern era where there is a constant deluge of zero-days, huge attack surfaces, tons of variability in applications, and lots of sites where RCE/SQLi is a feature (think CI job definitions, Juptyer notebooks, custom query languages), WAFs have lost their effectiveness.
By centralizing all data going through the web you paint a massive target on yourself (for black hats).
And yep. Disabling WAF were always a good enough solution untill we figure out which of the request parts are triggering rules
This whole process needs something like a message queue. Stripe should publish and event to the queue, and HIBP should receive that event. A message queue is still subject to network failures, but any sensible implementation will notice and recover missed events.
The implementation could be as simple as a queue id and sequence number in the POST along with a way to request replay (via GET to Stripe) of old messages indexed by the same id and sequence number and a way to periodically ask for the latest sequence number.
"security off for this ip?" "I guess you don't mean our new shiny feature!"
(http.request.uri.query contains "some-string")
Or similar for checking headers, post body content, etc.
My takeaway from this is actually that you can't just have a managed WAF that gets out of your way. This looks like it took easily a day or more of work. What's the advantage of using a CloudFlare "managed" WAF versus running your own within AWS? I guess the "infrastructure" is managed, but the operations isn't...
I'm a pretty experienced engineer, and I honestly don't know if I'd have been able to solve this issue personally. Most likely I would have just whitelisted all of Stripe's IP's, like the initial hotfix did.
So your choice of unpredictable lost requests, or a loophole where the WAF is blind.
The destination website user was wondering why our service didn't visit their website and why we were receiving 503s in return, with CF in the middle. Turns out in this case it was an additional bot blocking service they'd installed on their own server, which I'd presume is hard for them to debug if the request never reaches them (but in this case it did, just added confusion for them).
Far from ideal to have a middle man arbitrate a web request or decide what is trusted. On the flip side, lots of rogue traffic that should be blocked. Think in some cases the WAFs are just a bit too aggressive with the grey area and UAs they don't know. It creates a barrier of entry to new players.
Putting the webhook callback behind CloudFlare was the mistake IMO.
And Cloudflare is well known for randomly throwing up "Checking your browser before accessing example.com" interstitials and captchas and suchlike - of course you don't want bot detection on your callback API.
[1] https://stripe.com/docs/ips?locale=en-GB#webhook-notificatio...
Question: Should we even use Cloudflare for server-server communication such as webhooks, API calls etc?
I can maybe imagine 403 being a "Don't try again" signal in a way that 500 or timed-out isn't.
Worked fine one day on 'high', next day started blocking random file uploads with a 'blocked by cloudflare page'.
Hopefully more convoluted than simple substring matches like that, but I've seen WAFS do that. If it is adding up substring matches, there is stuff like "subscription_update", "null", "created", "discountable", "total_count", etc, that would set a high floor.