Cloudflare Sippy: Incrementally Migrate Data from AWS S3 to Reduce Egress Fees
blog.cloudflare.com
blog.cloudflare.com
I have to admit that Cloudflare has been killing it recently with DevX / OpsX. If I wasn't against that company's role in modern internet (as a user of Tor, their firewall is annoying to no end), I would have tried them out already.
Likewise, TOR access is similarly configurable. Companies choose to block it because more often than not it IS bot traffic, and the few potential real customers who use TOR are deemed not worth the headaches of the rest of the network.
Cloudflare's WAF is really pretty granular, with a lot of toggles and overrides: https://developers.cloudflare.com/waf/managed-rules/
Anecdote: For small businesses with limited web resources, these are just everyday tradeoffs they have to make in order to keep hosting and security fees reasonable. At the place I worked at, previously we were spending tens of thousands a year on hosting and thousands more for a competing WAF that cost like 10x more and didn't work very well. Cloudflare let us move to a much lower hosting plan and cost like $240/yr and drastically reduced bot traffic. Not a single customer complained over the next year or two. It was a huge improvement in both performance and costs.
“I don’t like Cloudflare because they’re trying to centralize the Internet and block me”
It’s not as though Cloudflare goes out and randomly inserts themselves in Internet traffic and has some blanket policy of ruining TOR or blocking you.
Cloudflare has customers (site hosts) that have choice in the marketplace and choose them. The customer configures whether their services use Cloudflare or not. The customer configures TOR access, CAPTCHA level, geoblocks, and any other number of hundreds of parameters.
Then people get mad at Cloudflare when a site/host selects Cloudflare and configures it in a way that blocks them?
Cloudflare is selling what people want to buy and providing the service in the way they configure it. If you have a problem with that take it up with the site/host/CF customer, I truly don’t understand how/why they can or should be blamed for their success.
I think what you’ll find is that many Cloudflare customers are practical and pragmatic. Want access to our site over Tor? Sorry but Tor is 99.999% shady/malicious traffic we don’t care about. The risk vs reward isn’t there so blocked. Maybe if a customer says something we’ll enable it but that has never and will never happen so blocked.
Our PCI scans and auditing systems are showing weird traffic from Asia even though we have no customers or business there? Blocked.
Repeat this for any other number of factors and you can start to understand why Cloudflare has double the market share of their nearest competitor (AWS Cloudfront).
They offer a product suite site owners and hosts love. The collateral damage from a tiny fringe of legitimate users who get stuck in the CAPTCHAs, use tor, etc just don’t matter to the site hosts. If they did they would configure Cloudflare differently or leave them altogether.
I think cloudflare is even the only one that supports the "onion routing" to improve the situation for real Tor users.
I understand they don’t want get hacked but believe it or not, customers travel.
Security and convenience are always tradeoffs. I think 2FA is annoying as heck too (much prefer passkeys these days), or ridiculous password requirements, email passwordless login, etc., but those are all choices some admin or manager made on behalf of their business.
I feel like geoblocking is the easy way out, because if developing countries suddenly started waving their cards en masse, these merchants would find a way to let them in.
Speaking of card waving, it likely only appears that developing countries are not a large customer base because to merchants, they look like US customers.
Most African countries don't have access to Visa/Master cards, so often they'll have a US account where they can transfer some of their money. Others might earn in the US (like remote workers), and spend considerably in the US.
Then, because most merchants don't ship outside the US, these customers would use shipping forwarders, like myus.com.
So when making the decision to "block Nigeria because we don't really have any customers there", they're likely not considerig this potentially large customer base they're alienating.
Even worse, these are usually customers that do not have access to credit, only debit, so for example, when buying large ticket items (like a car), they tend to pay for it all upfront, so likely great customers.
Then there are the business customers, the ones what want to buy containers full of merchandise. Those too get blocked.
It's not usually a benefit to a business if a customer pays upfront.
Whether my customer pays by debit or credit, I get all of that money upfront before I let the transaction proceed.
Some businesses, like car dealers, actually make more money if the customer buys using debt, because they get incentivized by the loan company.
And lastly, the sheer scale of the US economy means that it's really not worth the hassle. All of Africa would be equal to one of the larger states (Wikipedia says $3T, Texas is 2.1T and Cali is 3.5T).
So it's vastly simpler, cheaper, and easier to deal with say 30m Texans or 40m Californians than literally 1.3 billion people in Africa or India, and you get roughly the same total addressable market and a fraction of the bots & scams.
Hence why many sites simply block non-North American traffic.
I wish we lived in a world that was more fair and open, but a couple of bad actors can really ruin things for everyone.
Secondly, I also get it, there's only so many things a business can worry about, and supporting geographies with historically high fraud rates is not high on the list, this is why my gripe here is with CF that does not make it easier to improve this even though they know they control such a huge chunk of the web.
100%. Richer people tend to be better customers. But that's another strike in favor of geoblocking non-US visitors.
When I was a kid growing up in Africa, I dreamt of a world where everything was accessible and purchasable and learnable everywhere, all the time, to everyone. Hopefully the internet turns out to be an equalizing factor and we get there someday.
Right now it's not really fair to expect business owners - most of whom are in non-tech businesses that require 100% focus - to keep up with the tidal wave of scams, hackers, and regulators originating from outside their sphere of concern.
Sure. Even by default, Cloudflare won't block entire countries. That's a CHOICE some businesses make if the default blocks aren't enough, and they don't have the time or resources to configure more nuanced WAF rules. (OWASP isn't exactly straightforward). Edit: For example, at that job I was talking about, we had different rulesets for different regions... China and Russia were completely banned, Africa was put behind stricter JS security checks and CAPTCHAs but allowed in, Europe had a medium security level (we did occasionally sell there, but very rarely), while the US had entirely custom WAF rules. It just depends on who we wanted to sell to or not.
It goes the other way around, too, you know. I've seen European and Asian sites that geoblock US customers. It's not out of malice, they just don't want to deal with the edge cases. Even if a foreign customer can access your website and buy stuff, dealing with international customs, consumer laws, credit card fraud, wire transfers, etc. can be a pain that's not worth it for smaller merchants. And if the foreign buyer is using a reshipper anyway, well, the reshipper can just buy the whole thing for them and deal with payments, etc. as an intermediary, like how Tenso/BuyFromJapan/JapanRabbit work.
Big companies have proper international presences, but for small local businesses, the amount of effort it takes to support international buyers just isn't worth the profit they typically bring in. Even on eBay, with its built-in international payment and shipping rules, sellers often won't want to bother.
This isn't really a matter of security rules, really, but just business cost/benefit decisions.
Besides, it helps businesses in each country stay local! Do you really want Amazon taking over everywhere...?
And yes! it's better to buy local, and Africa can't blame the US because our economy isn't there, and we aren't building all the things we should be building. But that is an entirely different discussion isn't it?
We also aren't just talking about blocking DDoS and other common vulnerability scanning. Depending on your business there are other potentially costly fraud and abuse scenarios that you are blocking just by blocking other countries outright. Until there are tools to block all this that are as easy to apply as a geoblock, this will probably remain the unfortunate state of things. A lot of businesses just don't have the time or resources to manage all of this without applying geoblocks.
i would think that it adds on a huge cost?
It is not just about customers, you have to thibk about ecosystem as a whole.
Wow, why? Extortion? Competition? Collateral damage?
But if you don't have any DDoS protection set up, either of these attacks will essentially be L7 DDoS attacks when deployed at scale.
You need DDoS protection when someone does not want your service to stay up.
Further, cred stuffing is often automated by a botnet.
The two things are distinct, but have similar means and end results - they aren't completely different.
I suspect a lot of the malicious traffic coming out of Africa is not direct attacks from cybercriminals but residential machines that have also been compromised to send malicious traffic. The only difference between that and cloud providers is that you cannot afford to block all of Amazon or Google. They have a level of economic privilege that the entire continent of Africa lacks.
Regardless of whether or not it’s the intention of the attackers to disrupt the service, that’s the effect it can have, and it’s something DDoS protections service will usually mitigate, especially the services that incorporate WAF functionality (which I think is pretty much all of them?…).
When bot fight is on, we don’t notice anything.
That'd be one distinct advantage of using Cloudflare over AWS, regardless of how opinionated it may appear; and you get to fine tune some of the settings if you're a paying customer.
Anthropomorphizing technical services is probably not going to lead to good conclusions. Better suggestions will.
Let's try. I suppose you know 1.1.1.1 but:
- Their WebAnalytics
- Flexible SSL ( instead of letsencrypt)
- Their free Hugo setup ( for your blog) -> Cloudflare Pages
- Buy DNS domains at cost
- 500 Cloudflare worker scripts for 5€ / month. Or 100 Cloudflare worker scripts for free
- Cloudflare tunnel - instead of ngrok or others. You can link it to your subdomain, other options have a paid option if you want to link a subdomain.
Cloudflare doesn’t “block” anything universally everything is completely configurable (other than obvious exploits like mass-blocking the recent http/2 attacks)
It's due to their users and associated behavior that tor Exit nodes have an elevated bad reputation.
> https://developers.cloudflare.com/support/firewall/learn-mor...
> Due to the behavior of some individuals using the Tor network (spammers, distributors of malware, attackers, etc.), the IP addresses of Tor exit nodes may earn a bad reputation, elevating their Cloudflare threat score.
Customers of cloudflare have an option to improve experience for Tor users
> Beyond applying firewall filters to Tor traffic, Cloudflare users can improve the Tor user experience by enabling Onion Routing. Onion Routing allows Cloudflare to serve your website’s content directly through the Tor network, without requiring exit nodes.
Email the sites where you have issues and ask them to enable Tor routing.
I'm also not understanding how enabling Tor routing prevents bot traffic from hitting the site. The traffic gets served over a .onion instead, cool. But how does that prevent the bots?
How many of us deal with automated password attacks is to issue questions that only locals or people with specific knowledge could answer. Change the questions and do everything custom.
> How many of us deal with automated password attacks is to issue questions that only locals or people with specific knowledge could answer. Change the questions and do everything custom.
If I'm understanding what you're saying, this sounds horrible. What if I'm visiting an area where I don't have local knowledge? What about for the year or so after I move in to a new city? What if your assessment of what locals do and don't know is just wrong? There are a ridiculous number of failure modes in this questions-oriented approach. The only place this could possibly make sense is in some sort of internal company software, but even that context has better options available.
At the country level (and for applications where you have enough control over your infrastructure to use a real firewall) I question both the efficacy and accessibility of a system like you propose—it's not that different from the old style "what is 2+2" CAPTCHAs, and there's a good reason why most applications have moved on from those. They're not a serious alternative to behavioral rules like what OP describes.
...that's on by default and so used by the vast majority of Cloudflare customers making it effectively a Cloudflare configuration.
And everyone knows it because that's what the lived experience of trying to access cloudflare blocked sites on tor browser (or any other browser that's not made by a megacorp). It doesn't matter what cloudflare's intentionally ambiguous and probably disingenuous wording might try to imply. The only people who think otherwise have never actually tried using tor to surf the web.
This one shows you how to set it up in cloudflare.
While the info provided in both is indeed similar.
There’s not much AWS can do about it because they must make untold billions from those sweet, sweet S3 egress fees.
I’d be willing to bet S3 egress fees make up about 60% of all AWS revenue.
Like when you set up an RDS instance, the “prod” template defaults to multi-AZ (a good idea tbf), but completely elides the fact that if your app is in a different AZ, you’re going to start racking up $0.02/GB.
Same with NLBs and cross-AZ routing. Sure, it can be helpful, but yeesh.
Or EKS, since Topology Aware Routing is in no way a default.
Eg. For cloudflare workers. If you're worker is making an outgoing request ( db / rest) it's not considered cpu-time and it's not counted towards that either ( in cloudflare ofc).
While this is a hidden profit of many cloud providers :)
Wait time isn’t calculated as compute?
https://community.cloudflare.com/t/how-is-cpu-time-per-reque...
Cloudflare is only billing that actual cpu-time.
Edit: this is a better resource https://blog.cloudflare.com/workers-pricing-scale-to-zero/
An API call can take a lot of seconds. While Cloudflare only bills cpu-time ( eg. 10 ms. ). Other providers bill those seconds too as "duration", while the CPU was just sitting idle.
I thought that this was possible because Cloudflare eliminated cold start delays.
So there's no RAM reserved either, I guess.
( can someone correct me if I'm wrong?)
Then, one month, I got a ~$500 bill out of no where.
Docker had changed an api causing my service to return 5xx errors all month. Each error was individually logged to CloudWatch - which racked up a ~$500 bill.
I moved to Cloudflare Workers that day and haven’t moved back.
The Cloud really loves logging ( bills :p ).
It would be nice if Cloudflare implemented "Open telemetry".
It could reduce the cloud bill by at least 2. Logging is really expensive.
A good number of people end up using cloud watch for all of the above, even though it’s (comparatively) mid.
I’m in a seemingly small subset of people that is very happy with AWS for side projects. Granted I’m not doing anything that requires many resources.
If you want to lock ec2 access to cloudfront only you can do it in SG with "managed prefix list for CloudFront".
We have tens of TB in AWS that we'd like to move to CF, but I'm reluctant to without being able to know if we're going to get a call demanding we switch to enterprise.
I don't want to take advantage of them and get on their abuse list since this is production. I'm happy to pay more! I just don't want to deal with negotiating an Enterprise plan. They ask you so many questions like "how many Page Rules do you want? How many Worker requests?" I just want R2. And this response confuses them too because they say "well R2 is pay-what-you-use..." I would honestly be happier with a $5000/mo "excessive R2 bandwidth" fee. But they don't seem to want to implement that.
Enterprise deals start at 100-200k/year, minimum, IME.
Then your 24k a year sales person number is assuming that the sales person will sell one license the entire year, an insane assumption.
probably someone who spends no more than one man-week on this sale, which doesn't seem that hard if the customer is a small org and only wants to buy one feature? sure, if you expect to spend months negotiating and years implementing, you'll need to charge a lot to make up for the time, but we're not talking about that case here.
Enterprise salespeople working in that segment understand you're small-fry and they'll kick you a small contract not because they're going to get rich off the commission but because they know that depending solely on whales is a high-risk approach that sooner or later gets them fired. Small contracts pad them out and provide more reliable monthly numbers as long as the sales process doesn't turn into a tarpit that isn't worth the money.
Do not think its free for all usage.
>>
PUT, COPY, POST, LIST requests (per 1,000 requests) = $0.005
GET, SELECT, and all other requests (per 1,000 requests) = $0.0004
>>
"You pay for requests made against your S3 buckets and objects. S3 request costs are based on the request type, and are charged on the quantity of requests as listed in the table below" https://aws.amazon.com/s3/pricing/#:~:text=You%20pay%20for%2...
“DELETE and CANCEL requests are free.”
> You can use Amazon S3 Inventory to help manage your storage. For example, you can use it to audit and report on the replication and encryption status of your objects for business, compliance, and regulatory needs. You can also simplify and speed up business workflows and big data jobs by using Amazon S3 Inventory, which provides a scheduled alternative to the Amazon S3 synchronous List API operations. Amazon S3 Inventory does not use the List API operations to audit your objects and does not affect the request rate of your bucket.
>
> Amazon S3 Inventory provides comma-separated values (CSV), Apache optimized row columnar (ORC) or Apache Parquet output files that list your objects and their corresponding metadata on a daily or weekly basis for an S3 bucket or objects with a shared prefix (that is, objects that have names that begin with a common string). If you set up a weekly inventory, a report is generated every Sunday (UTC time zone) after the initial report. For information about Amazon S3 Inventory pricing, see Amazon S3 pricing.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/storag...
Still use lifecycle policies (I visualized how it ages out 1 billion items here[1]), but list request prices are not a factor worth mentioning.
1. https://tomforb.es/visualizing-how-s3-deletes-1-billion-obje...
- click on the R2 link in the Cloudflare dashboard and add a bucket
- upload files to be shared (max 300MB)
- click the settings tab and add custom domain that you want files shared from
This prevented me from having to upgrade to Vercel Pro and saved me $240/y. A bit more on the topic: https://sometechblog.com/posts/don-t-overpay-for-bancwidth/
In particular the following line should enable it:
> Unlike most Cloudflare products, the Developer Platform can be used to host content.
Also see https://blog.cloudflare.com/updated-tos/
> Over time, Cloudflare’s network became larger and more robust and its portfolio broadened to include services like Stream, Images, and R2. These services are explicitly designed to allow customers to serve non-HTML content like video, images, and other large files hosted directly by Cloudflare.
> Video and large files hosted outside of Cloudflare will still be restricted on our CDN, but we think that our service features, generous free tier, and competitive pricing (including zero egress fees on R2) make for a compelling package for developers that want to access the reach and performance of our network.
(of course this is not legal advice)
Our last month invoice was only $0.26 for 26GB+ of managed storage.
https://blog.cloudflare.com/aws-egregious-egress/ comments https://news.ycombinator.com/item?id=27930151
https://developers.cloudflare.com/images/image-resizing/
https://developers.cloudflare.com/images/image-resizing/url-...
Example: src="/cdn-cgi/image/width=80,quality=75/uploads/avatar1.jpg
About /cdn-cgi/image/
It's a fixed prefix that identifies that this is a special path handled by Cloudflare’s built-in Worker.
---
Price at Cloudflare : 50,000 monthly resizing requests included with Pro, Business. $9 per additional 50,000 resizing requests.
See: https://www.cloudflare.com/plans/#add-ons
As a reference: Vercel is $5 per 1000 source images. So Cloudflare is a whopping 25 x cheaper.
Price at Vercel: https://vercel.com/docs/image-optimization/limits-and-pricin...
I checked quickly and cloudinary works with credits. Cloudflare mentions 1$ / 100 k. Images served, which seems pretty cheap to me.
The pricing I mentioned was for image manipulation, eg. Resizing ( not storage), which was what was asked.
I doubt they are cheaper though.
Vercel : 5$ / 1000 requests
Cloudflare : 9$ / 50.000 requests
Google Cloud is already in the bandwidth alliance, while AWS is not.
So you can send your files already to cloudflare at a discounted rate.
https://www.cloudflare.com/bandwidth-alliance/
Additionally, there are reasons to keep your data duplicate within AWS. There's no one-size-fits-all here.
The point is to make use of cloudflare free eggres to get your data where you need it ( since others have egress costs, but no ingress costs)
Eg. Train ml models in AWS, GCE or Azure ( where compute is cheapest) and make use of cloudflare to provide them with the data.
This should create a huge cost saving, since there's no eggres anywhere.
Example: https://blog.cloudflare.com/cloudflare-r2-mosaicml-train-llm...
At the end, you can still decide to migrate all data or to abandon the not-used-till-now data.
Getting that data is really high. It should only be used if you are sure it will never be accessed.
I've read quite a few posts where they complained about the huge bill, when they needed to get their data...
Not sure how it is right now, it's still obscure on their pricing page and before, you had to check the gotchas in their FAQ .
aws and cf are the perfect combo, and them competing on quality and price is great for consumers.
aws is the control plane cf is the data plane and entrypoint. what a time to build.
The only problem that remains is their support...
I'm writing a book about Cloudflare (launching very soon) where I share this and many other things to scale faster all while saving big on your cloud bills. You can join the waiting list here: https://kerkour.com/subscribe
The most consistently amazing thing about Cloudflare is the clarity of their product positioning.
You have this common problem, we built a thing to fix it.
No 'change your problem into this other problem' gymnastics. Just 'pay us once you exceed the free tier, and it's no longer a problem'.
And furthermore, they seem to have clarity of platform vision, in that each piece does something very specific to help them compete efficiently against AWS/Azure/GCP (who have much larger resources) AND has synergies with their existing platform. E.g. edge compute, free/cheaper network traffic from compute/storage
Critically, Cloudflare seems like the only competitor to the majors that has their eyes on competing on price by capturing enough of the market of {some thing} that they can still make profits at extremely low price points.
Also, just glanced at their financials again, and they look exactly like you'd want to run a large company if your eye was on order of magnitude growth. They just pivoted to positive FCF in 2023, biggest expense is sales and marketing (over half their gross profit), and have exponential revenue growth.
https://softwarestackinvesting.com/cloudflare-net-q2-2023-ea...
Their support is superbe, but it takes a while to access.
The only time we needed them, the chat option seemed relatively quick.
+ they pointed us to a tls connect issue at Azure with a very detailed analysis of why.
Thing is. If you see a cloudflare error page, it's probably you're hosting provider and not cloudflare...
We'll see if NET survives public investor expectations.
Is your info still up to date? ( I'm not following this topic too much, but I do remember some things passing by).
---
Additionally, most of their investors are companies and not private.
There's a lot going on. One of the improvements that they did was in the sales department.
If those previous sales that were severely underperforming are now replaced by even average sales. Then expect a big rise in sales for Q3.
Reference: https://softwarestackinvesting.com/cloudflare-net-q2-2023-ea...
This is about R2
> Storage: $0.015 / GB - 10 GB free
> Class A operations (mutate state): $4.50 per million - 1 M. free
> Class B operations (read state): $0.36 per million - 10 M. free
Cloudflare workers:
> $0.15/million requests per month ( 100 k. / day included)
> Up to 30s wall time per request
Min. 5 $ / m. when exceeding the free tier.
Info that there are "zero egress fees" is only available on R2 product page and not pricing page.
IMO R2 pricing page look like it only displays quick info and that there might be fine print somewhere, but there is no link to more details. It could be that's all there is to it, but somehow design feels off to me. Especially because of the "zero egress fees" info being displayed only on the product page.
Workers product page shows "Maximum number of scripts": 30 free, 100 paid. But on workers pricing page it shows "Up to 100 Worker scripts" for free and "Up to 500 Worker scripts" for paid.
Links to different sections (Pricing, Products...) don't have an option to open in new tab. IMO the whole website is weirdly organized. But maybe it's just me.
Your are billed every month for your total storage. Every month you don't need to pay for the first 10 GB.
I'm not really sure if eggres has to be mentioned if it's widely known that Cloudflare doesn't bill for eggres. But I get your point.
Concerning Cloudflare workers:
Workers has a free tier and a paid tier at 5€/month.
The free tier has a limit of 100 workers and the paid tier has a limit of 500 workers.
Perhaps just scroll down a bit more on the pricing page of Cloudflare workers. I'm assuming you are checking it on Mobile and missed that.
About but being able to open pricing in a new tab. I noticed the same.
---
They also have a minor UX issue that if you want to go to the Web analytics page, the menu goes to the first child and hides. So you'll have to click it open again and click on Web analytics ( again, just an issue on Mobile)
EDIT: I'm testing on desktop.
Notified them on their Discord of workers, let's see if it gets picked up tomorrow.
Edit: It's going to be escalated and fixed ( Got a response within 44 minutes ... On a Sunday, nice).