Comparing AWS S3 with Cloudflare R2: Price, Performance and User Experience
kerkour.com
kerkour.com
A business contact of mine asked Cloudflare for a quote for setting up a CDN for 7B requests per month. They dragged him through five sales calls to eventually deliver an offer with request costs 30% above CloudFront's public pricing. He said the costs per GB were ok though.
The cheapest reliable CDN I've found is bunny.net. Unlike OP, who is selling a book about Cloudflare, I have no conflict of interest by recommending Bunny, other than being a customer. I serve around 100 TB per month through them, at $0.005/GB. You can put it in front of a bucket to have cheap (globally distributed) egress with all the benefits of your favorite object storage provider. In my case, the buckets lie on Linode/Akamai. Bunny uses CDN77's infrastructure, which I have also heard good things about.
While Cloudflare may reach out and say 'you should be on enterprise' when that happens on R2, the fact they also handle DDoS and similar attacks as part of their offering means the likelihood of success is much lower (as is the final bill).
It should be possible to use the service, especially common ones like S3 with little knowledge of architecture and stuff.
S3’s simple setup (which denies all public access) is not flawed in the manner being discussed here. Allowing public direct access to an S3 bucket is a supported option, but for years has been both non-default and strongly recommended against.
It doesn't take anything near DDoS. If you dare to put up a website that serves images from S3, and one guy on one normal connection decides to cause you problems, they can pull down a hundred terabytes in a month.
Is serving images from S3 a crazy use case? Even if you have signed and expiring URLs it's hard to avoid someone visiting your site every half hour and then using the URL over and over.
> AWS is just charging you for how much it served, it doesn't make sense to hold them to a fault here.
Even if it's not their fault, it's still an "inherent vulnerability of S3 pricing". But since they charge so much per byte with bad controls over it, I think it does make sense to hold them to a good chunk of fault.
Do you expect DDoS protection to kick in from one person downloading a single digit number of images per second?
[0] https://aws.amazon.com/about-aws/whats-new/2024/08/amazon-s3...
AWS made S3 as data storage, then added CloudFront for CDN purposes. The CDN is an optional addon which may or may not make sense... E.g. an internal data storage staying in AWS doesn't need the CDN.
Comparing CloudFlare to S3 is apples and oranges, in my option. Comparing CloudFlare+R2 and S3+CloudFront is more appropriate, I think.
---
Additional thoughts:
It's hard to evaluate the bias of the author.. The author clearly dislikes AWS, and while some of the plants are generally what I would agree with.. I question if the author is properly evaluating AWS in the comparison and trying to sell their book. Would the book be any better, or would it echo chamber the typical AWS perceived negatives. For instance, the author notes S3 intelligent tiering... But if you know that the data is not going to be accessed, you could likely skip the overhead of intelligent tiering and directly put it into a cheaper storage class... And in general the same with other S3 data types.
I do generally agree that AWS bandwidth charges are extortion... But I still would want less bias in a product review/comparison from something posted here on hacker news.
I have a few of the S3-like wired up live over the internet you can try yourself in your browser. Backblaze is surprisingly performant which I did not expect (S3 is still king though)
https://observablehq.com/@tomlarkworthy/mps3-vendor-examples
It's misleading to look at S3 as a CDN. It's fine for that, but it's real strength is backing the world's data lakes and cloud data warehouses. Those workloads have a lot of data that's often cold, but S3 can deliver massive throughout when you need it. R2 can't do that, and as far as I can tell, isn't trying to.
Source: I used to work on S3
1: https://blog.cloudflare.com/how-cloudflare-auto-mitigated-wo...
Cloudflare has the network to do that, but they charge money to do so with their other offerings, so why would they give that to you for free? R2 is not a CDN.
neither of these tell you how fast you can serve static data.
Yeah, I'm sure they use a completely different network infrastructure to serve R2 requests.
S3 can't saturate user bandwidth unless you make many parallel requests. I'd be (pleasantly) surprised if R2 can.
If we are talking about storage, well, SATA can't give you more than ~5Gbps so I guess the answer is no? But also no one else can do it, unless they're using super exotic HDD tech (hint: they're not, it's actually the opposite).
What a weird thing to argue about, btw, literally everybody is running a network layer on top of storage that lets you have much higher throughput. When one talks about R2/S3 throughput no one (on my circle, ofc.) would think we are referring to the speed of their HDDs, lmao. But it's nice to see this, it's always amusing to stumble upon people with a wildly different point of view on things.
There are any number of ways they could implement R2 that would allow it to run at full wire speed, but S3 doesn't run at full wire speed by default (unless you make many parallel requests) and I'd be surprised if R2 does.
I have some large files stored in R2 and a 50Gbps interface to the world.
curl to Linode's speed test is ~200MB/sec.
curl to R2 is also ~200MB/sec.
I'm only getting 1Gbps but given that Linode's speed is pretty much the same I would think the bottleneck is somewhere else. Dually, R2 gives you at least 1Gbps.
But if you look at https://aws.amazon.com/s3/ it says things like:
"Object storage built to retrieve any amount of data from anywhere"
"any amount of data for virtually any use case"
"S3 delivers the resiliency, flexibility, latency, and throughput, to ensure storage never limits performance"
So if S3 is not intended for low-latency applications, the marketing team haven't gotten the message :)
Personally I think you have a point.
Cloudflare wants to "protect" the world from the evils of DNS services other than themselves even knowing what geographical region people are in, so they strip all geographical information, even general, broad location, from DNS lookups. This has the effect of increasing latency for non-Cloudflare CDNs sometimes, since data will sometimes end up being served out of the wrong region.
I've wondered since I first heard about this if this is their way to enshittify CDN deliverability in general and make their latency look better in comparison.
Source: I used to work at Netflix, building systems that pull TBs from S3 hourly
So apart from ToS abuse cases, do you know any other cases? I ask as a genuine curiosity because I’m currently paying for Cloudflare to host a bunch of our websites at work.
Put another way, if Cloudflare really had free unlimited CDN egress then every ultra-bandwidth-intensive service like Imgur or Steam would use them, but they rarely do, because at their scale they get shunted onto the secret real pricing that often ends up being more expensive than something like Fastly or Akamai. Those competitors would be out of business if CF were really as cheap as they want you to think they are.
The point where it stops being free seems to depend on a few factors, obviously how much data you're moving is one, but also the type of data (1GB of images or other binary data is considered more harshly than 1GB of HTML/JS/CSS) and where the data is served to (1GB of data served to Australia or New Zealand is considered much more harshly than 1GB to EU/NA). And how much the salesperson assigned to your account thinks they can shake you down for, of course.
> Cloudflare’s content delivery network (the “CDN”) Service can be used to cache and serve web pages and websites. Unless you are an Enterprise customer, Cloudflare offers specific Paid Services (e.g., the Developer Platform, Images, and Stream) that you must use in order to serve video and other large files via the CDN. Cloudflare reserves the right to disable or limit your access to or use of the CDN, or to limit your End Users’ access to certain of your resources through the CDN, if you use or are suspected of using the CDN without such Paid Services to serve video or a disproportionate percentage of pictures, audio files, or other large files. We will use reasonable efforts to provide you with notice of such action.
https://www.cloudflare.com/service-specific-terms-applicatio...
For example maybe you have ~500 GB of data across millions of objects that has accumulated over 10 years. You don't even know how many reads or writes you have on a monthly basis because your S3 bill is $11 while your total AWS bill is orders of magnitude more.
If you're in a spot like this, moving to R2 to potentially save $7 or whatever it ends up being would end up being a lot more expensive from the engineering costs to do the move. Plus there's old links that might be pointing to a public S3 object which would break if you moved them to another location such as email campaign links, etc..
And let's be honest. If the house burns down, the computers are the third thing I get out of there after the wife and the dog. My external backup is peace of mind, nothing more. I don't ever expect to need it in my lifetime.
I backup to Glacier as well. For me to need to pull from it (and pay that $90/TB or so) means I've lost more than two drives in a historically very reliable RAIDZ2 pool, or lost my NAS entirely.
I'll pay $90/TB over unknown $$$$ for a data recovery from burned/flooded/fried/failed disks.
I'm not saying its a bad deal, I very much wish I'd gone with deep archive (especially with the recent addition of "update if different" to the API)
With 10TB drives being available for $80-ish... I expect a lot of semi savy people to upload tens of terabytes then cry when they see the bill.
Numbers rounded a bit for simplicity.
https://community.cloudflare.com/t/r2-object-versioning-and-...
Seems like a bug. Had to crawl through documentation to find out the only support is on Discord (??), so I had to sign up.
Go through some more hoops and eventually get to a channel where I received a prompt reply: it’s not an R2 issue, it’s “expected behaviour due to an issue with “the CDN service”.
I mean, sure. On a technical level. But I shoved some data into your service and basic standard HTTP semantics where intermittently not respected: that’s a bug in your service, even if the root cause is another team.
None of this is documented anywhere, even if it is “expected”. Searching for [1] “r2 http range” shows I’m not the only one surprised
Not impressed, especially as R2 seems ideal for serving Parquet data for small projects. This and the janky UI plus weird restrictions makes the entire product feel distinctly half finished and not a serious competitor.
1. https://www.google.com/search?q=r2+http+range&ie=UTF-8&oe=UT...
I mean... that's wrong? If you come across such software, do you at least file a bug?
It’s not clear how you’d expect to handle a webserver trying to send you 1Gb of data after you asked for a specific 10kb range other than aborting.
To be clear: most software does handle it, because it detects this case and aborts.
But to a user who is explicitly asking to read a parquet file without buffering the entire file into memory, there is no distinction between a server that cannot handle any range requests and a server that can occasionally handle range requests.
Other than one being much, much more annoying.
Whenever a new incumbent gets on the scene offering the same thing as some entrenched leader only better, faster, and cheaper, the standard response is "Yeah but it's less reliable. This may be fine for startups but if you're <enterprise|government|military|medical|etc>, you gotta stick with the tried tested and true <leader>"
You see this in almost every discussion of Cloudflare, which seems to be rapidly rebuilding a full cloud, in direct competition with AWS specifically. (I guess it wants to be evaluated as a fellow leader, not an also-ran like GCP/Azure fighting for 2nd place)
The thing is, all the points are right. Cloudflare IS different - by using exclusively edge networks and tying everything to CDNs, it's both a strength and a weakness. There's dozens of reasons to be critical of them and dozens more to explain why you'd trust AWS more.
But I can't help but wonder that surely the same happened (i wasn't on here, or really tech-aware enough) when S3 and EC2 came on the scene. I'm sure everyone said it was unreliable, uncertain, and had dozens of reasons why people should stick with (I can only presume - VMWare, IBM, Oracle, etc?)
This is all a shallow observation though.
Here's my real question, though. How does one go deeper and evaluate what is real disruption and what is fluff. Does Cloudflare have something that's unique and different that demonstrates a new world for cloud services I can't even imagine right now, as AWS did before it. Or does AWS have a durable advantage and benefits that will allow it to keep being #1 indefinitely? (GCP and Azure, as I see it, are trying to compete on specific slices of merit. GCP is all-in on 'portability', that's why they came up with Kubernetes to devalue the idea of any one public cloud, and make workloads cross-platform across all clouds and on-prem. Azure seems to be competitive because of Microsoft's otherwise vertical integration with business/windows/office, and now AI services).
Cloudflare is the only one that seems to show up over and over again and say "hey you know that thing that you think is the best cloud service? We made it cheaper, faster, and with nicer developer experience." That feels really hard to ignore. But also seems really easy to market only-semi-honestly by hand-waving past the hard stuff at scale.
You wouldn't build a cloud from scratch in this way.
We’re using it to power the OpenTofu Provider&Modules Registry[0][1] and it’s honestly been nothing but a great experience overall.
[0]: https://registry.opentofu.org
[1]: https://github.com/opentofu/registry
Disclaimer: CloudFlare did sponsor us their business plan so we got access to higher-tier functionality
The magic that moves the region sounds like a dealbreaker for any use cases that aren't public, internet-facing. I use $CLOUD_PROVIDER because I can be in the same regions as customers and know the latency will (for the most part) remain consistent. Has anyone measured latencies from R2 -> AWS/GCP/Azure regions similar to this[0]?
Also does anyone know if the R2 supports the CAS operations that so many people are hyped about right now?
https://cloud.google.com/cdn/docs/using-signed-cookies
As far as I can tell, this feature is also supported by Akamai here:
https://techdocs.akamai.com/property-mgr/docs/cookie-authz
I am pretty sure you can implement this on CDNetworks using eval_func:
https://docs.cdnetworks.com/en/cdn/docs/recipes/secure-deliv...
With AWS Cloudfront, I'd think you--worst case--pull out Lambda@Edge?
IIRC its essentially the same as the aws style signed urls and header bearer token auth. I _think_ lambda@edge is only relevant if you want to do the initial sig generation in the cdn instead of your api/app endpoint.
Edit: actually GP mentioned Cloudfront already, so yes works as theyre asking for AFAICT
But most importantly for an indie dev like me the cost became $0.
Cloudflare’s documentation just says “we offer 11 9s, same as S3”, and that’s that. It’s not that I don’t believe them but… how can a smaller organization make the same guarantees?
It implies to me that either Amazon is wasting a ton of money on their reliability work (possible) or that cloudflare’s 11 9s guarantee comes with some asterisks.
Amazon has an entire automated reasoning group (researchers who mostly work on formal methods) working specifically on S3.
As far as I’m aware, nobody at cloudflare is doing similar work for R2. If they are, they’re certainly not publishing!
Money might not be the bottleneck for cloudflare though, you’re totally right
The 11 9's is for durability, which is really more about the redundancy setup, erasure coding, etc. (https://cloud.google.com/blog/products/storage-data-transfer...)
fwiw availability is 4 9's (https://aws.amazon.com/s3/storage-classes/)
I think I overstated the case a little, I definitely don’t think automated reasoning is some “secret reliability sauce” that nobody else can replicate; it does give me more confidence that Amazon takes reliability very seriously, and is less likely to ship a terrible bug that messes up my data.
I've seen complaints of users about R2 having erratic upload speeds.
For example: no regions, no replication (and no AZs either), limited lifecycle management, no versioning, no MFA protection, no intelligent tiering, no customer encryption, no IAM, etc.
> As explained here, Durable Objects are single threaded and thus limited by nature in the throughput they can offer.
R2 bucket operations do not use single threaded durable objects but did a one off thing just for R2 to let it run multiple instances even. That’s why the limits were lifted in the open beta.
> they mentioned that each zone's assets are sharded across multiple R2 buckets to distribute load which may indicated that a single R2 bucket was not able to handle the load for user-facing traffic. Things may have improve since thought.
I would not use this as general advice. Cache Reserve was architected to serve an absurd amount of traffic that almost no customer or application will see. If you’re having that much traffic I’d expect you to be an ENT customer working with their solutions engineers to design your application.
> First, R2 is not 100% compatible with the S3 API. One notable missing feature are data-integrity checks with SHA256 checksums.
This doesn’t sound right. I distinctly remember when this was implemented for uploading objects. Sha-1 and sha-256 should be supported (don’t remember about crc). For some reason it’s missing from the docs though. The trailer version isn’t supported and likely won’t be for a while though for technical reasons (the workers platform doesn’t support http trailers as it uses http1 internally). Overall compatibility should be pretty decent.
The section on “The problem with cross-datacenter traffic” seems to be flawed assumptions rather than data driven. Their own graphs only show that while public buckets have some occasional weird spikes it’s pretty constantly the same performance while the S3 API has more spikeness and time of day variability is much more muted than the CPU variability. Same with the assumption on bandwidth or other limitations of data centers. The more likely explanation would be the S3 auth layer and the time of day variability experienced matches more closely with how that layer works. I don’t know enough of the particulars of this author’s zones to hypothesize but the s3 with layer was always challenging from a perf perspective.
> This is really, really, really annoying. For example you know that all your compute instances are in Paris, and you know that Cloudflare has a big datacenter in Paris, so you want your bucket to be in Paris, but you can't. If you are unlucky when creating your bucket, it will be placed in Warsaw or some other place far away and you will have huge latencies for every request.
I understand the frustration but there are very good technical and UX reasons this wasn’t done. For example while you may think that “Paris datacenter” is well defined, it isn’t for R2 because unlike S3 your metadata is stored regionally across multiple data centers whereas S3 if I recall correctly uses what they call a region which is a single location broken up into multiple availability zones which are basically isolated power and connectivity domains. This is an availability tradeoff - us-east-1 will never go offline on Cloudflare because it just doesn’t exist - the location hint is the size of the availability region. This is done at both the metadata and storage layers too. The location hint should definitely be followed when you create the bucket but maybe there are bugs or other issues.
As others noted throughput data would also have been interesting.
Maybe it was an old thing? The changelog [0] for 2023-06-16 says:
"S3 putObject now supports sha256 and sha1 checksums."
[0]: https://developers.cloudflare.com/r2/platform/changelog/#2023-06-16Also, has CF improved their stance around hosting hate groups? They have strongly resisted pressure to stop hosting/supporting hate sites like 8chan and Kiwifarms, and only stopped reluctantly.
Their job isn’t to investigate and punish harassment and criminal behavior, but they certainly don’t have to condone it via their support.
If they are known bad actors, let the police do the job of policing the internet. Otherwise, all bad actors are ultimately arbitrarily defined. Who said they are known bad actors? What does that even mean? Why does that person determining bad actors get their authority? Were they duly elected? Or did one of hundreds of partisan NGOs claim this? Who elected the NGO? Does PETA get a say on bad actors?
Be careful what you wish for. In some US States, I am sure the attorney general would send a letter saying to shut down the marijuana dispensary - they're known bad actors, after all. They might not win a lawsuit, but winning the support of private organizations would be just as good.
> they certainly don’t have to condone it via their support
Wow, what a great argument. Hacker News supports all arguments here by tolerating people speaking and not deleting everything they could possibly disagree with.
Or maybe, providing a service to someone, should not be seen as condoning all possible uses of the service. Just because water can be used to waterboard someone, doesn't mean Walmart should be checking IDs for water purchasers. Just because YouTube has information on how to pick locks, does not mean YouTube should be restricted to adults over 21 on a licensed list of people entrusted with lock-picking knowledge.
Cloudflare protects these organizations. Cloudflare goes far above and beyond what most other companies do, and I personally can't wait to see Cloudflare held liable for content they host for which they ignore legitimate complaints.
Imagine if I were to call your phone every hour, on the hour, and my position was that it's not illegal until someone reports it as a crime and the police or a court contacts me to tell me to not do it. That's Cloudflare. They're assholes, and insisting that they're not responsible for anything, any time, until there's a court order is just them being assholes.
Is this reflective of Cloudflare, or instead reflective of other companies’ willingness to play judge, jury, and executioner?
I like knowing my business on Cloudflare won’t be subject to extrajudicial punishment regardless of the grounds.
> They're assholes
Being an asshole is your constitutional right.
Every large company does business with many people / orgs I don't like. I'm not defending or attacking AWS or CF, but merely stating the deeper you dig the more objectionable stuff you'll find everywhere. There are shades of gray of course, but at the end of the day we're all sinners.
Yeah this is really annoying. That and replication to multiple regions is the reason we're not using R2.
Global replication was a feature announced in 2021 but still hasn't happened:
> R2 will replicate data across multiple regions and support jurisdictional restrictions, giving businesses the ability to control where their data is stored to meet their local and global needs.
https://www.cloudflare.com/press-releases/2021/cloudflare-an...
+1
[1] https://aws.amazon.com/blogs/aws/free-data-transfer-out-to-i...