From S3 to R2: An economic opportunity
dansdatathoughts.substack.com
dansdatathoughts.substack.com
It allows you to incrementally migrate off of providers like S3 and onto the egress-free Cloudflare R2. Very clever idea.
He calls R2 an undiscovered gem and IMO this is the gem's undiscovered gem. (Understandable since Sippy is very new and still in beta)
Cloudflare has huge ingress, because they need it to protect sites against DDOS.
They basically already pay for their R2 bandwidth ( = egress) because of that.
Additionally, with their SDN ( software defined networking) they can fine-tune some of the Data-Flow/bandwidth too.
That's how I understood it, fyi.
Some more info could be found when they started ( or co-founded, not sure) the bandwidth alliance.
Eg.
A CDN keeps the data nearby, reducing the need to pay egress to the big bandwidth providers.
( not an expert though)
You setup your website and preferably DON'T have it talk to anyone other than the CDN.
You then point your DNS to wherever the CDN tells you to. (Or let them take over DNS. Depends on the provider.)
The CDN then will fetch data from your site and cache it, as needed.
Your site is the "origin", in CDN speak.
If Cloudflare can move the origin within their network, there is huge cost savings and reliability increases there. This is game changing stuff. Do not under estimate it.
In the CDN case Cloudflare has to fetch it from the origin, cache (store) it anyway, and then egress it. By charging for R2 they're moving that cost center to a profit one.
It's an option that's often been attractive if/when you didn't want the hassle of building out something that could provide S3 level durability yourself. But with more/cheaper S3 competitors it's becoming a significantly less attractive option.
I used the storage box from hetzner before but they only had 1TB or 5TB (and higher) choices so I had to pay for 5TB (€12 per month) without using most of it. Having rsync support was nice but rclone works fine with S3.
The customer (because they are lazy, don't know better, aren't capable of, or all three) opts in to use various "convenient" CSP "services". These services could look convenient (and are always pretty to extremely expensive), they quickly becomes an integral part of the customer's badly architected "system".
The end result is complete vendor-lockin, the inability of the poor (stupid) user to leave and the continued gang rape of their bank account (also via additional, incompetent developer and devops "resources").
Throw in average modern "devops" who are hired to handle this. They aren't like the sysadmin of yesteryear, they no longer have experience with, or understand the bits and bytes. They are glorified UI clickers and YAML editors, they even lack any reasonable system level debugging skills. For every problem they encounter they first immediately run to google in search for answers.
In addition, I would argue that CSPs are a huge, huge waste of computing, space and power resources, because their systems completely encourage people to just do things, without understanding what they are doing, screw the consequences and just pay.
Result, the business suffers greatly (on so many levels), the CSP wins big and continues winning.
What happens here is that a system, if designed right from the get go, could have been run on a SINGLE, modern, high end, well positioned and connected server to the Internet, is now replaced with tens to hundreds of "instances" and random assorted CSP provided services -- what a colossal waste.
Books can be written on negligence, lack of understanding, utter tech stupidity and ultimately the costs which are absurd.
It's also a huge waste of human effort managing the complexity introduced by the cloud provider's arbitrary bullshit.
At this point multiple generations of engineers have little understanding of underlying layers of technology, having only really learned how to use cloud services. No TCP/IP, no UNIX, just a bit of bash and a ton of AWS.
Cloud providers do hide most of the low level complexity, which could be seen as a benefit (at least that seems to be what's touted as a main benefit, along with instant scalability.) Unfortunately they replace all of that with more arbitrary complexity which is ultimately (in my opinion, at least) a much bigger burden than the fundamental complexity that is abstracted away.
Of course you can (like Snap) but it's a MASSIVE engineering effort and initial expense.
Amazon uses $/gb as a price gouging mechanism and also a QoS constraint. Every bit you send through their pipe is basically printing money for them, but they don't want to give you a reserved fraction of the pipe because then other people can't push their bits through that fraction. So they get the most efficient utilization by charging for the stuff you send through it, ripping everybody off equally.
Also, this way it's not cost effective to build a competitor to Amazon (or any bandwidth intensive business like a CDN or VPN) on top of Amazon itself. You fundamentally need to charge more by adding a layer of virtualization, which means "PaaS" companies built on Amazon are never a threat to AWS and actually symbiotically grow the revenue of the ecosystem by passing the price gouging onto their own customers.
It is pretty much exclusively to lock you into their services. It heavily impacts multi-cloud and outside of AWS service decisions when your data lives in AWS and is taxed at 5-9 cents a GB to come out. We have settled for inferior AWS solutions at times because the cost of moving things out is prohibitive (IE AWS Backup vs other providers)
To them "that's just what bandwidth costs" but anyone who's worked with this stuff (sounds like you and I both) can do the quick math and see what kind of money printing machine this scheme is.
Some people want to host a lot of warez and pirate movies and stuff but that doesn't monetize very well per GB consumed so pricing bandwidth high means those people never show up, thus saving a lot of trouble for AWS.
I remember when salesforce.com announced a service that would let you serve up web pages out of their database, it was priced crazy high (100-1000x too much) from the viewpoint of "I want to run a blog on this service" but for someone who wanted to put in a form to collect data from customers it was totally affordable. Salesforce knew which customers it wanted and priced accordingly.
The issue isn’t charging for egress, but charging excessively.
First a lot of these roads are 'free' and yet you're still being charged for it. If two large networks come to an agreement then they connect the two networks (ie build that road), but no money changes hands.
Second if there is a paid peering agreement in place (ie say AWS had a cost to push your data out), that still wouldn't be billed to them in the way they're charging you. Instead they'd be paying for the rate of traffic at something like the 95th percentile of the max. This means that you could download a petabyte of data from them when the pipe isn't busy and cost them nothing, or you could download a gigabyte when it's busy and push up the costs.
I almost wish they had some kind of sustainable usage-based charge that was much lower than AWS.
Feel free to tell me why I’m wrong! I’d love to jump onboard - it just seems too good to be true in the long-term.
There probably needs to be an abuse prevention rate limit (and probably is), but it's not quite as crazy as it sounds to just rely on their CDN bandwidth sharing policies instead of charging.
I do think there are “soft limits” in place like you say - it’s just my personal preference to have documented limits (or pay fairly for what you use). IMO it helps stop abuse, and prevents billing surprises for legitimate heavy use-cases.
But that's really no different from the guarantee you get from most CDN services. If you're using cloudflare in front of S3, for example, you'll end up with the same behavior.
But in my mind it’s also comforting that something like Cloudfront has a long-term sustainable model (I should also add with fewer strings attached like hosting video).
I do think the prices ant AWS are too high, but it discourages bad actors from filling up the shared pipes. ISPs are sometimes a classic example of what happens when a link is over subscribed.
Cloudflare’s “soft limits” are also somewhat of a dark pattern if you ask me. I like to know exactly how much something will cost, and it’s really hard to figure out with Cloudflare if you’re a high-traffic source. Do I hit the “soft limits,” or not? It’s really hard to say with their current model.
FWIW, I think Cloudflare is a great product right now - I am just skeptical they can keep it up forever.
The original post also includes a link to a more recent Cloudflare blog post on AWS bandwidth charges: https://blog.cloudflare.com/aws-egregious-egress/
And, if you're in the same boat as someone down-thread complaining about Cloudflare's uptime in recent weeks, you can keep S3 + Cloudfront (or Lightsail Buckets + Lightsail CDN) and S3 + CacheReserve on R2, which is what we do, and flip between them with DNS.
> let’s remember that the internet is 1-to-many. If 1 million people download that 1GB this month, my cost with @cloudflare R2 this way rounds up to 13¢. With @awscloud S3 it’s $59,247.52.
Maybe abuse isn't the right word but definitely making the most.
I am a bit scared about being turned off overnight though.
This translates to roughly 1.2 gigabytes per second (every second of of the month), and 3240 terabytes of data per month - in or out, the choice is yours.
Things scale down as you buy more bandwidth, or commit to a longer contract.
Many would say that $1000 per month is literally "nothing" in terms of costs of service for most real businesses our there, and if you're a happy CSP user, you're probably paying a hell of a lot more than that per month for your infra.
Cloudflare doesn't pay for egress and neither does AWS.
The greatest trick AWS ever pulled was convincing the world you needed to pay for bandwidth.
Not saying not to trust him - he’s probably a very reasonable and standup guy - but you should know this about him before taking his word on a topic like this.
No disrespect @eastDakota
Yes, but every customer would never do this, so what is your point?
You have to think more in terms of averages for things like this
Which is never going to happen for legitimate use cases.
And Cloudflare has DDOS protection for ilegitimate ones.
Though not everyone is Netflix.
I feel R2 should charge something for transfer though, otherwise people could abuse it. Hetzner charges ~1.5% of AWS egress fees which I feel is right thing to do and likely profitable.
I'm also a Backblaze B2 customer, which I also highly recommend and has slightly different trade-offs (R2 is slightly faster in my experience, but B2 is 2-3x cheaper storage, so I use it mostly for backups other files that I'm likely to store a long time).
What could you have a petabyte of that you're pretty sure you'll never need again? What kind of datasets are you storing?
It doesn't have to be nearly that stark.
If we factor out egress, since it's the same for everything, the bulk retrieval cost for glacier deep archive is only $2.50/TB.
That means that a full year of storage ($12) plus four retrievals ($10) is roughly the same price as a single month of normal S3 storage ($23).
Plenty of other people storing images, video, etc. a PB is really not that much stuff when it's not just for personal consumption.
2. R2 doesn't support file versioning like S3. As I understand it, Wasabi supports it.
3. R2's storage pricing is designed for frequently accessed files. They charge a flat $0.015 per GB-month stored. This is a lot cheaper than S3 Standard standard pricing ($0.023 per GB-month), but more expensive than Glacier and marginally more expensive than S3 Standard - Infrequent Access. Wasabi is even cheaper at $0.0068 per GB-month but with a 1 TB billing minimum.
4. If you want public access to the files in your S3 bucket using your own domain name, you can create a CNAME record with whatever DNS provider you use. With R2 you cannot use a custom domain unless the domain is set up in Cloudflare. I had to register a new domain name for this purpose since I could not switch DNS providers for something like this.
5. If you care about the geographical region your data is stored in, AWS has way more options. At a previous job I needed to control the specific US state my data was in, which is easy to do in AWS if there is an AWS Region there. In contrast R2 and Wasabi both have few options. R2 has a "Jurisdictional Restriction" feature in Beta right now to restrict data to a specific legal jurisdiction, but they only support EU right now. Not helpful if you need your data to be stored in Brazil or something.
I do have to wonder if that leaves R2 customers one minor compromise away from losing their whole data store.
If you already use AWS for lots of other things, yes.
Every cloud provider has outages sometimes but CF has been horrendous.
We were actually planning on migrating some other parts to R2 but we are just ditching CF altogether and just going to pay a bit more on AWS for reliability.
So if R2 has been impacted even a third as much as CF images, that would definitely be an important consideration.
I found https://isdown.app/integrations/cloudflare/cloudflare-sites-...
That said we don't use any queues, KV, etc. Just pure JS isolates so that probably contributes to the robustness.
We do use the Cache API though and have ran into weirdness there. We also needed to implement our own Stale-While-Revalidate (SWR) because CF still refuses to implement this properly.
Overall CF is a provider that I would say we begrudging acknowledge as good. Stuff like the SWR thing can be really frustrating but overall reliability and performance are much better since moving to CF.
I don't understand. You say that you used a very small subset of their offering in a very specific and limited way; and with that you conclude that their offering is "good"? Shouldn't you make that conclusion after reviewing at least 50% of their offering?
And they won't increase it unless you become an enterprise customer in which case they'll generously double it.
If you don't mind having your bits reside elsewhere, Backblaze B2 and Bunny.net single location storage are both cheaper than Cloudflare.
Another (that probably contributes directly to the write latency issues) is region selection and replication. S3 just offers a ton more control here. I have a bunch of S3 buckets replicating async across regions around the world to enable fast writes everywhere (my use case can tolerate eventual consistency here). R2 still seems very light on region selection and replication options. Kinda disappointed since they're supposed to be _the_ edge company.
Same with data that is aggregated into smaller data set within AWS before you egress it.
Otherwise, I've been using R2 now in production for wakatime.com for almost a month now with Sippy enabled. The latency and error rates are the same as S3, with DigitalOcean having slightly higher latency and error rates.
1. No Object history and locking. So there is absolutely no way to recover files when you do any kinds of mistakes.
2. No object tiering and storage is not that cheap. Although R2 egress is free, R2 is only 35% cheaper than S3 in terms of storage, but it is not cheaper than other alternatives. Furthermore, R2 is a lot more expensive than S3 infrequent/cold tier.
For example, Backblaze B2 is 4 times cheaper than S3, and B2 offers history/locking. When B2 egress is free up to 3x monthly storage, B2 is much better option than R2 for most cases if a considerably high egress is not needed.
The last time I benchmarked B2 was years ago but it wasn't as reliable as I wanted at getting me files in under two seconds.
That egress money was going to be spent with or without sippy. It's not "just spreading" the pain, it's avoiding adding any pain at all.
Details: https://twitter.com/DatabendLabs/status/1719580350677237987
It currently feels a little limited and… bolted on to the Cloudflare UI.
https://twitter.com/tomlarkworthy/status/1711846776905293967...
Next up: moving our cloud edge (NAT Gateways, WAF, etc) to Fortinet appliances, which licenses we purchased bundled with our on-prem infra.
I know Corey Quinn always harps on AWS' egress pricing but you really can't emphasize it enough: it's literally extortionary!
At small data scale this adds up.
And..... it's 11 cents a GB from Australia and 15 cents a GB from Brazil.
If you have S3 facing the Internet a hacker can bankrupt your company in minutes with simple load testing application. Not even a hacker. A bug in a web page could do the same thing.
(Assuming your company can be bankrupted for ~$20k.)
If you are having lots of egress: R2 is the cheapest (15$/TB/month, free egress)
R2 can get somewhat expensive if you have lots of mutations, which is not a typical use case for most.
You can proxy things through cloudflare and get unlimited free egress thanks to bandwidth alliance between cf and b2
Interesting. What sort of companies can take advantage of this?
I work for a non-profit doing digital preservation for a number of universities in the US. We store huge amounts of data in S3, Glacier and Wasabi, and provide services and workflows to help depositors comply with legal requirements, access controls, provable data integrity, archival best practices, etc.
There are some for-profits in this space as well. It's not a huge or highly profitable space, but I do think there are other business opportunities out there where organizations want to store geographically distributed copies of their data (for safety) and run that data through processing pipelines.
The trick, of course, is to identify which organizations have a similar set of needs and then build that. In our case, we've spent a lot of time working around data access costs, and there are some cases where we just can't avoid them. They can really be considerable when you're working with large data sets, and if you can solve the problem of data transfer costs from the get-go, you'll be way ahead of many existing services built on S3 and Glacier.
Basically, using R2 allows you to undercut competitors' pricing. It also means I don't need to build out a separate CDN to host my files, because Cloudflare will do that for me, too.
Competitors built out and maintain their own equivalent CDNs and storage solutions that are more ~10x more expensive to maintain and operate than going through Cloudflare. Basically, Cloudflare is doing to CDNs and storage what AWS and friends did to compute.
But reality is a bit more complicated than that. Migrating data + pointers to that data, en masse, isn't super easy (although things like Sippy make it easier).
In addition, there's all the capex that's gone into building systems around the assumptions of their blend data centers, homegrown CDNs, mix of storage systems. There's a sunk cost fallacy at play, as well as the inertia of knowing how to maintain the old system and not having any experience with the new system.
It's not impossible, but it'd require a lot of willpower and energy that these companies (who are 10+ years into their life cycles) don't really possess.
Having seen the inside of orgs like that before, starting from scratch is ~10x-100x easier, depending on the blend of bureaucracy on the menu.
And the difference is that you will fail your customers when that time comes because you'll just get suspended (we've seen some cases here on the forum) and you'll have to come here to complain so the ceo/cto resumes things for you.
In their docs they explicitly state it as an attractive feature to leverage, so that’d surprise me.
That being said, I’m not planning to serve particularly large files with any meaningful frequency, so in my particular case I’m not concerned about that possibility. (I’m distributing low bitrate audio, and small images, mostly).
If I were trying to build YouTube or whatever I’d be more concerned.
That being said, with their storage pricing and network set up as they are, I think they make plenty of money off of a hypothetical YouTube clone.
I do think they’ll raise prices eventually. But it’s a highly competitive space, so it feels like there’s a stable ceiling.
> I’m distributing low bitrate audio, and small images, mostly
This means the cache-size would be much smaller though.
Re cache-size, maybe I've misunderstood what you mean by cache size limiting, but yeah that's my point – I don't need a massive cache size for my application. My data doesn't lend itself much to large and distributed spikes. Egress is spiky, but centralized to a few files at a time. e.g. if there were to be a single day where 1TB were downloaded at once, 80% of it would be concentrated into ~20 400MB-sized files.
> They were also seeing ~30TB of daily egress on a non-enterprise plan, which would absolutely never happen in my case – 1TB of daily egress would be a p99.9 event.
I don't understand what media company you'll be competing against if you'll use just 30TB/month of bandwidth.
Minio to seaweedfs around 2020 because our minio servers had problems serving very large number of small files.
Then this year we migrate to B2 because it's way cheaper and we don't have to rewrite our apps.
Still my hat goes to S3. It is so massive that every open source or competitors need to have compatible API and it give us the ability to move to any vendor or selfhost just by changing endpoint.
The real issue is how that data get's into S3 in the first place and what else you need to do with it.
S3 and DynamoDB are the real moats for AWS.
You always had available plenty of space on dedicated servers for way cheaper before the cloud.
You could make an argument about the API being nicer than dealing with a linux server - but is AWS nice? I think it's pretty awful and requires tons of (different, specific, non transferable) knowledge.
Hype, scalability buzzwords thrown around by startups with 1000 users and 1M contract with AWS.
Sure R2 is cheaper but it's still not a low cost option. You are paying for a nice shiny service.
It's certainly quite cheap for a set of typical "requirements" for media hosting companies.
But yeah, if you're storing data for mainly archival purposes, you shouldn't be paying for R2 or S3.
This would be very useful in genomics, where pretty much everything is stored on S3 but always a pain to connect to apps.
Although I do wonder if that would be considered a bait and switch.
Compare with say Oracle cloud which tries to compete by having 1/10th the egress charge. But nobody uses it anyway and they DO offer all the other services.
My use case is image storage + serving for a service that users will upload a lot of images to. Currently using Cloudflare + storing all files on disk but space will soon become a concern.
My assumption is that "at least three" means "exactly three" in practice.
Perhaps i’m too hasty with my judgement, hope so….
Cloudflare may well be on their way to becoming a monopoly, but they certainly show they don't care about abuse. Even if it weren't a simple matter of principle, in case they aren't successful in forcing themselves down everyone's throats, I wouldn't want to host anything on any service that hosts phishers and scammers without even a modicum of concern.
AWS has zero interest in S3’s API being a universal standard for blob storage and you can tell from its design. What happens in practice is that everybody (including R2) implements some subset of the S3 API, so everyone ends up with a jagged API surface where developers can use a standard API library but then have to refer to docs of each S3-compatible vendor to see figure out whether the subset of the S3 API you need will be compatible with different vendors.
This makes it harder than it needs to be to make vendor-agnostic open source projects that are backed by blob storage, which would otherwise be an excellent lowest-common-denominator storage option.
Blob storage is the most underused cloud tech IMHO largely because of the lack of a standard blob storage API. Cloudflare is in the rare position where you have a fantastic S3 alternative that people love, and you would be doing the industry a huge service by standardizing the API.
For example, "Can one append to an existing blob/resume an upload?" leads to lots of questions about data immutability, cacheability of blobs, etc.
"What happens if two things are uploaded with the same name at the same time" leads into data models, mastership/eventual consistency, etc.
Basically, these 'little' differences are in fact huge differences on the inside, and fixing them probably involves a total redesign.
Heck, HTTP already provides verbs that would cover this, it would just require a vendor to carve out a subset of HTTP that a standard-compliant server would support, plus standardize an auth/signing mechanism.