How Canva saves Amazon S3 costs
canva.dev
canva.dev
Tens of millions wasted.
Canva is likely trapped in S3 never to exit. The cost of getting their data out makes it impossible.
S3…. the Hotel California of the cloud. You can check in any time you like but you can never leave.
S3 is 9 cents per GB egress fees.
Cloudflare R2 charges zero egress fees.
At list price of 9 cents per GB it would cost Canva about USD$20 million to export their 230 petabytes. Let me know if my calculation is wrong.
We realized that instead of using scaling for peaks, by having additional dedicated servers available full time - it still worked out significantly cheaper. We also moved a lot of our internal processing to run in specific windows of time where the expected load was low to maximize server utilization in those periods.
You only save money if the differential cost based on all the factors is lower. AWS is expensive for some things so it often does save some money, but not always, and if you haven't done a proper analysis you can't know.
The cloud vs. on-prem argument often seems to ignore the (enormous) middle-ground. Just because one portion of your architecture would do well to be run outside the cloud, doesn't mean you take out all the other parts you don't want to deal with yourself. Furthmore, "on-prem" might mean in your building, in someone else's building co-located and you rent space and control it, or in someone else's building where they deal with most hardware, etc.
That said, maybe Canva has considered a non-AWS solution and decided against it. Or maybe they've gotten certain better deals from AWS. We can't really know for sure on the outside.
- Consistency. Instead of saying "Oh look in AWS for this app, Azure for this one, and hetzner for this app, except it's test env is in AWS", it all just lives in AWS. It massively simplifies docs, onboarding, and reduces the amount of one-person specialised knowledge.
- Engineering Costs. Similar to above, but in terms of engineering, there's less to know and understand. Instead of needing to know how the AWS load balancer routes/connects to a VM somewhere else, and how that VM gets it's blob-storage-data from azure, we only need to understand AWS concepts.
- Vendor Lock In. Yeah, it's there. If we have a service that uses data from S3, there's egress costs from S3 to <other provider>, but not with EC2. We've consciously accepted this lock in for the time being.
Now, we're a 50 person company so YMMV, but the above tradeoffs plus an "opinionated" setup in AWS (everything on ECS, logging to Cloudwatch, RDS for DB) drastically reduced the "ops" overhead on our side after the initial setup. If I started over, I'd make the same decisions again.
Okay? I never said that sticking to a single cloud provider isn't appropriate for some (or maybe even most) people. It's good that you have a setup that you believe works well for you.
Very few people should really be using S3 at any serious scale is my thoughts. The cost savings are enormous (plus cloudflare for example replicates your data a lot closer to users for no extra cost, significantly improving performance). The cost savings can be absolutely enormous for very little/no additional complexity given how many providers are compatible with S3, and the fairly 'boring' nature of S3 compared to other technologies.
Because I've done small scale and can tell you I'd run S3 in the future.
This is where I think the FCC should take action.
To the extent that this issue is a mutually agreeable arrangement between you and Amazon, it seems obnoxious but does not seem like it rises to the level where regulators should take action. But it affects third parties too: specifically, it prevents non-AWS-hosted vendors from effectively marketing their services to you. In that regard, I think the FTC should try to put a stop to this. AWS should not be permitted to effectively subsidize its and its partners’ services over outside competitors.
(And the US Government should never have accepted cloud deals with excessive egress costs. Part of the bidding process should have been a requirement for networking outside the winning provider to be priced competitively with internal networking)
and easily lower costs by 50% depending on balance storage/bandwidth
This is the traditional “sell” for cloud computing and I don’t buy it at all.
The clouds would have you believe it doesn't make sense to run your own systems because you need too much specialist expertise to run your own systems.
The clouds sales pitch is the if you go cloud then you don’t need all these specialists.
That’s rubbish. Cloud operations need the same or more headcount if technical specialists, they’re just doing different things.
The old “don’t run your own systems, it’s cheaper and easier to go cloud” is just sales fiction.
Don’t believe it.
You don't need to 'believe'. You need to do an analysis of what you need and how much each option will cost. Then you can know.
If your argument is "I believe onprem saves money" or "I bet AWS is cheaper" or "Jim Morrison came to me in a dream and said I should use Azure" then you haven't done enough research.
the research isn't free either. The more indepth and time it takes to do such research, the slower you come to a decision and ship.
Cloud allows you to ship fast. It allows you to go without research - just accept the marketing, and pay up.
You pay above-cost (compared to on-prem) when you grow to a certain size. But this is usually a worthy trade off tbh.
Even the same number of "technical specialists" doesn't mean it's the same if one option lets you move faster or remain more reliable.
To be fair most businesses probably don't need it, but it is worth taking into account.
However, as someone put it, put your money where your mouth is. Become a contractor offering to reduce storage costs in exchange for 10% of the savings. You should be a millionaire in no time, according to your narrative.
It's a capex vs opex question. For financial quackery reasons, accountants and stock markets prefer having less people on the direct payroll, which is why everything not part of the "core business" is outsourced - even if it is more expensive in either short or long term.
Cloud makes a ton of sense for high growth companies. If your infrastructure is mostly static, that's when it makes sense to go to a datacenter. Or if you have one very specific use case, like Dropbox.
Do you recall what the big savings were in? Compute, storage, DB etc? Be really interested to understand why this case is so different to many I've read.
Data Center (per month)
Servers: $6K
Cabinet (x3): $15K
Bandwidth: $2.5K
Support: N/A
Total: $23.5K
EC2 (per month)
Servers: $13K
Storage: $1.5K
Bandwidth: $1.1K
Support: $1.2K
Total: $16.8K
At the time we and Heroku were running the two largest Postgres clusters on EC2.
This is an important factor to remember when evaluating could costs. If you want to survive a cloud outage, you need a multi-AZ or multi-region deployment, and that costs developer hours. And you need to deal with the potentially catastrophic cost of inter-AZ or inter-region traffic, which can be catastrophic and/or cost more developer hours to mitigate.
RDS databases are one click to set up in multi-AZ, and if your stack is Kubernetes based or at least EC2 autoscale capable it isn't much more work to make it multi-AZ as well.
Multi-Region deployments however, these are indeed expensive and nasty to set up.
Of all the possible related problems, object storage is the easiest.
More of a "good problem to have" (and you can solve it when you need to).
Also, it seams that it's 230PB in total, not per month, which fits in an apartment-sized server room (you want to have several of such places, but it's not that big).
Then there was a contract which had a server room twice the size and mostly contained an AS/3<mumble> and a couple racks of Intel hardware. You could probably get a ping pong table into that one. Giant rooms that are mostly white and too cold are a little freaky, and reminiscent of 2001.
There have been a number of stories of people not being able to utilize their floor space for more servers because they ran out of space on the roof for AC units. At one point I recall Google was working on power dissipation because they had hit the max electrical code, and so no one would agree to run more power into their buildings. More servers meant more compute per watt and per ton of chillers.
That's obviously not something you can do on a whim, but at the same time we're talking about a seven figures investment plus the time to build the in-house team to manage this, so it's not going to happen overnight anyway.
> - System hardware: $500,000
For one PB? Gold plated HDDs and cases I guess…
> Assume 15% annual system maintenance over five years
Do not run your hard drive cluster on the same floor as vibration-inducing machines, folks. /s
According to Backblaze, even a 8 year-old hard drive is still twice more reliable than this, and 15% failure rate is what to expect over 6 years of consecutive use!
> Assume the space needed to store 1PB is between 25%-50% of a standard rack (42U)
You can fit almost 10PB in a rack though, so that's another figure that they've inflated.
> Assume personnel costs to manage a 1PB system are 0.5 IT/storage admin FTE
This one is probably true (and it's even an underestimation) if you have only 1PB of data, but it's not going to scale linearly: you're never going to need 130 FTE for 260PB.
Overall, if a supplier (who's going to use the same kind of off-the-shelf tech as you would) claims that his price is 4 times cheaper that what it would cost you to run it by yourself, he's probably just lying to sell you stuff, just saying…
Oct 2020, StorONE announced a 1PB All-Flash array for $499k.
Other examples which are close to their pricing. https://www.reddit.com/r/storage/comments/vdr6ql/pricing_exa...
No doubt about that, when you have tens of thousands of drive you end up replacing some of them all the time, but not 15% of them every year…
> Oct 2020, StorONE announced a 1PB All-Flash array for $499k.
With Optane and all, right? The kind of things that's a massive overkill if your workload allows you to use Glacier…
> Other examples which are close to their pricing.
In that list, the only offers that are close to this pricing (but still almost 20% cheaper) also includes several years of support (which you probably don't want if you're Canva's scale).
mind you the price of flash storage has dropped significantly since 2020.
Hard disks are still being used, but having all flash systems is not uncommon.
- if they had a sunk cost CAPEX of 100 PB of storage, all they could do is sell the storage servers for a fraction of the original cost (and folks like me snapping up pre-owned server equipment for said fraction of cost for our home labs).
- on the other, since S3 is OPEX (not sure if they are locked in to 1 year billing cycles? but still better than 3-5 year CAPEX runs), they could cut the cost of that storage "immediately".
Similarly, if they needed to expand, CAPEX means placing an order for the storage, waiting for delivery, racking, burning in etc, versus just paying almost "instantly" for more storage.
So long as they are still earning money, why wouldn't they stick to this OPEX model?
Not saying this is good or bad, it depends on your business model, ultimately.
The suggestion was to move to a different / cheaper provider. You can stick to OPEX.
The scarier part is Canva actually does use Cloudflare CDN, so it's not so much the cost of adding a new provider. It's just that the CDN team that manages the Cloudflare account is likely a different team :)
Too much isolation, silos and politics...
Cloudflare has been reliable on storage adjacent uses (caching) for a long time. Presumably you're more likely to be the cause of errors than R2 is by a few orders of magnitude.
You can't really check the SLA as being realistic
Edit: Note also that Snowball egress is $0.03/GB. Slightly higher egress, much lower setup cost. You'll have to do the math but they're both clearly attractive options vs. full price $0.09/GB egress.
Someone else in the comments will correct you with the correct answer.
This in effect.
Same if someone is calling out all of the general problems with solving a problem but not providing answers where you can't get them to bite.
See the Docker x Cloudflare case study where improving cache-hit ratio by 2% (by moving to R2) decreased S3 egress fee by 66%: https://www.cloudflare.com/en/case-studies/docker/
Which in 99.9% of the cases is the right decision. Sorry for all you infra techies, but that's the way tech matures.
YMMV.
Directly, yes maybe. Many companies don't have in-house experts (and get shafted by vendors as a result, as they're lacking the competence required to assess the quality of work).
But indirectly, they all have. In Germany, at least, you're required to have electrical appliances inspected by a licensed professional (an electrician or otherwise qualified person) at least every two years - most companies opt for using dedicated external companies, but you can also train someone for that task. On top of that come all the electricians hired by the landlords - the ones dealing with complaints and remodels, or servicing lifts, escalators, HVAC, datacenters, ...
Electricity as a Service, which includes payments for expert maintenance.
Thread is about AWS and merits of in-sourcing, which is not how the world operates.
Anyways.
As long as dev's don't seem to understand real world limitations, (for instance, a network cannot be instantanious, and latency is not consistent) and thus leave tons of performance on the table, i think the "infra guys" will be just fine.
Also, don't forget there is an entire world out there of people making sure connections to the cloud are even possible across the internet.
How did they get so much data?! They say they only have 75 million stock photos. They must have tens of petabytes that are just dark.
If you're doing a one-time export you can use a Snowball, which would cost about $70K for 230PB, assuming you had somewhere to load it off to once you got the data so you could give the device back.
- Since non transactional data is likely to be on GitHub or somewhere, it should be easy to re-hydrate that on R2.
- Once enough time passes, there could be a sunset on unused data, or a strategy planned to 1-time exit write-off for S3 with those assets not on R2.
Standard migration procedure.
Apparently, it would be easy for them to save tens of millions, in turn making millions for the consultant.
One only says this if they don't have a good understanding of the differences between capex & opex, the amount of money spent on tech labor & compliance projects, as well as the level of assurances AWS give you regarding durability of data. All of this in 2023 when we have almost 20 years of data showing more and more business moving to the cloud and staying in the cloud.
I think anti-cloud sentiment aligns with hacker mentality on decentralization=good (which it is), anti big corporation feelings (boo amazon) and so it leads people in these threads to make emotional arguments over something that is clearly going in the other direction. It's an excellent example of confirmation bias, trying to look for any and all justification to say the cloud is worse, when the majority of businesses have decided otherwise.
Step 2, write articles like this as part of your newly contracted "engagement" with Amazon to help justify the lower price.
So pay Amazon a bunch for the privilege to have lower prices. How about just starting off somewhere where you don't have to commit to paying a lot of money to get a reasonable price?
Canva grew very quickly. I don't think it was a terrible call to use the cloud at the time.
Cloud became a fad a while ago, companies like Amazon/Google/MSFT rushed in burning cash to provide incentives. Now said companies adjust fees to milk businesses due to high exit cost, just because they can. This will inevitably create outflow of customers, however we are far from this point currently.
Even if you don't use any other cloud tech in your stack, my advice is, use S3 (or equivalent)! That level of durability, availability, and scalability, is not trivial to pull off, and is best made somebody else's problem. Even when you're as big as Canva.
FWIW, I have an EU region cluster (3 nodes per inner-region in the North, West and South). It’s all one logical region to S3 libraries and holds nearly half a terabyte.
It’s probably one of the most stable parts of my infrastructure. I don’t know (or need to know) what the replication lag is between regions, but IIRC, there’s some configuration of consistency.
[1]: https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecy...
More seriously, I think the sweet spot is to buy 2x o 3x what you expect to need but do it off the cloud.
At Canva's scale you'll spend hundreds of thousands, save literal millions and have way more headroom and stability than with crappy auctioned instances on AWS.
Desperately trying to work out if they can delete this resource or that because who the hell knows which bit of critical code is using it. Or spending days trying to manage a chewing gum ball of IAM policies.
> Desperately trying to work out if they can delete this resource or that because who the hell knows which bit of critical code is using it.
How is this unique to the cloud?
The point is that you will spend millions anyway. Might as well have something better than whatever home-grown dumpster fire your team will come up with.
You also seem very ignorant of how AWS works. I've never had to wonder if I could delete a resource or not, because all our resources are created via CloudFormation. Same for IAM policies.
At this point even the not-infrequent public cloud outages are essentially shrugged off with "Eh, we use $CLOUD, that's what everyone uses. What are you going to do?"
The ops team can actually say "It's Amazon, not our problem." while they sit around powerless to do anything about it.
Yes there is the "survive anything AWS related" geo-distributed multi-region approach that you pretty rarely even see because it itself requires an army of AWS experts and drastically increased cost to actually implement and operate while not shooting yourself in the foot left and right - actually decreasing reliability. Not to mention the uber-unicorn "multi cloud" which at least I've never seen or actually heard of in practice. These don't usually survive even an initial cost review.
and some of them just did. Both know names and small shops. Big names like cloud-flare R2 or Backblaze b2.
Having been a part of these negotiations for services other than S3, DTO discounts can be generous if you have a good forecast (3-5 years) on your data egress.
Did not think to try the first link though.
There's tons of OSS to manage all layers of this stuff now.
50-ish drives in a single 4U server is done.
or, if you go fibre channel, just replicate the writes across multiple storage systems and be done with it. (this requires a fibre channel network though, which is neither easy nor cheap).
RAID5 gives you really the worst of all worlds. You get poor performance (n/4) due to the number of operations required and you get the fragility of a RAID storm that comes from a cascade RAID failure when a drive dies. A "stripe of mirrors" (RAID10) gives you the safest and the highest performance array (at scales where these decisions are being made which most people would see anyway). The common argument against is the cost (usableCapacity = totalCapacity/2) but my counter argument is that RAID5 is a timebomb with a secret countdown. RAID5 is great if you don't need the data on the array.
Source: >20 years running complex large arrays
But I do agree with you, it does make the problem significantly easier/different.
ZFS and Ceph are marvelous creations.
What are your thoughts on implementing RAID 0 over four RAID 1 arrays, each consisting of three disks, resulting in a total of 12 drives?
S3 keeps multiple copies in multiple geographic locations.
Having more locations is actually helpful for this. Because of the independence of different locations, the risk of data loss is in some ways lower than a local RAID system where all drives might be fried at the same time by a single event.
From there you install Hadoop, maybe a query tool like Spark, Impala, and Presto. Then you install some reporting tools like Apache SuperSet.
At least that's how we did it at my old job with Petabytes of data.
List price of Backblaze:
$0.005/GB - Base Storage
$0.01/GB - Egress (Use a data partner to bring this down to $0)
$0 - PUT API Requests
$0.004/10,000 - GET API Requests
No delete penalties a/k/a minimum storage
List price of Glacier Instant Retrieval:
$0.004/GB - Base Storage
$0.09/GB - Egress
$0.02/1,000 - PUT API Requests
$0.01/1,000 - GET API Requests
Billed for 90 days of storage
How may of those countries are profitable? My sense is Internationlization is a relatively fixed cost, but how does customer support in Burma, for example, not trump any realizable revenues there?
There's quite a few anti-cloud opinions in the comments but for anyone who is hosted on S3 and would like this same level of automated analysis for cost savings on S3 (and other AWS services), we basically profile for this out-of-the-box on Vantage and it's free to get an optimization check to see savings: https://www.vantage.sh/
Not meant to be as shameless of a plug, but seems very relevant to the topic at hand for onlookers.
Every single day I read the trope on HN that "premature optimization is the root of all evil".
So, these guys literally followed HN's creed.
Oh, my. That's a lot. What does Figma? 300?