Presumably the author would do much better with a VM or something from OVH, they'll just shut you off or limit you before it becomes a problem (not that they would care about 30 TiB).
Presumably the author would do much better with a VM or something from OVH, they'll just shut you off or limit you before it becomes a problem (not that they would care about 30 TiB).
For someone coming from my perspective, that would be a huge and unpleasant surprise to be billed $2600 for a service that I know costs $180 elsewhere.
Edit: especially so because people keep repeating the mantra that one should use the cloud to save money by not paying for what you don't use. Obviously, my 3 Xeon servers with 64 GB RAM each are way overpowered for sending out some GB of static files, but I wanted to have a bit of redundancy. But with my setup, there should be plenty of obvious inefficiencies for "the cloud" to eliminate
=> It feels like cloud should be cheaper than my dedicated servers. But it's not, and that is the unpleasant surprise.
But it's totally unfair to say they're "overcharging". The pricing is set up to encourage using the service properly. For example, did you know you get unlimited free (very fast) transfer between s3 <-> ec2?
You are free to use S3 however you want, but the pricing is set up such that people use a proper cdn backed by s3, do as much as they can between ec2 <-> s3, etc. instead of making s3 the backbone of their public site.
If you want to use it in a way it's not intended it will cost you more $$ which is how it should be.
Amazon's systems should have notified at least twice and required direct confirmation from the customer for their continued use.
We don't just make business decisions like this based off our self-interest. In this case, a failure to be resourceful on Amazon's part gave the customer one of two options: track everything manually minute-by-minute, or be surprised by an excessive bill.
I can actually think of possible/real workloads that would burn 2700 usd out of blue in two days and those customers would not be happy if AWS blocks their account because AWS thinks they did something wrong.
AWS Support quite often happily refunds these amounts and you don't need to post to Twitter. I've seen them refund a lot more.
Why haven't they done something to prevent this mistakes? They already refund these things, it's just those refunds is a drop in bucket compared to services that enterprise customers are paying for, and AWS engineers are busy building for them.
AWS gives you a lot of power compared to what you can do on your average hosting provider, but sadly there's also a lot of room to shoot yourself in the foot if you don't know what you are doing.
Citation needed.
Network throughput is a finite resource. And spikes from noisy neighbors could absolutely degrade throughput performance for everyone else.
So it makes perfect sense for me that part of Amazon's pricing calculus would be based not just on the cost to deliver the feature but the amount of utilization of the overall network they expect for the use cases their pricing supports.
That's a very charitable take. A more cynical take is that they do it to encourage lock-in. If you already have your data stored on s3, they can get away with overcharging on compute because they know that if you switched to GCP, you'll end up paying more because of the egress costs.
>but the pricing is set up such that people use a proper cdn backed by s3, do as much as they can between ec2 <-> s3, etc. instead of making s3 the backbone of their public site.
I'm not sure what your point here is. Cloudfront costs 8.5 cents per GB of egress, meanwhile ec2 costs 9 cents per GB of egress. half a cent doesn't make the egress charges any less outrageous. That's not even including the per-request charges that cloudfront has, which probably can make it cost more than 9c/GB depending on your usage profile.
Do you believe that little gem applies to all products and services, or do you reserve it to downplay AWS's notorious price gouging?
The most common services like S3, Kafka, RDS (pg and mysql), Redis should be enough to cover most use cases.
With k8s and a dedicated smart team, it would be a possible adventure.
Moreover it can work great also on-premises. Several old medium business have their own physical infrastructure and they are not yet ready to move to the cloud.
Which is what AWS is anyway...? Also, I don't think AWS' data-centers exist in isolation: https://www.datacenterdynamics.com/en/news/wikileaks-publish... Definitely not the Cloudfront PoPs which ought to be leased, surely?
Also see: https://www.tatacommunications.com/solutions/cloud/infra-dc/
And: https://www.vapor.io/vapor-io-partners-with-cloudflare-on-na...
I worked on a project a few years back where the client were paying $10k/month as a "managed service" fee, on top of about $4k/month worth of platform at full on-demand prices. I showed the client how they could have it all running on reserved instances for Prod and spot instances for dev/staging for under $2k/month - but no, somebody had signed up for $14k+ per month just to have someone to blame/shout at 24x7 if something went wrong. (And that company was ~85% likely to call me and blame it on the app before they even bothered looking to see if the platform was working...)
How much additional profit would it generate if I could cut your cloud costs by 20%?
Usually, the savings are even higher, but 20% is low enough to be easily believable, yet large enough for them to invite me to learn more.
The reason I do it manually is that many companies are afraid of the workflow interruptions usually associated with moving off the cloud. That's why I first analyze their deployment to calculate possible cost savings and then I offer them to hire me to take care of the migration. I'm pretty sure those companies would not feel comfortable switching to a different cloud by themselves.
The cloud providers, however, have a trick: they can eat these costs and make serious $$$ on the actual compute, storage and bandwidth.
If you’re trying to compete with AWS et al based on price for bandwidth etc “but with common services such as S3”, you’re gonna have a really hard time convincing a customer your expensive managed Kafka is reasonably priced. And then you’ll just find yourself competing on just VPS servers...
https://aws.amazon.com/s3/faqs/#Durability_.26_Data_Protecti...
It's "Simple Storage Service", and it does that one thing (storing practically unlimited amounts of objects) spectacularly well. Like 11 nines well. It doesn't do "inexpensive egress charges" well. Amazon never claimed that though.
I think a lot of people don't care about those 11 nines, and would love to buy something with perhaps only 5 or 4 nines of durability ("Maybe I'll need to upload it again once every few years, that's cool!") but which optimises for cheap bandwidth costs instead. Right now, I think that thing is an inexpensive VPS with an unlimited traffic (Hetzner were offering "unlimited traffic at 1GB/sec for the first 20TB, then throttled to 100MB/sec" for under 10EUR/month last time I went looking. And they're well above "the bottom of the barrel" pricing in cheap VPS land.)
Seems to me S3 is more designed for highly reliable data storage. If you wanna store a bunch of data in a way that's highly reliable and resilient to hardware failures and also not have to worry about managing RAIDs and clusters of servers for the data size, S3 is just the thing, and probably priced pretty fairly.
It's also a capable and flexible service. It's possible to use it for things it isn't really designed for and have it behave pretty well. Well enough that you can mostly ignore that the other thing isn't really what it's designed around. The "mostly" is important, though. If you start hitting any extremes while using it in a different way, then you certainly can run into some pathological cases in the pricing structure and pay way too much for something that would have been cheaper in a service dedicated to that.
Looking purely at the traffic price, I believe the word choice "overcharging" is warranted if you can buy the exact same product elsewhere for 90% less.
Their cost is less than .01 cents per gb and look at their published pricing for customers.
And I’m including networking equipment as well as head count in my cost, not just the cost of pipes.
You don't need to manage any storage to get your bill to drop, S3 can be fronted by your cached proxy (minio or just plain nginx works easily) it is trivial to setup on OVH or hetzner or any other VPS provider.
You get all the benefits of S3 without the egress bill of a major cloud
He used a CDN, just didn't read the fineprint:
“Things I should’ve known but didn’t.” Did you know that “The maximum file size Cloudflare’s CDN caches is 512MB for Free, Pro, and Business customers and 5GB for Enterprise customers.” That’s right, Cloudflare saw requests for a 13.7 GB file and sent them straight to origin every time BY DESIGN. Ouch!
It's another case of misunderstanding and misusing a tool/service in my opinion. Cloudflare (and Cloudfront) are intended to be website CDNs. Their use case is _not_ serving 13.7GB objects - and so they didn't. (Whether "passing those requests straight back to the origin every time" is a better/more customer friendly thing than " blocking those requests since they're not being caches due to an object size rule" is a good question...)
If you want "inexpensive large file hosting/serving", you don't use a combinations of "a service with 11 nines of durability" and "a web-infrastructure and website-security company, providing content-delivery-network". As you rightly point out, duct taping this together on VPS providers competing on low-cost bandwidth is the "right thing" for this use case. The gold-plating of 11 nines of durability and a CDN with over 200 globally distributed POPs almost (perhaps "should have"?) cost this guy $2.5k for "taking the easy way out without thinking through the consequences".
In my head, it's like he had a Tesla Model S in the driveway, and needed to transport 6 tons of lead bricks across town, and thought "I know, I'll just load them up on the back seat and drive 50 miles!" and then getting surprised the car needed expensive repairs afterwards.
The b/w costs are very high on the cloud , 5-6x higher than it is outside, I don’t think there any value they provide that justifies it for vast majority of users .
I am pretty skeptical on anyone at all in the world needing 11 nines of durability. You are 100’s of times more likely to have a meteor strike .
All the major service providers have gone down recently. Cloudflare went down just this month . At best they are delivering 4-5 nines practically speaking (loss of one PoP or region is still down )
CDN is exactly the service which needs to be significantly cheaper at scale, if you really need 200 PoPs then you have at minimum 10,000s of users at which and you will have consumption
If the industry’s position is they only want/can service enterprise customers for whom current pricing is within the budget, then their online pricing models/ marketing are incredibly misleading on who they are targeting .
To me it looks like their business models depends on predatory pricing prosumers y hooking them on with the ease of use.
This is very much like credit card industry. If everyone paid on time , CC providers will never make money, predatory loan pricing at 36% or higher is their real revenue source.
Paying on time / reading the ToS only lets them off the hook legally , morally they are both praying on vulnerable users .
But I guess simply running Debian on some server and keeping that up to date just isn't cool anymore. Everyone running a blog about their dog needs Cloudflare, S3, heroku, micro services and docker. Obviously with no limit on what kind of bill they will generate should something go wrong. You can just vent on twitter and generate enough attention that the vendor will make an exception for you and pay you back to limit the negative publicity. Who wouldn't prefer that to the tedious work that is maintaining a Linux install on a server?
Now that said it looks like his job is in cloud evangelism too so I’m sure a large part is that he wants to maintain a personal account to learn/hone skills. I’d recommend anyone do that, getting a look at aws without whatever your company is doing on top of it is pretty awesome for learning. Just don’t host massive S3 images and shoot yourself in the foot ;-)
If you just host static content you don't even need to care about security updates. Just deploy a new instance, "apt install nginx", done, forget about it. I would say less work than setting it all up on AWS.
When is the last time there was an remote exploit for nginx serving static files or in Linux kernel? You dont even need to worry about security updates in this scenario.
> I’m sure a large part is that he wants to maintain a personal account to learn/hone skills.
Yeah, he seems to be responsible for OpenShift (marketing?) at Red Hat and he only learnt how expensive AWS is when he got a bill? Bit embarassing.
But yes you need to run, for example, a certbot as well to setup certificates.
so it's: apt install nginx apt install certbot certbot certonly -d your.domain -d www.your.domain # again, when you apt install certbot it installs the cron/timer automatically for automatic certificates renewly and nginx reload
Still less work than going the CDN route I would say.
It really seems to me you don't trust default linux settings and need to take care of everything and fine tune everything but you trust CDN providers thus making the CDN route less work for you. If you trust default settings of a linux distro, it's less time consuming setting it up yourself on linux. Yes that's how user-friendly it really is today.
The problem is that, at least in the case of public clouds, there’s a real risk of your bill exploding. I guess the author learned a lesson here, but I don’t think it’s the right attitude to start blaming the guy for “it makes no sense to tinker with AWS services” over here. Learning is probably one of his goals.
I wonder if a potential solution is to have two billing modes that get hardlocked at signup (or require a key or something to change): one is the standard model with alerts etc. The other is a personal model that kills all of your stuff when you go over some limit. I would feel much safer if the latter were in place.
https://docs.aws.amazon.com/awsaccountbilling/latest/aboutv2...
I do this with my startups, so why not do this with your sandbox account?
AWS does not support a 'pre-pay' model, and to my knowledge there's no water-tight way of capping your costs. Yes, you can build an watchdog to nuke all your instances if you go over-budget, but there's still the risk of missing some unexpected source of costs, or misconfiguring your watchdog, or perhaps not getting there in time, etc.
AWS could support pre-pay, but they don't. I think it's a reasonable criticism. There are plenty of horror stories about surprise AWS bills. [0][1]
Never managed to get the money back. They always seem to focus on building tools around reporting (ie “budgets” being just a report, rather than an actual enforceable budget).
But I still can’t escape the fear of accidentally triggering something that costs a lot. It still happens sometimes, a sudden $1k EFS bill being the latest.
(I wouldn't have any problem with this, just seeking clarification.)
The whole article just read as a "I don't even now what I'm doing"-thing.
As for Cloudflare, hosting multi-gigabyte files is not what their service is for, I can’t see how you can blame them for having a limit on how large files they cache.
You survived the DDoS, that's great, now go see your bill.
The article doesn't use the term "piracy", but I'm curious what Microsoft's license says about public redistribution.
Some sources define piracy as redistribution outside the terms of the license. That implies that you can pirate free software, if you violate the license.
Without getting bogged down on definitions, I believe MS could issue takedown requests to anyone hosting their free trials, when the license forbids it.
That's actually what I love about these dedicated providers. Often I prefer to be surprised by cutting the service instead of a gigantic invoice. For some business-critical applications I can understand the need to have it scale (and price) accordingly, for many other use cases it is better just to have the serve switched out when the traffic limit is hit off.
I really hate long-form articles that feel the need to explain the weather, someone's clothing, or their family member's eating habits.
COME TO THE DAMN POINT. This is the Internet. 99% of content is garbage, and I'm not interested reading through pointless content-free filler that could be generated by a neural network just to find out if your article somewhere contains useful/interesting information.
My hypothesis is that the client connected to Cloudflare and performed a HTTP range request for a portion of the 13.7 GB file. For an unknown reason, Cloudflare did not preserve this range request to S3 as the origin. It transferred the entire file, returned the range requested by the client, and dropped all bytes transferred into /dev/null because it is not caching.
The end result is that Cloudflare pulled down 30 TB of data while delivering 67 GB to clients in the one month period shown in the screenshots from the blog post.
See also https://blog.vbgn.be/2019/06/20/nextcloud-cloudflare.html
What's not stated here, and the OP learned, is that requests for larger files are always passed through directly.
https://support.cloudflare.com/hc/en-us/articles/200172516-U....
So when you request the large file, the CDN (cache) doesn’t have it, so it immediately passes through to the original source. The details of how this is implemented don’t really matter in the sense that they are not going to host the large file(s) in the CDN, so they will always fall through.
It sounds like he was operating in tiny land and expected bills to be commensurate, so it just seems irrational to expect a low or free tier of another service to do this for nothing.
If you literally put a massive file on the public internet, what the actual fuck are you complaining about? The quality here is such shit.
Let's snipe at amzn, everyone hates them because they're winning, because of hn readers dependence on AWS to run their shit web apps... It's like: ope! It's been a day since we had a "amzn scammed me" post. Nevermind that the people complaining are either a) too ignorant and footgunned themselves, b) minimizing or lying about their own complicity, or c) outright scammers themselves.