AWS Customers Rack Up Hefty Bills for Moving Data
theinformation.com
theinformation.com
Turns out that by default Google's Docker container registry only stores the Docker images in the US. So each time I launched a VM the Docker image was downloaded from the US. I wrote more about it here: https://www.mattzeunert.com/2019/10/13/reducing-docker-image...
The billing interface didn't show that the Cloud Storage cost was related to the Docker images. I was investigating my normal Cloud Storage use, but it didn't explain why I was being charged so much. Only after a few days did I get the idea that it might be the Docker images that were causing it.
That's an expensive disk drive.
You could do this even just pegging the line 9 minutes per day and otherwise leaving it unused. ;)
People forget just how affordable it can be to maintain your own infrastructure. You can have the hardware and network capable of supporting 10X your average traffic loads and still have it operate far more cost effectively than the equivalent traffic on AWS.
At my work, we slashed our overall hosting costs by moving a data warehouse off of AWS and on to our own self-maintained infrastructure.
But with a bloated inefficient IT department or non-savvy negotiations with hardware vendors or transit providers, it can also be more expensive than AWS.
What you get from AWS is the logistics pipeline is already built as is the infrastructure should you suddenly require to serve factors of traffic more.
The ROI of AWS comes from the backend and capex vs opex debates.
Example : The CFO can go to the board and explain we are getting ready to reduce opex by laying off 10 developers at 150k yr during any meeting. The capex cost is usually fixed and hard to explain away.
The other is stupendous burst activity, like you just need a thousand(s) cores for a couple hours. Of course this doesen't mean the baseload has to be in AWS, just easy for small teams.
It's generally a function of network cost versus infra cost, and so the 'equation' solves differently based on the longest tech cycles, from inception to maturity to skill pool to diminishing returns and then back again on some other mode — this really is a decade+ thing.
Some "always true" inherent advantages of one approach (e.g. availability for cloud, or resources for on-prems) would remain across cycles as permanent gains; disruption then occurs when some new approach (e.g. containerization) fundamentally upsets the order of costs.
It used to be OpEx was easy, CapEx was hard to get approved. As people took advantage and OpEx went through the roof people are getting alot more pushback on reducing OpEx.
Pushing back on OpEx in favor of CapEx might be seen as a more long-term strategy too. Basically rent vs purchase, even for deprecating assets like infosys, as I think the current DevOps trend (scaling in pure software, virtualization, etc) makes it easier than ever to squeeze every last FLOP on-premises.
Okay, like AWS
> and doing that is a service that it’s probably worth paying for.
Okay, but how many orders of magnitude?
They want you to keep all of your data in their cloud, do all your processing there using their services (because doing it elsewhere incurs expensive egress fees), and get paid handsomely for your need to actually serve the data to your customers.
This also makes a migration additionally expensive, because you would need to egress all your data (old logs, some fresh backups, etc) in a short while.
So it's basically a soft lock-in.
I think it's more than that because a lot of folks aren't aware of the costs until they need to or want to move and then they get hit with a massive bill.
So maybe they chalk it up to experience and pay the bill because they have no other choice or calculated it's still better in the long term.
To me that's a much difference scenario than walking into a store and happily buying a dozen eggs for anywhere between $1 and $1.50. In this case you know what you're getting into before you make the purchase and everyone around you (other customers and businesses) decided that's what eggs will sell for in the open market. With outgoing data fees, it's more like a "take it or leave it" price dictated by the provider while they already have your data and there's no price competition since they are the sole business with your data.
Whether or not it's the customer's fault for not doing enough research is debatable, but it certainly doesn't help that most providers make it pretty difficult to calculate costs.
The market isn't exclusively good, and has many failure modes. This is one of them.
Sure, it’s the market. Why? Is it a Giffen good? Is it a case of very poor visibility of services to consumers? ...? There is something interesting going on, let’s figure it out.
This is, oddly enough, similar to a debate people have about consumers TV or Internet: should pricing be "unlimited" or "a la carte"?
AWS is combining all your networking charges into one lump "outgoing data transfer" fee. So it's heavily marked up in comparison to what they're paying for the outgoing data transfer, and you're not sure how much is profit vs. whether it's going to cover all their other costs.
So it might be fairer if AWS broke out separate line items for internal, incoming and outgoing data transfer, plus all the additional systems a customer uses.
I think AWS's billing is probably already on the falling side of diminishing marginal returns. That is, it's complex enough that more information would tend to hinder customers from getting the best price. Right now, if I plan to reduce my data charges, I have one variable to tinker with. If we expand this, it would mean I'm having to balance incoming / internal and outgoing charges. That sounds simple, but in terms of engineering it can be very complex.
The next claim is that this biases customers not to move. Of course, Azure and GCP have the same arrangement, so while you pay to move out of AWS, you don't pay to move in to Azure or GCP. So all the vendors are attempting to lock you in to their product, and at the same time trying to extricate you from their competitors, overall it's a wash.
So, yes, part of the motivation for egress charges is that ingress is a loss leader. But it's also true that egress is a metric that does, for the vast majority of their customers, directly translate into customer value. If there's a compelling case for doing it differently, someone should do it and see if it works.
Cloudflare doesn't charge for bandwidth. I always throw cloudflare on top of anything I do, not because I really need a CDN or anything, but because the bandwidth cost would bankrupt me otherwise. The ceo of cloudflare gave the rationale on why they don't charge:
> There’s a fixed cost of setting up those peering arrangements, but, once in place, there’s no incremental cost. That’s why we have similar agreements to Backblaze in place with Google, Microsoft, IBM, Digital Ocean, etc. It’s pretty shameful, actually, that AWS has so far refused. When using Cloudflare, they don’t pay for the bandwidth, and we don’t pay for the Bandwidth, so why are customers paying for the bandwidth. Amazon pretends to be customer-focused. This is a clear example where they’re not.
They also do charge for Enterprise plans, but instead of transparent pricing I got high-pressure sales techniques and black box pricing offers - which then anchored our rate so that as we grow past our current contract, we're forced to upgrade at any point with pricing based solely on our original negotiation.
Frankly, while I save money using Cloudflare over Azure's CDN right now, it's left a very sour taste in my mouth and I'll be jumping their ship as soon as I have time to find a suitable alternative.
If you have the ability to shift your entire enterprise CDN away from them, why not first try renegotiating?
The example given above for comparison, Hetzner, also doesn't charge for inbound and internal transfer AFAIK. Nor "the additional systems a customer uses". You pay a charge for the server, you get some amount of traffic included, and if you go over, the additional traffic costs something like $1.1/TB. That's all you pay.
> So it might be fairer if AWS broke out separate line items for internal, incoming and outgoing data transfer
This explanation doesn't cut it for me - most (all?) "traditional" VPS providers don't charge for ingress traffic, and I doubt anyone, ever, has charged for internal traffic.
So what exactly is 'all the networking charges' comprised of, other than egress data?
For egress bandwidth costs, I'd assume it included, well, the egress bandwidth cost.
But I think level3 is associated with:
https://www.centurylink.com/business/hybrid-it-cloud/public-...
And while they have a call-us price list (if you have to ask...) - they at least state:
"Public and Private high-capacity networking options up to 10Gbps. Note: there is no charge for internal data center traffic. Cost on a per-GB-out model"
I have no idea what they charge pr gb for this cloud product however.
That companies don't charge for specific things doesn't mean those things don't cost them anything. It just means they're trying to work out a pricing scheme that scales with customer usage and is broadly understandable. So "data egress" is really just a proxy for "how much stuff you're doing with the networking subsystems of AWS."
Same thing with EC2, there are a whole pile of costs that are summed up with "time you rented an instance."
Of course there are is an internal cost of doing business, and peripheral infrastructure cost - but if I pay $100 for service "A" I reasonably expect that fee pays for service "A". Instead, egress bandwidth costs seem to be used to trick customers into thinking services are cheaper than they really are.
No, it's two variables - the egress charges you refer to and the actual cost to store the data.
We[1] have found that it is, as you might expect, quite a bit simpler to charge for just the storage and forget about metering the usage/bandwidth/transfer.
So we have typically had our price point higher than the B2s or Wasabis of the world, but there's just one simple number to think about - and no potential for surprises in the billing.
I will admit to having a bit of concern over adding 'rclone'[2] to our platform and the potential for users to just burn bandwidth using an rsync.net account as a "transfer host" but that is why we peer with he.net and their cheap an plentiful 10gb pipes.
[1] rsync.net
[2] ssh user@rsync.net rclone s3:/bucket gdrive:/blah/blah
Which is to say, each of our five[1] regional POPs have a single connection provided through a dumb switch one hop from he.net[2].
They have no interconnection or dependencies to one another.
No routers, no firewalls, no balancing, no failover. When rsync.net fails, it's a very, very boring failure.
We've had zero network outages in the last 60 months or so.
[1] Fremont, San Diego, Denver, Zurich, Hong Kong
[2] init7 in Zurich ...
And I think the charges for inter-AZ transfer are to incentivize customers to do that.
Of course, to make them fully independent, you have to replicate everything, so you wind up buying several redundant copies of your system...
Yeah, and keeping around warm systems ready to failover in case of a zonal outage seems like a preposterous waste of resources.
The alternative... to keep around multiple replicas of your system in different zones, all ready to accept traffic and which do serve traffic, seems more practical and less wasteful.
If instead of availability being the only value, there would be a more value provided from actually using such resources, more folks would adopt cross-AZ architectures which would be a win-win for both the customer (get HA for lower or no cost and go down less often and succeed in the market) and thus the cloud provider (keep raking in the steady cloud revenue as the customer grows).
This is one of those gotcha's that company's hit. They see the public pricing page and think "wow that is much cheaper than one my internal IT department charges for X", and then when they go to actually implement they find that "best practice" says they basically have to more than double or even triple the cost to get a reliable system (more because not only do you have to duplicate all the infrastructure into a second AZ, you are getting charged for the replication traffic between them).
This explains the cost.
Price probably should be based on value, not on cost. Why do they charge for it? Because they decided it's a good way to make money and profit.
People under pressure make short term optimizations at the expense of long term strategy. There's nothing more to it. AWS reduced compute and IT expenses in the next few fiscal years, so people jumped all in on it. Solve the problem now. Someone will fix it in the future.
Actually, it would probably just make some programmers even more intolerable.
Europe is probably the most competitive and best market in this case, almost all other markets are dominated by monopolies. Yes, even significantly worse than the Telekom monopoly.
Look at transit in Singapore, Japan or Australia. Or even in Brasil. You’ll go bankrupt even from trying to deliver a single movie to customers.
Hurricane Electric has POPs in both Singapore and Australia. Transport between Singapore and Australia is $1.50 or less. Peering ports are $0.20 per Mbps.
If all you want to do is push movies at customers, there are plenty of dedicated server providers who will sell you bandwidth on the cheap in both countries.
Obviously YMMV if you want better routes or direct interconnects with local monopolies.
Isn't DigitalOcean a cloud provider ($0.01/GB)?. Isn't Oracle a cloud provider (first 10 TB free, $ 0.0085 after)? Isn't OVH a cloud provider (free bandwidth)?
One thing folks like about AWS - you can actually know what you will be paying and there is no fake / hidden limits.
That's my somewhat outdated experience wasting a TON of time on this idea ages ago.
Give them a warning if you're going to spike your bill that high, but there shouldn't be any fake/hidden limits.
When you care about packet level SLOs... yeah, you start to shop around.
That's pretty much it - in my experience most of AWS' awesome features are designed to lock customers into an environment where Amazon increasingly provides all of the components and services you need to do business.
I would assume the endgame is to create an ecosystem where the vast majority of customers allow their IT function (infrastructure, developing software, engaging third-party SaaS vendors, etc) to atrophy entirely, after which they'll have no choice but to buy what Amazon is selling, at any price, in perpetuity.
I've always got full 1Gbps out of Hetzner, even across the ocean, and for periods of multiple days of transfers.
(Like on all machines, you need to set the right TCP settings for windows sizes to make it physically possible.)
If you are doing it as a replacement for traffic within an AWS region and availability zone, it seems like you will be both more expensive and have much higher latency.
Or is the application something else entirely?
I assume you didn't mean that literally because I can't see how that will ever work out in terms of cpu cost. I think breaking it up into blocks like what RAID4/5/6 would be better but will still impact the performance of reads.
The performance of writes is going to be worse. Not because of the parity calculation but because you will be taking the max latency over all the cloud providers.
I can't see people trading off that much performance for better fault tolerance (in a world where S3 guarantees 11 nines) or ease of switching.
I'd probably use commodity VMs for this rather than big clouds if it is indeed resilient.
I disagree on this one. The margins on egress are, well...egregious.
Even if a cloud provider had competitive transfer costs they likely wouldn't attract any new customers and would have less margin left over to subsidize the main cost customers look at, $ per instance hour.
The less attention is paid to transfer costs the better for AWS/GCP/Azure. Why hasn't a spot-market for transfer been introduced? Same reason why I can't sell my unused home internet bandwidth to my neighbors, the money is in controlling the means of transportation/communication and the providers want to keep as tight a control on that as possible.
(Source Open Guide to AWS - https://github.com/open-guides/og-aws)
Then again, I worked for AWS for years, so maybe I'm just used to thinking this way so I'm not really surprised by it.
You still see paying for bandwidth with residential connections, though some operators (like Comcast) are trying to do away with it.
This is just the static picture though. What's harder to predict are the consequences of some innocuous looking code change.
What I'm saying is that for a hosting architecture to make it difficult to predict the cost of any code change is a downside compared to an architecture that makes such predictions easy and intuitive.
Of course you will try to mitigate any downsides and learn what you can from any mistakes. But unpredictability makes learning far more difficult than it should, which inevitably means a waste of development resources.
The issue with per-bit pricing is that a fair agreement for network use would probably look like paying a fee that makes up for the amortization of the network equipment. Anything else is an artificially restricted market created in an attempt to extract more value out of consumers by having them bid against each other.
At some point, yes, we will run out of places to put the switches and routers and then the cost of connectivity will be closer to the cost of land use and will mimic rent, but we are a ways away from that.
Enron tried to create a market for this.
[0] https://www.wired.com/2001/11/enron-a-bandwidth-bloodbath/
A flat fee for the act of a human getting the data onto the physical medium ($200) + the cost of shipping (<$100?) + $15 per day you keep the snowball device past the first + price per GB of data you're transferring ($0.03 per/GB).
So if your getting out 30 TB of data that's $200 + ~$100 + ($0.03 * 30000) = ~$1200
I've been a developer for 10 years and AWS seems to me like its intentionally designed to be as messy as possible.
I am completely unaffiliated with this site, I just enjoy it
Hold onto the packet for as long as you can vs hand it off to your peer as quickly as possible.
Most networks do "hot", Google does "cold" since their network is almost always better than that of the peer.
For example, a cold potato network may have a link from Dallas to Chicago to New York, while a hot potato network could have a direct link from Dallas to New York.
Cogent uses cold potato and is frequently worse than other transit providers.
(Also, they offer the option to use hot potato and pay them less: https://cloud.google.com/network-tiers/docs/overview )
I'm a Google network SRE, but perhaps I'll get new business cards saying "cold potato engineer".
But their hot potato still costs $65+ per TB at medium volumes and $45+ per TB at high volumes. That is still extremely high compared to normal peering costs.
This seems more sustainable than Wasabi's model, but there's no way of knowing for sure.
AWS (US East 1, no free tier) - $92.07
Azure (East US) - $88.65
GCP (Americals) - $122.88
I'm quite surprised by GCP being the highest cost here, and by such a wide margin.
Some of their advice for saving money was to keep track of who created each resource and why, so there's less reason to doubt whether an apparently unused resource can be deleted, and to make some limits regarding how many resources can be automatically created (especially in dev environments). Some other ideas were to look for signs of inactivity like low CPU or bandwidth use, and consider deleting such little used resources.
There was much more to the talk, but those were some of the highlights that I can remember without digging out my notes. It was a good talk.
I'd believe this more if the pricing of bandwidth on AWS hadn't stayed pretty flat since its launch.
Plus, it's frustrating that AWS Lightsail (https://aws.amazon.com/lightsail/pricing/) offers a $3.50/month plan with 1 TB of transfer. That terabyte alone will cost you $92.07 on a normal instance, and the $3.50 includes storage and an instance!
Any chance some of those high costs are due to movement of data due to GDPR compliance? Maybe Apple did all its prep in 2017.
Also there are some misconceptions on inter-AZ data transfer as well: (https://www.lastweekinaws.com/blog/aws-cross-az-data-transfe...)
And they could negotiate the best rates on the planet given their scale.
ie a $20/mo instance on linode gets you 4TB of transfer -- $0.005 per gig. Scale enough of these in various DCs around the world and you have a pretty cheap self-hosted CDN.
https://servercheap.net/crm/cart.php?gid=10
Or if you don't trust "unlimited", Hetzner sells CX11 with 20TB/mo for 2.96 eur/mo
$20/mo for 100TB metered ($0.0002/gigabyte). $120/mo for 200TB metered ($0.0006/gigabyte). $170/mo for 1gbps unmetered ($0.00085/gigabyte at 200TB, $0.00052/gigabyte at full saturation one way).
or https://client.layeronline.com/cart.php?a=confproduct&i=0
$199/mo 1 gbps unmetered.
https://oneprovider.com/order/item/dediconf/2372
$40/mo (Seattle)
https://www.dedimax.com/bare-metal-server/16954
$40/mo (Seattle)
As far as I know, Apple would never share cost data like this, nor would a lot of others on this list. Apple doesn’t even publicly acknowledge that they are customers.
I'm not sure that "stolen" is the right word but, yes, it certainly appears based on what The Information wrote that either someone at AWS or a third-party with access to the info (probably not too likely) leaked the confidential numbers. Disgruntled former employee or... Who knows.
It goes down, but only to $14,131.11!
Nice find. I will file this away if I ever need to do a full remote recovery.
> 66.3. You may not use Amazon Lightsail in a manner intended to avoid incurring data fees from other Services (e.g., proxying network traffic from Services to the public Internet or other destinations or excessive data processing through load balancing Services as described in the Documentation), and if you do, we may throttle or suspend your data services or suspend your account.
It still costs $0.03/GB, which is on top of the $200 they charge you to use the service.
Either they put it together from public resources, or (more worrysome) someone at AWS leaked them line-item based expenses of their top customers.
This has got to be incredibly sensitive data for AWS, not something they'd want leaked. If I'm (say) AirBnB or Snap, I'd worry that this data leaks information that would be detrimental to their ability to negotiate for cloud computing with Google, Microsoft etc.
I am often interested in some dataset for which I can afford the storage in the form of a hard drive but not in the form of a download through my home connection.
If a service existed that simply offered the following:
* customer provides URL (and optionally hash checksum)
* customer pays, and later receives hard disk drive / SSD drive with the download contained
* democratic pricing for the media, or alternatively send your own media (hence at twice media shipping cost...)
* possibly eventually local brick and mortar locations / affiliate locations to drop off and pick up media, in the larger cities
* preferably without account, although an account is not a large impediment
* definitely not coupled to a financial account in a credit fashion, i.e. no qualms with topping up the account, but the service should not be able to withdraw money from my financial account. i.e. debit only (like the typical european bank cards, yes I know credit cards are available in europe as well...)
This would seem like a profitable side business for many programmer types who have high bandwidth connections (or have access to them and are allowed to use the connection for this purpouse).
If someone builds the software stack for a main portal such that affiliates can advertise their physical location, and their pricing for media, for download and for copying to media, then customers could compare and choose on the basis of price.
1TB HDD: 59 AUD International Shipping (assuming we can keep it under 1kg with packaging): 38 AUD
That is 97 AUD before considering any profit you might want to make (35 AUD for the movie + small popcorn and drink), fixed setup costs like a NAS to cache data sets or the bandwidth costs.
there seems to be plenty of opportunities here, like buy or rent an old bank building with the individually lockable drawers, put your drive in the locker which has a USB cable or ethernet cable, close the locker, use some app to set the URL / hash, pay, and you get a ETA, download complete notification, and a deadline to pick it up (or else incur a fee to unlock the drawer proportional to overtime).
I guess the idea could be pitched to those operating rentable local PO boxes, using similar lockers, but with internet connectivity.
Digital Ocean: 1 cent per GB
The really weird thing is this should be the absolute lead on all Digital Ocean marketing but they don't even mention it. It's their single biggest selling point.
(Although Azure is apart of the bandwidth alliance if you use cloudflare/b2)
[1] https://support.cloudflare.com/hc/en-us/articles/36001614391...
[2] https://cloud.google.com/interconnect/docs/how-to/cdn-interc...
I found it more interesting just to see a list of their top ten customers. In particular I didn't realize that Capital One had so much infrastructure.
Moving data into any network that is outbound heavy is free because both paid peering and transit is settled based on a peak percentile traffic (unless it is flat rate).
That's why the "gansta" position is to have a balanced in/out for any network as in that case you get to effectively double charge for the same pipes -- your eyeball heavy customers pay for incoming and your web farm customers pay for outgoing.
https://www.lastweekinaws.com/blog/aws-cross-az-data-transfe...
No connection to the blog.
Make it nearly free to move documents in.
Charge monthly rent per box of documents.
But charge like a wounded bull if customers try to permanently remove documents.
Do NDAs expire? AFAIK you sign an NDA and you’re bound to respect that forever. Unless the knowledge becomes publicly available.
As a quick example, this random free template I found online[0] uses "5 years".
What happens to them if the company goes out of business?
I have spoken to a lot of lawyers about contract terms and what I've been told for both NDA and Non-compete style agreements is that they have to have very clearly defined scope to be enforceable.
For a non-compete for example, it requires distance and specific field with a clearly defined time period (that's reasonable).
For an NDA you have to ensure that the covered information is explicitly labeled (which is why people have confidentiality labels in email footers) and that there is a defined time to expiration. An NDA without those two criteria place an undue burden on a person who is not being compensated for their compliance.
Reasonable time period is generally 2 years.
All they provide you is the bare metal in easily-purchasable quantities. Whatever time you saved not having to lug a server up to your datacenter and plug it in will be spent debugging CloudFormation stacks or figuring out how to auction off your reserved instances that you no longer want.
Are you saying that AWS service aren't providing value to their customers but it's fear-based marketing?