With the huge margins they must have on egress bandwidth, I'm not holding my breath though.
I don't think they have huge margins on egress at all. There needs to be some incentive for customers of cloud services to minimize bandwidth usage. It is a limited resource.
The $/Mbps prices there - about $6k/Tbps in the US - are based in reality and are absolutely reflective of what it costs, hardware, software, redundancy and all - for an effectively 1Tbps pipe.
If you're pricing as $/GB on top of that capacity and keep it reasonably heavily utilized—which can be hard given diurnal demand—the margins only get better! Products like Glacier (S3) exist to fill exactly those gaps.
(Note: currently work at Cloudflare, but wasn't part of this blog and I've been around a bit...)
5Gbps*1 month is ~1.5 PB. AWS is about 0.02/gb or ~30k/1.5pb.
Approximately 30x the cost.
These are old numbers for specific use cases so I'm not sure how much that has changed.
https://aws.amazon.com/shield/pricing/
but regardless, the EGRESS charges are what are absurd. The only logical reason for charging so much for data exiting is to make sure that you can't practically leave the AWS ecosystem.
Just one example is Hetzner, who include 20TB of bandwidth, with anything over that charged at just €1/TB.
Meanwhile, AWS is gouging at €80/TB.
Come on, Amazon is not serving traffic through a Comcast business connection, they're peering directly with other large operators for free or for next to nothing.
- A 42U cabinet - A 15A 120V circuit - An unmetered 1Gbps IP transit link
All at a proper datacenter, namely Hurricane Electric's fmt2 facility. Includes a /29 of IPv4 and a /48 of IPv6, allows me to announce the /24 and /36 I own, and there's a free internet exchange onsite which I have a 1Gbps port at.
You need to provision enough power for simultaneous startup after a power outage, unless you have some really smart PDUs and automation. We have a 10 year old DC with 30A@208V per rack and we have to leave racks half full because modern servers are so power-dense.
In any case, even at your optimistic 250W per rack unit, a full rack of servers with two switches is more than 10kW!
You could possibly run this fictional daemon on your management switch :)
I think this is the most common “enterprise” datacenter server type in 2021, mostly due to licensing constraints from VMware/Microsoft/Red Hat/Oracle/etc. Such servers give the most “bang” per dollar when licensing costs are included.
https://i.imgur.com/Zr5rq4X.png a graph if you're curious :)
They very much have massive margin on egress. Given some of the cost comparisons floated today comparing R2 to S3 egress AWS is likely hitting 1000s of times (likely more) their return on the actual bandwidth cost they pay for month over month.
Usually DoS attack doesn't exhaust the bandwidth, come on..
Remember those ISPs need to make margin, that's why they charge what they do.
AWS _could_ make it's margin on services alone. They just choose not to.
If your data is in S3 or DynamoDB, egress fees will encourage you to process the data within AWS. You’re not going to use BigQuery for that. If you want to add a search index on that, you won’t go shopping around for search products on other clouds, you’ll likely use AWS elasticsearch. That’s where AWS makes money - through this soft lock-in.
They’ll survive and maybe even thrive if they lose the revenue from egress. But who knows what’ll happen if the moat is destroyed? Every AWS service would need to compete on merit with every service on other clouds. They won’t be chosen simply by default.
Euh, 4-80x , the 80x is for EU and the US.
Seems like a pathological use case to run a query engine on one cloud provider datacenter (Google) against the disk storage at another cloud provider (Amazon).
Even if egress were $0, I still wouldn't want to do that. I want queries to run as fast as they can and the WAN link bandwidth is opposed to that.
Is there anything about BigQuery that would compel anyone to do that instead of just using AWS RedShift?
So generally speaking, latency and bandwidth between services is not a significant concern. It's all about egress billing.
Things like operational overhead don't always get a look in when a team has convinced someone with the purchasing authority that tool X is going to solve all their problems. Even if the entire org has zero experience with it and it's going to have flow on effects.
A recent example at one of my customers was a team deciding to outsource a platform to the provider. (Outsource, not SaaS it's a managed service hosted in AWS). I told them the network design and AWS build on our side to join the two would require significant effort and they said that's fine. Now we've spent almost their entire budget for the move on just working out how to connect their VPC to ours (there are some legislated controls we had to put in place and the vendor architects were less than helpfull). Of course it's all my team's fault because we are the ones who say "you can't just plug the two together" and it would be much better if we had a "can-do attitude like the other team instead of naysaying all the time."
You'll find data centers of all major providers within a few miles of each other in at least 10 locations around the world. Latencies are <5ms and the links between those data centers are cheap and can provide huge bandwidths (even though they don't want you to know this and still charge huge prices).
So from both a technical and an economical perspective there should be no reason why you can't shift terrabytes of data daily between GCS, AWS, Azure, Hetzner, OVH, Cloudflare etc. Only artificially high egress pricing set by the three biggest providers keeps people from doing that.
In Europe with OVH and Hetzner that's exactly what a lot of people already do (from my experience), also nicely visualized in their Weathermap (look at Frankfurt) [1]. These two providers are tiny compared to AWS and yet the 400gbit/s links between them are utilized ~50% pretty much 24/7. And from my experience I can say that there are quite a few use cases where mixing these providers is cheap and easy.
I had a huge corpus of data and ran some instances that were basically downloading+parsing+summarizing data for a couple days. Again, I didn't think of bandwidth as they were EC2 instances querying S3 objects, what could go wrong, it's Amazon to Amazon, right? Wrong.
Bill came up in the 10,000s USD range ... fortunately, it happened during a trial period where I could spend a lot of money for ~three months.
I moved my stuff into a small-ish cluster of unmetered VPS and the whole thing is 100x cheaper.