Comparing Bandwidth Costs of Amazon, Google and Microsoft Cloud Computing
arador.com
arador.com
The big cloud platforms offer a rich selection of different offerings, which (just like in every other industry) cross subsidize each other.
When I go to a restaurant, I don't expect that they will be making the same profit margin on every item on the final bill, and in fact, they almost never do. Drinks tend to have a very high profit margin, some labour intensive items may be a break even at best, and the complimentary bread sticks or chips and salsa (if offered) will certainly be a loss.
I guess I could write a very upset article about how my local mexican restaurant is SERIOUSLY SCREWING ME OVER with their drink prices, but if I don't write the companion piece about their cheap burritos (subsidized, of course, by the drink prices), it would only show half the picture.
The reality is that I'm buying a whole package (at AWS or a restaurant) and I should evaluate the whole picture. Yes, I can get bandwidth cheaper outside AWS (or a can of coke a lot cheaper from a big box retailer). But I can't really get the total package of integrated, managed services outside AWS (certainly not for the cost they charge), any more than I can get someone else to show up in my kitchen and cook a three piece meal and then do all the dishes. (Which is to say, I totally could hire a chef to do that, but it would cost me a lot more. I could BUILD an internal SQS clone if I had to, but my employer would never break even on the cost of getting me to do so.)
AWS is very cheap for some things and very expensive for others. Depending on your usage and workload it may or may not be economical to buy the package they offer. If it is, go for it. If not, don't. Just like, you know, every other good or service you purchase in both your personal and professional life.
It costs me less than $5 with my current configuration (which is also a global CDN). It's way over our "soft limit", but it's an awesome site so I don't care. The important part is, I don't have to care.
This isn't about slightly more expensive tacos. It's about spending $560 on tacos instead of $5. "Meh" wouldn't exactly be my first reaction to getting that bill at the taco cart (or the fanciest taco place in the world, for that matter).
I can get IP transit in datacenters from $240-$600/Gb right now. So even his $960 transit cost for datacenters is off by quite a bit. He's comparing with a pretty high price and it still looks ridiculous.
And soda has a 1,150% markup [1], so it actually is like that.
[1] http://www.businessinsider.com/products-high-markups-2014-7
P.S. You could host it on Github for even less, if you were really price conscious.
> P.S. You could host it on Github for even less.
I'm not going to attempt to save an inconsequential amount of money by hosting a site I'm hosting on another host.
But that strategy probably wouldn't end well. The $560 site uses 25x more bandwidth than Github's fairly low soft limit of 100GB (https://help.github.com/articles/what-is-github-pages/) and is likely above the 1GB site limit (which is also just the sum of any changes ever, because git) of a Github pages site.
From seeing who their CDN provider is (one that charges basically the same CDN rates as AWS), my guess is that Github is paying a lot more for CDN transit than I am. It's perhaps not a coincidence that the 100GB BW limit was quietly introduced after they started using the CDN.
We went from cloudfront -> edgecast -> keycdn and our bill dropped from $3000 -> $500 -> $120
Edgecast when you buy is through a reseller is quite cheap but we moved because we needed custom domain SSL which is quite expensive in edgecast
And mixed drinks are more like a 4x markup. Lower at the high end even. Not sure where they pulled that number from. If it's anything like the coffee total, they'er comparing the garbage ingredients people typically use at home against the highest quality ingredients in the best places.
Neocities isn't just for static websites? How much traffic are getting that it would $560 on s3?
Is the basic principle to use cheap cloud VPSes or dedicated hosting (e.g. of the Digital Ocean, Hetzner, or OVH ilk)? And then you can afford lots of servers, and have then geographically distributed?
I notice that I drink less (too little, even, if I'm out for an evening) because of these prices. If each drink costs a few euros for 200ml glasses, I'd be crazy to drink a liter if at home the cost is somewhere between negligible (tap water) and 1.50 for that whole liter.
What I don't understand is why they can't spread out the profit margin. It'd make me feel a lot better about buying drinks and I'd probably spend more than I do now. Courses are a few euros in ingredients but I don't hesitate to pay €9-15 for it because I know labor is involved and someone has to pay rent. For drinks there is hardly any labor or ingredient cost (assuming all I ask is tap water, which is often since Dutch tap water tastes better than bottled water, and it's not cooled). Prices might also be higher because people don't order it in great quantities, which they might if restaurants wouldn't inflate drink prices to pay for rent (and other such costs), which would be so if you just pay per hour for a seat or something. I heard some places in Italy do that and people were surprised to be charged for it, but it's just itemizing costs and in my opinion more transparent.
All that said, the operating margins on high end restaurants are ridiculously low. Franchises and fast food do better at this (economy of scale), but if it is a proper linen-table-cloths and local sourcing sort of place, they are playing with very thin margins at the best of times.
A basic 1Gbps commit on a 10Gbps port in a data center might cost you from $0.50/Mbps (something like Cogent) to maybe $1.50/Mbps (let's say Level 3), other providers could be $4+/Mbps. By the time you factor in all of the above overhead costs, the true cost of the bandwidth is much much higher on a per Mbps basis.
Don't forget to significantly over-build your stuff, or you might get knocked off-line for anomalies or DoS attacks.
Admittedly, the scale of Google, AWS, Azure makes the cost per Mbps much much lower, but when as others have pointed out, AWS, Google, Azure don't need to charge less than they do.
Another possible explanation is that the ratio of available (or build-out-able) compute power to available bandwidth they have is such that they just can't give you cheap bandwidth and their pricing reflects that.
I suppose part of the issue is marketing. The customer reads the bill as "bandwidth." Would consumers understand better if there was a separate "network infrastructure" line item? Would they accept differntial pricing for network quality or class? Dont see why we'd find out as long as theres no downward pricing pressure.
Most facilities and providers don't require you to have your own ARIN registration and block. In fact it's quite difficult to get even the smallest blocks right now. Instead, an ISP will lease you IPs usually at around $1/month.
A pair of x86-64 boxes running OpenBSD can easily virtual switch, route and firewall well above 1gbps. They can announce BGP if needed, if you have a delegation or are muli-homing. A pair of high quality 10g switches may push a few grand on the secondary market but you can ~get away~ with Ubiquiti gear.
This kind of setup can scale up pretty high by dropping the OpenBSD boxes from the forwarding plane and just using them as route reflectors into higher end switches.
There is a slight labor disadvantage that amortizes quickly. i.e. setting up switches and routers isn't hard, and if you aren't sysadmin skilled enough to do that in a few days you probably shouldn't be running VM infrastructure either and retreat to something like Heroku, Lambda or GAE where there is much less room for footshots.
So TL;DR $300/mo and 5k CapEx will run a company, double or triple that for DR of a serious startup. Reap significant saving as a larger company by partnering or outsourcing remote hands and basic ops with a company like yours for the day to day stuff.
I'm currently paying 0,017 USD for high quality transfer, with a managed network. I could get even lower pricing by switching to 95th-percentile. Maintenance on server hardware is damn cheap with remote hands and hardware warranty.
It would cost me over five times as much to use GCP with 50 TB traffic/month, all things considered.
Cloud providers are cool, but bandwidth is still very overpriced.
Also, your transit prices are way higher that reality, at least for Europe.
Cost of IPs? Memberships? Wtf are you talking about? Cross connects? This stuff is all free, not needed or cheap and a one time fee.
Let alone the cost for DDoS mitigation. Thats an easy 300k€ worth of equipment plus lets say 2x40 GBit/s links for lets say 40k€ each per link and month. Over the course of two years thats easy 1 million euros, just to handle DDoS. Even if you just need 10 GBit/s capacity, you might need 10x that capacity in a DDoS situtation.
Or at least I've never stumbled across a blog post with anything like a line item cost range for everything that goes into DC or cloud networking.
And in the absence of transparency, yeah, people are going to assume they're getting screwed.
(I understand why this is. The networking free market seems to do a decent job at fulfilling needs, but it turns all those things into secret sauce that shouldn't be shared.)
The costs of doing those (well) yourself are not cheap.
Getting them from a provider that's certified to do them well while giving you software control also isn't cheap.
You're comparing cost of gas per gallon to to expense of miles driven per gallon. Pretty sure on your IRS or corporate expense report those aren't the same.
It takes times and skill to turn the ingredients into something useful.
Amazon, Google, and Microsoft are IaaS/PaaS, not ISPs.
Sell food at cost!
/s
When you put it all together, you're not paying for bandwidth, you're paying for the power and convenience of a robust, high-quality, high-availability solution that provides extraordinary benefits when looked at holistically.
For example: A dedicated m4.16xLarge EC2 instance in AWS is $3987/month. You could build that same server for $15,000 through Dell, lease it at $400/month (OpEx), and colo it with a 1GB/s blended bandwidth connection billed at the 95th percentile for $150/month.
IMHO the big cloud providers only make sense if you have a highly elastic load or if you are making extensive use of their as-a-service higher level offerings. If you're running baseline load bare metal destroys them.
And then magically maintain it for free?
Yeah, at the larger sizes it might make sense, but not for everyone.
I worked for a MSP for many years, and was well acquainted with our regional competitors. Our uptime and theirs was nowhere near AWS-level.
You're getting a different engineer every time you pick up the phone. They most certainly do not know your systems like an internal team would. I regularly saw cascading/circular issues caused by lack of familiarity and/or poor change management.
Owning and administering your own iron makes sense at a certain scale, but it's a much bigger scale than most companies will reach.
What do you do when they're on vacation/sick/sleeping?
Thing is, I really want us to host our own stuff. For one thing if you don't host anything you lose all competence and then you really are beholden to your cloud provider.
But it seems to be that we are in this situation until someone figures out how to comodify servers - and server orchestration - to the point where this stuff gets cheaper to manage again.
This pendulum always swings. No sooner will we control our own hardware again than someone will come up with a new way to centralize it.
With a bare metal server you have to manage all of that.
It's not a bed of roses in the physical world either, but you're simply wrong to say it's easy in the "Cloud" and hard in a DC/Bare Metal.
I find it amusing that people pretend without bare metal, you'll get all of that person's time back.
It really, really depends on on your use case, but there are some scenarios where bare metal just makes more sense. Also there are ones where cloud makes everything much easier.
North America-oriented B2B app that's a JavaScript SPA witch connects to some lightweight services fronting a database? I think that's a great contender for the cloud.
500-node Hadoop cluster running 24/7 with heavy load? You are going to see some huge savings by maintaining a Datacenter. You can use the extra 1.7MM/mo. to set up some robust networking equipment and have few people on-site.
Like every IT trend, "Cloud ops" is the trendy mind-slug that eats everyone's brain and makes them say that everyone who doesn't jump on the bandwagon is going to die broke and unloved in the gutter.
Remember how XML meant massively reduced integration costs and firing all your developers because GUI tools would let business managers connect pretty boxes and then kick back to watch the graphs go up and to the right as business exploded?
Yeah, so AWS is great, if you're building something that works well on AWS. And that's a large class of apps - if you're throwing up a Drupal front end to a CRUD line-of-business backend, need to push notifications and send some email, sure. You're dead in the middle of their target market.
Doing something interesting? Doesn't even have to be as massive as the parent suggests - if you're doing something that blows any of the billable parameters out of the sweet-spot, like bandwidth, storage, latency requirements, things that depend on system-locality, etc., you're much better off building, and using AWS for the pieces which can be broken off that don't hit the pain points.
Minimum wage => You get what you pay for.
Using EBS volumes with provisioned IOPS give you some advantages that you don't get with a bare SSD -- like snapshot support and the ability to easily detach it and move it to another server.
http://www.aerospike.com/benchmarks/amazon-ec2-i3-performanc...
Given the price of the server and the data you have on it, it's important enough to warrant a hot spare.
But if you are comparing it to a purchased server, you ought to be looking at Reserve pricing, a 1 year reserved instance is $16K/year, 3 years is $32K (or you can pay monthly if you want).
Still not as cheap as owning the server, but not quite as bad.
edit: I had selected a dedicated instance in order to compare apples to apples. By definition you don't share a colo server.
https://aws.amazon.com/ec2/pricing/on-demand/ https://aws.amazon.com/ec2/purchasing-options/dedicated-inst...
For smaller instances, there is around a 10% increase for dedicated, for example, for an m4.4xl, the price goes to $0.88 for dedicated tenancy versus $0.80 for default.
There is a $2/hour fee for each region in which you have dedicated instances, but that gets lost in the noise if you're using significant AWS resources. (it's $2/hour regardless of how many instances you're running)
We haven't noticed any significant difference between performance on dedicated and non-dedicated instances, and only use dedicated instances where required for compliance/regulatory reasons.
My smallest physical servers have 8 physical cores.
It is the best extension.
2. Will Dell servers arrived with 0 problems on its components? Any time I ordered meaningful amount of servers, I usually get about 5% fail rate. VMs aren't perfect but you can destroy and recreate in different region almost real time.
3. Ever had to deal with difficult-to-work-with network admin? The cloud is significantly less pain in the ass.
There are network engineers in the world who don't live in caves and gnaw at the bones of administrative assistants.
And Amazon allowed to to achieve amazing things that never would have been possible and that's without the idea of things like Kinesis or Redshift that we couldn't have achieved easily/at all even without personnel problems. You simply could never get 3 new servers or drastically respec an existing one in under a day let alone minutes. We were a small shop so it's not like we had spare servers ready to go.
Networking is a great place for assholes to build empires and exert control. In 20 years of professional engagement with distributed and data center networks, most of these guys (and they are almost always guys) running network orgs are an impediment and spend more time helping out their vendor of choice than anything else.
I've run into awesome network guys in position of power only a few times, and they are 100x employees. The last major rollout of a service that I did was literally 3 months ahead of schedule solely because of the efforts of this awesome network dude snowman as recently put in charge.
Eventually replaced with someone helpful, it was amazing what we started to get done. Got lucky there we found the replacement and he stayed long enough to be there when needed.
Worst person I've ever worked with. When I tell stories at future jobs I fully expect people to not believe me. Hell, people who weren't there for earlier incidents often didn't believer them.
Just basic docker use (or something similar) would have made a HUGE difference. Parts of the department were sent to training but it didn't happen because it would have made him less indispensable.
At one point someone higher up (but not high enough) asked how many servers we had in Amazon and the answer like 10x our physical set.
"And only one of you manages them?"
"On the side, we don't usually have to touch them."
"Then why does it take 4 people for the other servers?"
(Collectively) "Yeah. Why IS that?"
Change control was a big problem too. "Random" changes, things that didn't change at all (even when we could prove it), stuff like that causes SO MANY issues, basically all avoidable.
Also, hiring people is a risk, and it doesn't happen instantly. When you have a gap of months without a sysadmin things are bad. Hiring a bad sysadmin is even worse. Not to mention the impossibility of hiring 1/2 or 3/4 of a sysadmin. There's scaling limits there. If you have the resources to DIY, that can be a key competitive advantage, but not everyone does.
It's no secret that AWS is a rip off on instance pricing. Use the competitors if you worry so much about pricing.
What are most people running anyway? Using this kind of behemoth server as a 'typical' example seems out of line.
AWS m4.16xlarge => $2592
Google base is slightly lower and it gives you 30% off automatically on instances that run 24/7 for the month.
That's typical of on premise deployments. You get big servers because they are so annoying to deploy you might as well fill everything possible in the box.
On the clouds, you would create a group of VM per service, with appropriate size for the service. It is significantly more efficient.
Cherry-picking a bit there. You're entering a 3-year commitment with the Dell lease, but not with the EC2 lease. Take a 3-year RI on that EC2 instance, and the price drops considerably. If you also take away the 'dedicated' requirement, the 3-year RI is about twice as expensive, not ten times.
Edit: correction from Johnny555 above: this top-tier server doesn't have a 'dedicated' option, so there's not an 'extra' cost for that. It's just twice the price, if you do a 3-year commitment (as you are doing for the dell)
It is interesting that they end up charging you this for _bandwidth_ though. Their staff time doesn't actually probably scale with bandwidth, but their prices do.
I think it's probably kind of a way of pricing based on a proxy for customer size, likely revenue, and ability/willingness to pay. Rather than as a % margin on actual costs.
1) remote hands
2) network switches
3) hardware components you'll need for redundancy, I've personally replaced network cards because fans burned out or HDDs because of usual failure rate.
4) when I'm done using a n1-highcpu-32 I stop using it and not pay for it, or when I need 5 of those for 2 hours I provision them and stop when I'm done running a workload.
5) upgrades to hardware or additional hardware, intel released skylake, should I buy those and try to sell my Haswell on eBay?
I'm not sure it becomes a simple price of server + colo cost comparison.
With GCP (I use daily), AWS (I used to use) or Azure, what I'm paying for is someone else to focus on things like colo leases, capacity expansion, redundancy, physical security and provide higher order services - load balancers, BigTable/BigQuery, CloudSQL/RDS, GKE, DataProc.
This allows me to run all of infra as close to capacity as possible while avoiding any wasted cycles
However I do miss going to data centers and hear my machines hum and blink, that is quite the feeling.
Timeline: 120 days procurement (including product evaluation and PoC), 30 days install activity, 10 days buildout. 10 days software build, 15 days testing.
That's a little over 6 months. A ton of vendor/build engineering/tester time. And we'll be at 50% capacity for another 4 months. Or... we could spin the whole thing up on AWS/GCE/Azure in < 30 days and only consume what we need when we need it.
I've analyzed the cost of cloud services to death (I've worked for a couple of them) and the only way they aren't great deals is if you don't need high quality operations (i.e. if you can deal with slow-downs or occasional outages then you can do better elsewhere). Otherwise, if you're small-scale then these marginal cost differences don't matter, and if you're larger scale then call up these cloud providers and get yourself a discount off the list price.
I personally don't see the outrage. AmaGoogSoft overcharges for data transfers because they know they can get away with it and that lowering it won't attract more customers.
Customers with transfer-heavy applications will always buy their servers from providers with unlimited transfers like OVH[0][1], where you can do hundreds of terabytes a month with no extra charges (1.5 Gbps * 3600 * 24 * 30 = 486 TB). Even if AmaGoogSoft lowered their transfer prices by 100 fold their pricing still can't compete with OVH.
Companies with enough engineering resources can always go with the best of both worlds: transfer-heavy servers on OVH, and "regular" servers on AmaGoogSoft. The expensive data transfers will only hit smaller outfits, but these customers won't switch because it's not worth the hassle to split your hosting across two providers.
[0] https://www.ovh.com/us/private-cloud/options/bandwidth.xml
I know a handful of scientists, myself included, who would consider cloud computing were it not for the expensive egress costs. The ease of spinning up lots of computing power for scientific modelling is useless if retrieving the vast quantities of raw output data is costly. I will have to investigate OVH though as a potential opportunity--thanks.
Scientific computing -- you've probably got a ton of data, and your 'revenue' from project-based grants doesn't really scale with the amount of data and other computing resource needs you have, at least not to the extent it does for ecommerce.
Possibly AWS could make money on scientific data at prices you could afford... but they probably make the _most_ money by targeting their pricing model at ecommerce, which has the most money to spend.
Possibly there's a business model for a cloud-provider focusing on scientific computing. I mean, there probably are such providers already I don't know about? Maybe? But it'd be a tough business, the people paying are pretty budgetarily constrained, their budgets don't neccesarily go proportionaly as their resource needs go up, they traditionally are (ironically) averse to innovation in IT, there might not be enough likely customers to give the provider the scale they need to make money, etc.
Capitalism!
In fact, the federal government in its current business-oriented mindset may even lean towards a contractor handled machine to replace ones like Yellowstone at NCAR or those at NOAA NCEP if the price was right.
Currently, I suspect a lot of IT-heavy grant-funded projects are basically relying on university infrastructure being provided without full (or in some cases any) project-based cost recoup, which universities are increasingly loathe to do. The indirect/overhead margin probably doesn't always really cover it. (And it's not like a university is going to let you use 'overhead' dollars for AWS, _even if_ it would actually be a cost savings compared to the local IT you are using instead!)
It's not really rational economic decisions being made all around.
Proper professional IT is expensive. People still think automating is supposed to save them money, but turns out, nope. Many industries are still relying on unprofessionally provided IT instead of paying for professional IT.
You're likely not running a model just for yourself. You have collaborators at other institutions that need data to compare or combine with something else.
Your grant may require you distribute the output for use by other researchers. That could either entail being hosted by you or by an agency or other entity. But you still have to get the data to them.
A reviewer for your publications may request the data.
That brings up another point that the review process can be upwards of a year.
Oftentimes you have to go through an exploratory data analysis, where you don't even know what your final analysis will entail.
I've done the math many times and it's orders of magnitude cheaper to colocate as long as you can afford an IT guy and the upfront cost of hardware.
What happens when a harddrive fails in your colo? Unless you have best practice backups, you will lose customer data and trust.
GCloud abstracts the issue of hardware so you can focus on real business development. And to some, that's worth the cost until they can properly afford all the hidden costs of independent hosting.
Really though in my experience the cloud is far less reliable than colocation. I've had AWS VMs "degrade" or randomly die countless times but I can't remember the last time my Colo boxes went down, probably not since I last upgraded the OS.
It may be. But for many people, it does not matter.
Let's say you work at the place (I do) where "compute" cost is like ~0.5% of the total costs. Almost all of the cost is employee's salaries and benefits.
BTW we get (and from what I hear, everyone does) offers from other cloud vendors to switch, and get 1 year free or whatever. And it gets ignored because it is a big risk for a tiny possible gain.
It's really just a new form of vertical integration. If you're working only at the high levels you're not only paying someone to do that rest for you, those people are making profit off of you. Profit that could be yours.
(From your browser to multiple Google zones)
- assume $100/TB for cloud data transfer
- assume one employee full time equivalent to manage colo'd servers ($10,000/mo), plus $30/TB data transfer
The break even point for the colo'd setup from a networking perspective is:
10,000 + 30X = 100X
X = 10,000/70 = 142 (TB/month)
At 1MB per "request" I believe this works out to about 50 requests per second average to reach this traffic level.Weaknesses of this model:
- Data transfer only. Depending on what else you're doing you could also save a lot on compute and storage.
- I don't know that much about how colocated data transfer would be priced ... i.e. do you need overprovision to guarantee availability, etc.
- one employee to handle servers to replicate the Amazon AWS experience ... could be highly variable depending on what AWS features you are using.
It's useful to keep in mind that in this type of scenario the goal is not to replace a cloud provider but to leverage capital to get better bang for the buck on select services. ex. using in-house compute farms and leveraging s3 for storage
For the employee cost I agree if you just have a couple boxes pumping out data your labor cost will be quite low. I'm imagining what else you might have happening on your back end using AWS services that you'd also have to replicate in your colocated installation to preserve the efficiency, which would take labor. For example, if you want some extra machines to do some load testing. Or if you need to store a bunch of data and want an S3-like system. Or if you want to wipe a machine and reinstall from an "image". The backup system, and backups testing. RDS, Aurora, DynamoDB, etc. Most of that can be done with a couple clicks on AWS.
Even then, cloud bandwidth is insanely expensive. For example, Hetzner offers 1.3$/TB (if you happen to exceed their generous 30 TB quota). In comparison, Amazon is 70x more expensive at 90$/TB.
Sure, if high availability is critical for you business and money is not an issue, it may be worth it. YMMV.
;)
Still, recouping the cost of the NIC by making internal bandwidth free and Internet transfers really expensive is kind of a weird way to do it.
In any case, the hardware components are fixed costs. What's not fixed is the cost of applying routing and segmentation rules for each packet. There are scarce resources involved here: the hardware has limited memory and limited cycles. You're paying for the use of those scarce resources (e.g., having a memory resident stateful firewall rule for a TCP session). Charging for bandwidth isn't perfect, but it probably correlates pretty closely with the underlying resources and is a much more intuitive unit-of-value for customers.
AWS, GCE, azure seem like the platforms of yesteryear in all dimensions when compared to something like packet.net. I think these providers could be in a rock and a hard place due to the unsuitability of native Linux containers for secure multi-tenancy. This does leave a nice runway for Joyent as both a provider and software vendor for at least a little bit, but I think packet.net is really going to change the economy of infra.
I wrote a blog post on Google Cloud latency and pricing across zones and regions which may be useful for others:
https://blog.elasticbyte.net/comparing-bandwidth-prices-and-...
To add to the mix - what if you need a multiple data-center deployment/replication? Both Amazon and Google will provide you a greatly discounted traffic there $0.02 / $0.01. And that's only start. You can easily migrate from one data center to another, with no or little cost attached to it (try that in colo).
Then it all repeats next week.
In order for this comparison to be valid you'd need to get 100% utilization of your colo or Google Fiber pipe. You only pay for what you use with AWS et al. And quite obviously the pricing of GF and Amazon Lightsail assumes less than 100%. Nobody's getting "screwed".
People fail to realize the true cost of operating on S3, specifically when hundreds of TB of usable data is in play
"By putting the "tax" on bandwidth, a lot of these business cases are solved. I see why Amazon does that. AWS is great, but as you get into high scale (specifically in storage - 2PB+), it becomes extremely cost prohibitive."
The real number to compare to is the google data for business rate. You can do lots of colocation in that price range. And that is why the cloud prices are unreasonable.
Take credit card churning and apply it to cloud data. Build tools to seamlessly move apps between cloud providers.
[0] https://www.linode.com/pricing
[1] https://www.vultr.com/pricing/Maybe worth renting an office in Provo just to get the deal.
Disagree. There's no false advertising here, they're making you pay for their service and convenience of using a combined [Paas, Iaas, Saas ..etc]. It's unfair to view these services as a singular function, you typically touch MANY features/products in production. The cost includes the convenience of offering everything under one roof, because, face it, doing everything by yourself at the SLAs provided by the giants is no trivial task.
Unless you're a BIG company that likes to distract itself with infrastructure instead of building and sharpening the core offerings, chances are that you will NEVER really build anything as reliable, inter-operable, configurable and manageable at cost.
A streaming video service, for example, might end up paying $0.10/user/hour. So a movie-a-night customer would cost $6. What multiple of reasonable do you think that is?
Primarily, until 2 years ago we did lean on S3 for all block storage, but most of the rest of the infrastructure (metadata storage, etc) ran in our own datacenters.
Your point I think you're getting at sounds like something I'd agree with though -- you can wait a bit the cost efficiency starts to be what is important/impactful to work on before shifting your usage away from some of these providers.
I don't know if it's true or not but I heard the story.
Fortunately, it's relatively easy to locate some of your very high egress services outside of AWS. I discourage that for casual optimization, but when you get to Netflix/20 scale, that makes sense to trifle with.
“The best way to express it is that everything you see on Netflix up until the play button is on AWS, the actual video is delivered through our CDN,” Netflix spokesperson Joris Evers said.[1]
[1] http://www.networkworld.com/article/3037428/cloud-computing/...
Let them charge for extra for the 'value adds', what is the justification for adding these costs to the bandwidth specifically?
I think people raising the issue are just seeking more transparent pricing.
There's a lot of koolaid being thrown around by the companies with cloud to sell. Unfortunately these also happen to be the big "market leaders" so it's hard to deny what they're saying and get taken seriously.
It's like six sigma, agile, or stack ranking. The big guys are doing it so it must be the right thing to do... Right? Until everyone realized it's mostly a ploy to make money selling books and conference tickets, or in this case rent out a bunch of excess capacity for huge profit.
I disagree with reliability as well. Most of my bare metal and colocated machines have uptime of many years. Most AWS VM's die after a year or two. With the cloud you have to worry a lot more about fault tolerance whereas with dedicated equipment simple offline backups are often enough to meet any reasonable SLA.
Before all of these providers existed:
1) The volume of internet traffic that exists today didn't exist then. Mobile devices didn't exist. Mobile devices that aren't always connected to the internet and consume hours of our day didn't exist. Large downloads (1GB) didn't exist. They couldn't exist because the infrastructure that exists now can properly support it...at scale.
2) Websites that had large volumes of traffic had top tier expensive admins to maintain them (surprise surprise, Amazon did and turned it into a service)
> Most of my bare metal and colocated machines have uptime of many years. Most AWS VM's die after a year or two.
3) What's your definition of "most"? Sources for those numbers please?
I would guess a single server can handle thousands of times as many users as it could handle years ago. HAProxy, Netty, Nginx, and others can handle over a million (simple) HTTP requests per second. That's more requests than Google.com gets.
Most as in I've been watching over 100 AWS VM's and maybe 30 on Azure for years and they die or crash far more often than than the VM's hosted here, at colo, or our old bare metal machines. It's anecdotal but it seems like AWS doesn't really care about warning you before shutting off your machine. Azure is slightly better but still goes down regularly.
I know everyone says "it's okay! Just make your servers fault tolerant!". Well that works great for load balancers and frontend, but doesn't work at all for SQL databases. ACID compliant transactions require a single source of truth and a true multi master SQL database is impossible. Failover yes, but you always risk losing data in the switchover unless you use two phase commit which actually makes your multimaster database slower than a single system. In practice the failover almost always causes some data loss and log conflicts you have to diddle with later. And God help you if the replica falls behind more than a couple seconds.
Anyways, for SQL databases system reliability is as essential as ever and it's a lot easier to get high SLA numbers when you control the hardware and the power switch. The closest you can get to the Holy Grail is running KVM VMs locally and doing live machine migrations when hardware starts to fail, but even that won't keep your database running if something really bad happens.
Really consider how important 100% uptime is though. Google and S3 have gone down multiple times without killing the internet or losing a ton of customers. Plenty of large SaaS providers still use maintenance windows. Heck, GitHub went down today. Not sure if you use ADP but that goes down for a couple days a week!
I know it's not the popular thing to do but you can get much better relative reliability by running a single database per tenant and running a limited number of tenants per VM.
> Most of my bare metal and colocated machines have uptime of many years.
Nobody cares about host uptime, that just means you're not applying security updates. You're lamenting that you have to worry about fault tolerance in the cloud when the entire point is to have throw away instances!
Um... Only kernel updates used to require a reboot, and production security-critical ones are fairly rare.
Solutions like ksplice, kexec, kpatch, or kgraft have existed since about 2011, and now the Linux kernel has first-class support for "no reboot" updates.
How do you handle segmentation of workload among thousands of servers? Provide firewall services? Meter bandwidth? Provide redundancy to protect against failure of a switch or any other single network component? Provide redundant transit via multiple ISPs?
The answer historically is that you would buy a bunch of stuff from Cisco, pay through the nose for TAC, and hire a team of network engineers who may or may not be idiots. One place I worked at lived with a 30 day eta in firewall changes.
Your server that has a long uptime is a risk in a large org. It's obviously not patched, nobody knows how to configure it again. Different customers have different needs.
Firewall is on the front end load balancers. Linux is one of the best firewalls you can get if you configure it right.
Redundant L2 switches past that running RIPV2 or OSPF for routing. I've found that crappy consumer quality switches usually work fine, as long as you wire them in parallel. You can do dumb things like wire them quad redundant and it just works.
Redundant transit handled by datacenter level multihoming.
Extra redundancy using dual data centers if you want with DNS round robin.
Buy dedicated lines to avoid bandwidth costs.
You can rent a cage and do all this with used last gen dell workstations running Ubuntu. Get some last gen fibre NIC's off eBay too. Total cost of maybe 2k in hardware to run top 500 site levels of traffic...as long as you're not using something piss slow like PHP.
Historically how shitty your setup is depended on how good your network guys are, the cloud just took that out of the equation.
Now everybody can have a great setup as long as they pay out the nose.
Edit: Also you can patch Linux without reboots, even the kernel. I don't remember the last time I had to reboot from patching.
You do need to reboot for major version upgrades but if you stick with LTS versions you're usually good for 5-10 years
I don't see any claims about false advertising. Just complaints about one aspect of the service being very overpriced. Sounds like you consider it worth it for your application, and that's great. But it is simply not priced appropriately for a lot of services, which naturally go elsewhere.
It seems obvious to me that bandwidth is Azon's proxy measure for "utility". Lots of companies do this - Oracle uses core-count as a proxy for the "utility" you get from their software, for example. But like any proxy for a different variable, it is going to be wrong; sometimes a little, sometimes a lot.
And you hardly need to be a "BIG" company to choose something other than cloud hosting. My last two looked at and rejected them. Neither qualify as even medium-sized. Both are extremely cost-aware. Running our own data centers is cost-competitive at my current gig and was much, much cheaper at the last, which was very bandwidth-heavy.
I'm merely speculating, but if you have a idea what is the possible difficult to price component, can anyone help with any suggestions?
Edit to suggest a similitude: my job is to help people understand the price of (print) advertising. This is often intractable. To a certain extent, a good deal of people involved just throw costs into commission rates and other factors that make understanding the price models difficult and sometimes completely opaque. Is it impossible to imagine that the cost of cloud bandwidth is nothing less innocent than "put our non itemized expenses here'?
(Note: I worked for Google until 3 days ago.)
Not only that, but Google's datacenters have 24/7 attention, lots of redundant providers, etc.. A bargain basement colo won't be nearly as reliable.
The benefit of this has can't be understated. In SE Asia connecting to their Taiwan region from most places I'm literally just at the mercy of a few local hops at level 3 due to their extensive remote peering. Staying on the Google network is also interesting to observe when connecting between Google Compute data centers.
AWS and Azure behave considerably different.
Not everyone need a ritzy ultra reliable network. It costs way too much and hardly any customer is going to get comparable value out of it.