EC2 Price Reduction (C4, M4, and R3 Instances)
aws.amazon.com
aws.amazon.com
Honestly, the high bandwidth costs have killed a number of ideas I would have liked to tinker with on AWS. I find the cost of EC2 to be quite good, and storage cost is tolerable, but $.10/GB of bandwidth-out is horrific.
http://www.zdnet.com/article/cloud-firm-linode-resets-user-p...
I have been a big fan of linode in the past, I still have a fair amount with them but honestly at this point I'm seriously considering AWS even at the increased cost (which compared to my time isn't that large really and reputational cost which is).
I found out on Slashdot/Reddit first and got the email from Linode about 3 days later. I have known ever since then that Linode is the type of provider you don't want to be involved with. A VPS provider has one responsibility above all else: be honest and transparent.
For example, the $20 a month VPS on Linode offers only 250 Mbps connectivity (outgoing bandwidth). So as soon as you saturate that, you need one more VPS to handle the load. So this leaves a lot of bandwidth quota unused. Especially since in the real world traffic tends to come in spikes. AWS is similar, depending on the instance size the network connectivity differs. This makes the calculations for estimating costs for hosting even more difficult.
Softlayer no longer includes 5TB on every virtual server and their bandwidth is equally outrageous so I'm still evaluating a plan for new servers.
That is to say, the reason you can negotiate a discount once you get well past that point (or if you intend to grow well past that point) is because staying on AWS makes increasingly less sense for you.
Not anymore. In the US, you can now write off $500K/year in equipment costs immediately.
> It also means shifting expenses from monthly as-you-go to mostly upfront.
Dedicated server provider or lease the equipment instead of buy.
That is a lot of negotiation that needs to be done. Will they get anywhere close to that kind of discount?
How big a customer do you need to be?
(Speculation based on how AWS prices things: Build big, charge for everything, charge what it actually costs with a small markup, and put every high-margin provider out of business)
So if what you're saying is right, that would be just another reason for most small companies to avoid Amazon. If you're wrong, other comments in this thread have also given a good reason for small companies to avoid Amazon.
Some people even take that "confidential" offer letter for volume pricing and shop it around to other cloud providers. Shocker!
There are some growth incentives (eg, prepay $3MM for next year, get 6% off all services), and there are some tiered savings (e.g. 5% off on more than $500k of RIs in a single region), but I've never seen a plausible, verifiable incident of people successfully negotiating AWS rates downward.
Do you have any evidence of this at all?
Maybe they're milking your investors. Maybe you've built some absurd contraption that they know you won't be able to move elsewhere. Or maybe you need to hire a better negotiator.
Given that we couldn't get it; and that nobody from our company's entire VC network could get it; and that nobody on my CTO email list could get it... I'm going to continue thinking that the anonymous HN commenter is talking out of his ass.
But if you can prove me wrong, I'd be thrilled to be wrong.
Anonymous commenter versus 11-hour old throwaway account... who wins?
Looking at a few words in your rebuttal above and combining it with the tone you're using, I can come up with quite a few reasons you might not have gotten the discount you sought.
I'd suggest you put down the "CTO email list" and let an actual business person handle this for you. Handshakes and existing relationships still matter in this world. Good luck!
That said, no point in arguing further with a liar and/or troll. You're not getting 10x discounts on your AWS bill. You're not fooling me or anybody else.
What would use this much computing power, besides bitcoin mining, or popular top 100 site? Maybe an online game? Just curious.
When situations like this come up I always suggest people read the history of Western Electric. Talk about scale!
As such, $500k/mo isn't a terribly elite club. Essentially every SaaS company that does more than $10MM/mo of revenue is spending as much or more than we are.
I quit Softlayer because every time I sent a change order in something awful would happen. At one point the issue tracking system broke in a way that I could not put in new tickets and got to talk to four different people until I talked to a wizard who punched a few commands into the SQL monitor and told me he saw something "amusing".
I had a "near miss" at data loss because one of their techs botched adding another hard drive to the machine, plus I was dealing with an expensive and balky backup system so I immediately moved my data into S3 then all the servers into EC2.
Softlayer had a crack sales guy call me to try to get me back and I told him I had a day job and a night job and I don't have time to talk to minions to fix the problems they make for me. He brought up the egress cost issue and I told him flatly that "I make $1000 a month in ad revenue and I pay $30 in egress charges so I don't care."
My business situation has changed in many ways since then but I'd say that my egress charges tend to run between 5-10% of my total spend so it is not a concern for me. If there is any AWS service that I don't like the value of it is RDS and I am mostly off it since I have been using SSD-backed instances and first running local copies of MySQL and then ditched MySQL.
We have been working on a POC in AWS and it's such a breath of fresh air. Things work as advertised, the provisioning process is quick and everything can easily be done with an API call. No more waiting for support tickets. The freedom you get with your network routing in an AWS VPC is worth it to me.
At the risk of sounding too much like an AWS fanboy we are actually looking at scrapping it all and just bringing everything in-house running on our own physical boxes with some type of hypervisor on top.
We have done just that, with Opennebula (qemu,kvm; networking built on top of openvswitch and storage from distributed iSCSI NAS appliances). It's a perfect middle ground. All the easy provisioning niceness, less than half the price of AWS.
One rollup just prior to Softlayer was ThePlanet and their efforts to onramp us from their vanilla dedicated to their new managed service several years ago was a complete debacle. And our environment was super simple--only 5 boxes on a rack and a couple network devices.
But, the guy who was "leading" the effort was completely inexperienced. When we finally announced that we'd had enough, they brought in more management and senior tech guys to advise and save the deal (they were angling for an investment so wanted to book more customers before quarter's end; thus our little business mattered).
They talked us into staying and one of our conditions was that they replace the "lead". Oddly, they asked if they could leave him in place because he was "a young guy, just getting started, and the blow would set him back". Of course, I felt for the guy, but that struck me as a horrible thing to ask of a customer. Didn't want to hurt him but had to insist nonetheless.
In any case, we stayed with them for some time at close to legacy prices on fairly dated metal, primarily because we didn't have time to switch. Over that time, services became decidedly "less managed", especially after the rollup to Softlayer. Of course, by then they were also pushing their cloud. When we finally found time to switch, we moved to AWS and never looked back (except in relief).
We cut our costs by two-thirds. Better, it struck me that AWS's automated processes are an order of magnitude better than the "managed" services we were by then receiving from Softlayer.
There's plenty of dedi providers that are big enough to prevent real outages though. And it doesn't protect you as a customer, it just protects you from other Amazon customers getting ddossed, but Amazon doesn't keep your website reachable if you're getting ddossed (no dedi providers do).
I think you are seriously underestimating the size of Amazon's or Google's networks.
And Amazon, I guarantee you, doesn't have that much unused network capacity. That would be very, very expensive at their scale.
Also, I'm not sure I buy the argument Amazon has small unused capacity in absolute value, relative maybe. They could also have 5% across 23 locations so if your application is distributed it can have even better resiliency.
If it's expensive for Amazon to have multiple Tbps of unused capacity, imagine how expensive it is for any other provider. To match the absolute spare bandwidth of Amazon having only 1/10 extra network capacity, another cloud provider might need to have keep its network utilization at only 5%. Maintaining a network capable of serving over 20x your current utilization "just in case of DDoS" is bloody expensive.
[0] http://www.businessinsider.com/nobody-gets-fired-for-buying-...
http://googlecloudplatform.blogspot.com/2015/06/A-Look-Insid...
One major difference between Google and AWS is that Google will carry your packets between data centers on its backbone by default. Google will also carry packets as close to the customer as possible, whereas AWS will dump it off as quickly as possible.
So, even between Google Cloud and AWS it's not an apples to apples comparison.
It sounds like the reason for having that outside of your primary infrastructure (or more accurately, inside a cheaper bandwidth host) are lower bandwidth costs with the trade-off of some slower requests getting sent to your origin server when the cached resource expires/is invalidated.
Also, correct me if I'm wrong, but Moore's law doesn't say anything about ancillary costs like power/air conditioning. I have no idea what the trend is for those.
http://perspectives.mvdirona.com/2010/09/overall-data-center...
According to James Hamilton who speaks on behalf of Amazon, Servers make up ~57% of the cost of running a DC. CPUs are only a fraction of the cost of a server, maybe 25%.
So moores law only really applies to ~15% of the total cost of EC2. That's before factoring in all of Amazons capex to build the software that is EC2.
Edit: http://perspectives.mvdirona.com/2010/09/overall-data-center...
The capex on software is negligible, because the cost of duplication is effectively zero. AWS can scale to whatever size they want (as long as their software is architected correctly, of course), and not have to spend more cash.
That said, I'm not faulting AWS for charging what they charge. They have a remarkable service. It's just hard to stomach a post bragging about dropping the price by 5% when they're effectively printing money.
I have no education/insight at all into this matter, but I would keep that in mind and reconsider your assessment of Moore's law's direct influence on AWS' bottom line.
Doing exactly the same on a dedicated provider: 10$ per month. With more traffic than I did on ec2.
Also, it's an (atom powered) dedicated machine. Performance is far better. 4G memory instead of 1G. Disk space : 1T (of the rotating kind though), but of course I can have ramdisk now for most of the stuff. Compared to 100G SSD for 10$ on AWS. But egress traffic, that's what's costing me.
Downtime since switching : 0 (but I will agree that it's lower quality. Not 12x lower though).
They come into their own when you're running tens to hundreds of servers, with dependencies between each other, and a use for supporting services such as RDS and S3.
If you check out 3x redundancy, what does it get you ? Well, you can correct any 1 bit error (not 2, because you wouldn't know which version is the correct one). Hamming(7,4) with column encoding gets you the same (better in some ways even). Therefore would you really be lying to your customers if you told them you gave them 3x redundancy if you used Hamming(7,4), column encoded ? I'd say no. Because it gets them the same : any disk can fail, and you can rebuild the data.
If you intend to serve your customers correct bitstreams in the case of bitflips on the disks, you'd need to read 2 disks even in the case of 3x redundancy, exactly the same as in the Hamming case. Of course, people might choose not to do that, but then you only have backups, not redundancy. What can go wrong with 3x replication reading from one disk is that your system updates the 3 disks based on information read exclusively from disk 1, which may turn out to be wrong data.
But Hamming only costs you 175% storage, not 300%. That brings it to ~7 months. And with precomputed lookup tables Hamming decoding is far, far faster than reading from disk (even without I bet it would still beat it).
Another huge advantage Amazon has is that EBS means they don't have to allocate SSD space unless a customer actually uses it, not just if they reserve it (and they pay for it when reserving it). So in practice you do what ? 100% overprovisioning is prudent ? Let's say compression, given that these are operating system images mostly, gets you another 30-50% or so. If they dedupe, they could get far more.
On the other hand the newegg figure doesn't include power to actually use those disks (SSDs are cheap though). Amazon of course doesn't pay anywhere near full price there either. Then, actually putting stuff onto an EBS ... amazon charges for that. And of course, Amazon needs to develop a lot of software to make this happen. So there's various other things not counted here.
So hardware costs for Amazon would be at most 2-3 months or so until they're repaid, no more.
An SSD might use less power but you need more of them and the rest of the storage server won't change at all.
Finally, I'd love a citation for any compression + dedupe savings at the level you're seeing for large heterogeneous deployments, not to mention reliable performance at their scale.
Also, in our experience t2.* instances can't sustain any reasonable network traffic, but they do work wonderfully for lightweight RESTful systems. So any workload which serves mostly cached data and needs 6-7 GB per node is best off with m4.large. At least for Ireland, in our experience a single m4.large can keep up with bursts of ~120Mbps and sustain around 65Mbps. We have two as edge nodes for one of our public services, and will probably add a third one soon. Cutoff point for sustained bandwidth is slightly above 70Mbps, after that it starts to stutter. The t2.* instances choke and throttle bandwidth way earlier.
Finding the right instance type for a particular service always takes some time and experimentation.