Why we ditched Amazon AWS
blog.testingbot.com
blog.testingbot.com
I realize the cloud is just a marketing term and it doesn't really mean something, but if I move my single heroku dyno app to my own dedicated server sitting in my basement, did I just move it to my own private cloud?
If I start two VMs on my desktop and we call that a 'cloud' then the word 'cloud' isn't terribly useful in describing anything.
The whole point of a cloud to a business is that you can adapt to changing needs without having to pay for all the infrastructure all the time.
(You didn't think the suits talk about "the cloud" all the time because they fancy the technology, right?)
Amazon is paying for all the servers in the EC2 server farm. Does that stop EC2 from being a cloud?
TestingBot is spinning up VMs dynamically to scale out test infrastructure on demand. Their use case is basically the ideal cloud use scenario. They've built a system of servers that enables them to do all of the things typically associated with clouds, which means they've built a cloud.
You're stuck thinking of TestingBot needing a bunch of servers, and seeing the tradeoff as dedicated servers vs cloud servers. In reality, it's TestingBot's customers who need a bunch of servers. So TestingBot built a cloud for their customers. They aren't setting aside N servers for each customer. They're building a cloud with the capacity for X servers and carving out of that on-demand.
(Mind you, I'm still only discussing the financial side here, as this was the reason the original post cited. In technical terms, of course they are still a cloud, but so is the mainframe emulator on my Raspberry)
The point of a cloud is that you can adapt to changing needs by reallocating resources based on demands. For a public cloud, that's exchanging money for more storage/compute capacity/etc. with near-zero latency. For a private cloud, its not adding total capacity, but reallocating between different uses on demand.
Having worked in large organizations still primarily concerned about allocating physical servers to tasks -- which, just like one with a private cloud, still keeps substantial excess capacity to accomodate changing needs --I can immediately see that being able to reallocate resources with minimal latency is valuable even if its from a fixed pool where increasing the total size of the pool has substantial latency.
Latency of provisioning is time-to-market is financial.
> Sure, you can save a few physical servers by being smarter in allocating load
Saving physical servers isn't the main benefit. Being able to reallocate capacity with minimal delay and staff effort is the issue.
Once you plunk down all the money for networking and hardware yourself you've lost elasticity and will have a step function in cost the next time you want to scale. When you turn off VMs, you aren't saving anything at all. You just have expensive idle computing resources.
On day one, buying equipment, labor and software to stand up a computing environment is a risk. You don't truly understand your needs.
After a couple of years of operations, you should have a better idea of your typical run state and growth patterns. At that point the risk may shift. If I can deliver compute with AWS for $X, but can also deliver that compute through some other means for $X-Y, you have a new risk/cost evaluation to do.
The key is continuing to use the cloud thought process.
I also don't really agree with your claim that the key attribute of "the cloud" is offloading the risk and cost of scaling. From my perspective, "the cloud" is primarily about simplifying and automating scaling, thus reducing the human and time costs. It's not just about dumping the cost of physical machines onto someone else. I see taking less risk on machine purchases as a nice benefit, but not the core purpose.
[1] I haven't actually used TestingBot, so this isn't an endorsement. I have no idea if they actually do a good job or not.
Their customer can scale easily, they can not. Their customers don't have to pay for all the hardware needed, they do. Hence they are not using a cloud, they did not move to a cloud - they are using servers, they moved onto their own servers. The title is wrong.
Why is that important? Not only because it is a basic misconception, but because saying "we moved to our own cloud" misrepresents their current abilities regarding their ability to scale and their cost structure.
Yes it is. You represent the internet as a cloud on shitty visio diagrams. That is where putting stuff "in the cloud" comes from.
What you really needed to do was design algorithm that keeps machines running over period in time, according to a general trend. Not shutdown, and instantly bootup for every customer deciding to run tests.
I'd much sooner have a clean VM to run my tests on than have to think about the ways the previous user might have left things in an insecure state.
Not a Google shill by the way - I don't have a preference yet. However, for that particular complaint, Google has the solution.
Or so I remember from reading the ToS yesterday.
The command-line cloud utility is so much easier than Amazon's. Everything is in one place. ssh keys are managed automatically. In less than 1 hour I was able to do everything with Google, whereas it took me a couple days to feel comfortable with AWS.
http://weblogs.asp.net/scottgu/archive/2013/06/03/windows-az...
In what way?? I mean thsi honestly.
We use linux vms on azure with no problems at all. What non-msft tech stacks doesn't it "play well with"?
From my reading, azure seems to be pretty awesome as far as integration with visual studio, but is seriously lacking when it comes to features on their load balancer etc. Their security model is kinda weird too, where defaults open your firewall to the rest of azure commonly.
So that's just my ideas, what are your experiences?
have you seen AWS opsworks?
Fortunately there is also another side of Microsoft. A side which is doing fantastic work in the open source community as well as with things like the Linux support in Azure. I would be happy to talk more about Linux support in Azure and I am more than capable of helping out with any setup required, regardless of your operating system of choice.
If you're still interested, you can reach me at tistrimp@microsoft.com
I'm glad they found a solution which seems to work out really well for their use case. Changing the algorithm to use the machine for an hour would imply that some tests may wait to be run, and that could be a bad thing for TestingBot (I'm not sure what their SLA is).
Going to their own "cloud" though will require them to either overbuy capacity or eventually make a push to increase utilization, so they may end up making that algo anyhow.
There is another situation, which I would argue is the "enterprise" situation, where you will want to pay your programmers to make the changes because they're fixed overhead rather than capital expenditures. This scenario comes because of an internal push to save money (adding significant capacity costs too much), or because adding any sort of capacity increases operational complexity more than it increases code complexity. It's a much larger scale than where TestingBot is at, but a problem they probably hope to encounter one day.
I use S3 and DynamoDB regularly and think the pricing is better than most. I don't use it for pricing though. I use it because I don't have to worry about load balancing, adding multiple servers, running out of disk space. Set it up and forget about it.
Of course that's assuming you haven't locked yourself into amazon's databases/load balancer/... But that would be a horrible idea anyway, right up there with running a microsoft stack licenced per year.
http://www.holovaty.com/writing/aws-notes/
He just about had me sold.
Anyway, I'm definitely in the wanting to avoid sysadmin stuff if at all possible camp. Does anyone have any thoughts on AWS vs Heroku for Django? Does Adrian's solution seem reasonable?
Most developers should have no problem in AWS for small to midsize environments until you're scaling up.
I'm very much looking forward to moving our manually managed cross data center MySQL replication/admin to RDS. But I also don't have to scale MySQL to ridiculous levels either, I just need very high availability.
An EC2 instance still requires someone to set it up, install packages, button up security issues, deploy configurations & applications, etc.
E.g. you can have a nice app running with elastic beanstalk, file storage on S3, database on RDS, full text search on CloudSearch, without worrying once about apt/yum, incompatible JVM versions, backups, stuff in /etc or setting up iptables and ssh, which is what most people think of as sysadmin work.
Bad idea. Extraordinarily bad idea. Of course, I'll gladly charge you $450 an hour to recover your data when you accidentally drop a production table. It's part of how I make a living.
> setting up iptables
i.e. setting up security groups - same solution, better gui. Someone still needs to understand what it's doing.
Elastic Beanstalk... sounds cool, but based on the "let you take over management when you're ready", it's still a EC2 instance under the hood that's just configured for you. Scale at all, and you're probably going to have to start concerning yourself with all those bits you've been able to ignore until then.
AWS' automated offerings are good, but if you're doing anything at scale, they're frequently not good enough. And my experience has shown me that the line is really really painful to cross when you get to it with a fully staffed team, let alone a sole developer.
> i.e. setting up security groups - same solution, better gui. Someone still needs to understand what it's doing.
Yes, but like being able to install wordpress doesn't mean you know how to program, setting up beanstalk+rds is not on the same level of expertise as setting up N machines from the ground up. When I start an RDS/CloudSearch/DynamoDB instance the security groups are already setup in a sensible way.
Sure, there are plenty of cases in which what AWS offers is not good enough. Heck having an ops team is always better than not having it.
But I replied to AWS only getting rid of "someone who can physically access the box". That seems reductive.
The Xen hypervisor actually has pretty good resource allocation between virtual machines, and although there are academic attacks [1], I'm interested to hear what evidence you've seen of neighbors hogging your resources.
But I think if EC2 didn't let you borrow CPU, nobody would use it because they would realize how absurdly underpowered EC2 offerings really are.
yeah, that's pretty much the crux of virtualization; it's a lot like buying bandwidth. Yes, your upstream is oversubscribing. Yes, if they do this right, 99% of the time, you won't notice the oversubscribe; you will get more service for less money vs. something that isn't oversubscribed. But, oversubscription needs to be managed carefully.
One thing I've noticed? If your upstream has a 1000mbps line, and sells 1000 unlimited 10Mbps ports, you are almost never going to see contention, even though it's a 10x oversubscribe.
If your upstream has a 1000Mbps port, and sells 10 1000Mbps ports? that's the same 10x oversubscribe. But I /guarantee/ that you will hit contention at least once a week. Probably more often. If people expect 1000Mbps reliable off that, they will be very unhappy. (Of course, if you setup your QoS properly, and tell the customers that it's 100Mbps CIR and up to a gigabit of best-effort burst? it can work out just fine. But nobody is going to get a reliable full gigabit out of that deal.)
90% of your users are using like 10% of your resources. But you've always got a few who are running torrents (or, in the case of CPU, mining primecoins) In the 1000 10Mbps port situation? it doesn't really matter if you've got a few bittorrent users. In the 10 users who can all completely fill the pipe situation? it matters a /lot/
That's the thing about CPU sharing, though; most of the time you don't put that many guests on one machine, and often you give each guest the ability to use the whole machine (when it's otherwise idle) - so you are in the situation of selling 10 1000Mbps links when you only have 1 1000mbps uplink, which ends in tears if anyone actually expects a reliable 1000Mbps uplink. (Now, if everyone understands that it's actually 100Mbps CIR that can burst to 1000Mbps, then sure, people can be happy. but you have to be careful with those expectations. With the 1000 10Mbps links on a 1gbps uplink, customers can treat their 10Mbps link as 10Mbps dedicated, and 99% of the time, they will get what they expect.)
I like that example.
Reminds me of account receivable and bad debt.
Better to have 1000 customers that owe you $30 each rather than 10 customers that owe you $3,000 each. I'm not factoring into this example the cost of billing or customer service. Strictly that if you have 1000 customers it's much less aggravating and you don't loose sleep at night worrying about a big customer that doesn't pay a bill.
Of course, burstable uplinks all suffer from the 'best effort' issue... if you go beyond your CIR, well, there's usually headroom, but not always.
I've run a lot of tests in this area, and even presented some of the results at LISA last month. When it comes to I/O, EC2 is considerably less consistent than many others[1] and that's what really hurts their users. Their prices might be only 2x someone else's if you're only looking at the average case, but it's more like 10x if you consider their extreme variability as well.
[1] Side note: this is a hint that Amazon has an unusually high oversubscription ratio. The same practice has been evident in shared web hosting, with the same effect, since that industry was created.
On the other hand, people wanting to set up their own server clusters... VMs... Linux container stuff... that's heartwarming :)
it's real seamless and nice just ends up being expensive after a while if you are trying to do enterprise stuff & their add-on services aren't particularly cheap. The customizability of AWS & private stuff means you have to do a bit of server admin but it is generally cheaper & can give better performance.
EDIT: Oh also note that (my biggest gripe) I see big performance swings / queue-ing issues that aren't really correlated with traffic. Plus they introduce some platform/API changes intermittently that make their admin UIs kinda buggy. Or you have to change your workflow to integrate with their services (though the APIs for this are usually ok). I don't know, I feel like they make a lot of money draining people on threads since the performance of RoR overall is kinda questionable, along with the performance of their platform. I'm going to try with Play framework soon & see if there are less of these issues.
The gain might not be much however, given how powerful today's CPUs are.
Now we have a simple nodejs daemon which spawns and destroys VMs through libvirt.
Plus, Amazon has servers here in Brazil, so that a plus to me.