AWS: the good, the bad and the ugly
blog.awe.sm
blog.awe.sm
We're a big user of AWS (well, relative, but we run about $10K/month in costs through AWS), so I'd like to supplement this outstanding blog post:
* I cannot emphasize enough how awesome Amazon's cost cuts are. It is really nice to wake up in the morning and see that 40% of your costs are now going to drop 20% next month going forward. (Like: Recent S3 cost cuts). In Louisiana, we call this Lagniappe (A little something extra.) We don't plan for it, nor budget for it, so it is a nice surprise every time it happens.
* We've also completely abandoned EBS in favor of ephemeral storage except in two places: some NFS and MySQL slaves that function as snapshot/backup hosts only.
* If the data you are storing isn't super critical, consider Amazon S3's reduced redundancy storage. When you approach the 30-50TB level, it makes a difference in costs.
* RDS is still just a dream for us, since we still don't have a comfort level with performance.
* Elasticache has definitely been a winner and allowed us to replace our dedicated memcache instances.
* We're doing some initial testing with Route 53 (Amazon's DNS services) and so far so good, with great flexibility and a nice API into DNS.
* We're scared to death of AWS SNS - we currently use SendGrid and a long trusted existing server for email delivery. Twillo will is our first choice for an upcoming SMS alerting project.
* If you are doing anything with streaming or high bandwidth work, AWS bandwidth is VERY expensive. We've opted to go with unmetered ports on a cluster of bare metal boxes with 1000TB.com. That easily saves us $1000's a month in bandwidth costs, and if we need overflow or have an outage there we can spin up temporary instances on AWS to provide short term coverage.
Can you elaborate on this a bit? Why are you scared of SNS? Data loss / latency etc?
FYI http://www.1000tb.com/ is the website for a landscaping company, not about bandwidth...
We switched to Elasticache recently as a test, and found that it works fine but don't see any compelling reason not to just manage our own memcached servers.
Edit: Also, what scares you about SNS? We use it for our oncall alerting and have seen no reason to be concerned about reliability - but I'd love to know if we should be!
Disclaimer: I work at SendHub, but I've done it both ways and I definitely prefer ours ;)
Intel E3-12303.2 GHz8GB2 x 1TB100 TB$201.15 1Gbit Dedicated Port
So, for $200/month, they'll give me 100 Terabytes/bandwidth on this server.
250 megabit/second at 30 days in terabytes = 81 Terabytes. A decently peered/connected Pipe costs, at this volume, around $5-$7/megabit @95th, or $1250.month.
So either:
A) Their connectivity isn't hot.
B) If you actually use that 100 Terabytes per server,
they start to curtail/rate limit/traffic shape you.
C) Some other option I haven't considered? Maybe they
just rely on the vast majority of their customers not
using that bandwidth, and subsidizing (significantly)
those that do?
blantonl, appreciate any insight you can provide - sounds like you've got great control over your environment. From the sounds of it, 100tb.com provides bandwidth @$0.33/megabit.http://gigaom.com/cloud/simplecdn-takedown-a-cautionary-tale...
I don't want to try to guess why they did that, but it's something that should never be forgotten.
I'd be very interested though, in what people have been able to sustain per server at 100TB without getting dropped.
My guess is that it's somewhere around 32 Terabytes/month per server before 100TB gets a little angsty (100 mbits/second sustained), particularly if you have more than a dozen or so servers with them (filling out a GigE sustained)
It's pretty much impossible to do this in the weekend without even the passwords of the customers. Imagine receiving this email when you wake up (probably some time after it was sent) and having to get new servers immediately somewhere, emailing/phoning customers for their passwords and then also having to transfer potential TB's per server in the small remaining time frame. After posting on WHT (meaning after a PR disaster), 100TB did extend the deadline.
http://www.webhostingtalk.com/showthread.php?t=1218922
I was considering them for a high-bandwidth project, but after reading these cases I changed my mind.
I suspect it is C) - the simple reality is that most Web hosting use cases don't consume bandwidth at rates that saturate an unmetered port, other than those that violate the TOS from 100TB.com (i.e. running a CDN reseller)
We've never, ever had a single issue with 100TB.com. The servers we have with them are solely in place to provide MP3 streaming capabilities, and we certainly do get our money's worth.
yes, the transit is not super. i wouldn't run voice on it, but i'd surely run bulk transfers off of it.
of course they're counting on underuse, that's why it's sold at port level instead of usage
they may not be shaping, but guaranteed you'll see primetime saturation pushing you (way) below port speed
this kind of bw costs them about $0.50/mbit when bought via 10g xc
However - most people don't use anywhere near this amount - really, they are selling you a server connected to a gigabit switch.
As soon as over-subscription kicks in (it does immediately, since you don't start using 250Mbps the day you have your server installed) their bandwidth costs go to a fraction of what you are paying monthly.
In my experience, with large data sets (I'd say > 100 GB), you want to run your own database instances, if only so you have control over the small fiddly bits you need control over.
RDS: 1) Uses drbd for master failover (this introduces network overhead on disk writes), 2) Only gives you access to super commands through stored procedures, and 3) doesn't let you tune InnoDB to your dataset or performance requirements.
The point in the post that the AWS management console runs on EBS is fishy...
Don't get me wrong - the idea is great and it's great for really low end workloads where the gotchas don't matter or really big systems where you'd be managing against most of those issues anyway.
But in the middle of the two there is a dead space that I bet a lot of shops are stuck in - double digit instance numbers running a workload that a 2-4 servers could handle with plenty of leg room.
If you're deciding on a cloud provider in 2012 I think it makes a lot of sense to shop around. There are lots of people doing the on demand api deployment thing now with different trade offs. I like joyent a lot (local reliable io) or providers with cloud and a colo area even if it's exorbitant - as paying five hundred dollars a month for instnaces one ssd could replace sucks.
As parent says: EC2 does have its place, but for deployments in the 10-20 hosts range (or an app with "special needs") you tend to pay through the nose. Keep in mind that the amount that you pay for a fraction of a server on EC2 will rent you the entire box elsewhere.
Assuming that I/O is reliable and fast-ish was a bad OS design decision in the 1970s, as the Unix guys themelves soon realised, though it was an understandable mistake at the time. Continuing to develop and use OSes with an I/O interface that assumes I/O is reliable and fast-ish, in 2012 - it's pretty bad, isn't it?
With network filesystems (eg. NFS) you can choose to return an I/O error to the application when you hit a timeout or a network error (the -o intr mount option). This is rarely used since applications aren't used to dealing with them. So can't really blame the OS here either.
That said, I'm not sure I agree with the idea that we got away with a lack of error handling because disks had consistent performance. Magnetic disks have always had incredibly inconsistent random IO performance, and even inconsistent performance between different parts of the platter(s). And in the spirit of co-evolution, we engineered around it: OS disk caches are critical to decent performance on HDDs.
I think it's not that we found disks to be consistent, as much as our solution to their suboptimal behaviors was caching, rather than error reporting/timeouts. I believe this is because caching was the most transparent approach; a good cache makes a variable-speed disk look just like an ideal disk, so an application can be written assuming the disk is perfect.
Interestingly, as distributed systems have evolved, we've ended up having to engineer the error-handling constructs that might have been used for block devices; we see them in network filesystems (as mentioned) as well as most other network services. Applications have been designed to deal with errors. We just haven't propagated those constructs down to the disk devices in the recent past.
From a non-realtime app POV disk seek performance is pretty consistent: you get 5-25ms seek cies that center around 10ms. Especially In contrast to network backed where you get to contend with hiccups and contention with other users.
OS disk caching came about for a different reason.
Instead we rely on instances that use an instance-store root device. During the EBS outage, our instance store servers did not have any issues, while our EBS-backed servers really struggled throughout the day, with crazy high loads.
http://devblog.pipelinedeals.com/pipelinedeals-dev-blog/2012...
Also, do you backup your databases anywhere other than on other ephemeral instances?
We have 2 separate chef recipes for our DBs:
One is for a full-time slave. This recipe will set up the db to use the EBS volume.
The other is for a slave that will be promoted to a master. In this case, we do a little extra legwork to do a bit-by-bit copy of a recent EBS snapshot, onto the ephemeral disk.
Sounds like I have more blog posts to write!
Bluntly it seems like you must have spent some dev or ops time learning all this and migrating away from EBS etc. even if you didn't hire someone.
Frankly, if your bills are greater than $3K per month with AWS I question whether you are truly saving anything.
(I figured that midsized instances vs. dedicated servers are about 5:1 in terms of performance)
Our ops guy is much better than me, but his time is better spent working on higher-stack stuff like deployment automation, monitoring and efficiency tuning than on re-inventing a virtualization stack to save a few thousand dollars every month.
If we were bigger, it would be more worth the time and money spent. But without doing the math, we would have to be quite a lot bigger, I think.
I've got the Nobel Committee for Physics on line 2 should you accomplish that trick for DCs more than 30 km (18.6 mi) distant.
Edit: Nevermind (misplaced decimal).
Of course light doesn't travel that fast through glass and there is latency at the hardware on either end. ;)
There are real costs no matter which way you go.
PS, would be very surprised if you had even 45ms latency between AWS-east in Virginia and any of their facilities on the west coast...
And the < 10ms latency I'm talking about is between zones within us-east; latency to the west coast is a lot worse, but we only have emergency failover capacity in west.
The ability to whip out your credit card, and fire up new servers at AWS in minutes, becomes very attractive in that scenario.
And you can find essays of people who run their own rack who swear the same thing about moving to AWS.
I worked at an all-AWS bay area startup that ran a $100k/mo Amazon bill, and now I work at a bay area startup that runs all its own hardware.
It's about tradeoffs. It's not just about cost. (Though there are many cases where Amazon is cheaper and it has little to do with your monthly spend, and certainly not an arbitrary threshold like $3k/mo.)
Writing this at the close of my employer's Open Enrollment period, I'd say this: Comparing AWS with running your own hardware is like trying to compare two health insurance plans. Each brings its own seemingly impossible tradeoffs. All you can do is bring your experience and judgment to identify core issues, make the best call you can, and suppress the urge to dream of the road not taken when you're knee deep in whatever hell you're experiencing that would be a non-issue if only you had went the other way.
Though if you loathe the AWS "fanboyism" so much then maybe it's the company you're keeping: Few currencies in a startup are as valuable as flexibility. The younger the startup, the truer that is.
The comment around Ubuntu is interesting and I wish there was more detail there.
We use mdadm to run RAID across multiple EBSes. mdadm is great, but has a kink that it will boot to a recovery console if there the volume is "degraded" (i.e. any failure). This is even if the volume is still completely viable due to redundancy. This is obviously very bad, as you've got no way of accessing the console. It's an unfortunate way to completely hose an instance.
It's an easy one to miss, as you rarely test a boot process with a degraded volume. When it happens though - hurts a lot.
(If you'd like to check on this, make sure you have "BOOT_DEGRADED=yes" in /etc/initramfs-tools/conf.d/mdadm).
In terms of AZs, were are distributed roughly evenly across all AZs in east-1.
So what do you use now for your persistent storage? This might be the most interesting part.
PS great article
Except for t1 and m3, which have zero ephemeral storage.
Just a tongue-in-cheek comment: Those instances also cost over $2000 USD per month. For the same money you could buy a new Dell Server with 2-4 SSDs, every month. ;)
In this case, I had about 8 million output files from a Hadoop job that I needed to process with MySQL- it was perfect!
http://docs.amazonwebservices.com/AWSEC2/latest/UserGuide/In...
However, what do you do if there's a massive outage that affects most or all of your instances simultaneously? Diversification that fails when you need it the most isn't very helpful.
Of course, east-1 has fallen off the map entirely on at least one occasion -- for that reason we keep a "seed" set of warm databases in us-west; if east-1 were to have an extended outage we have a disaster recovery plan that involves transferring DNS to IPs in west and spinning up new app instances there.
However, our primary strategy for uptime is redundancy -- every db has at least one slave, and we are spread across multiple availability zones.
Stuff not in a database goes on S3.
And for that very uncommon requirements you have to pay per Kb and per hour rates for having no service or guarantee at all.
And no, you still need a sysadmin who understand how AWS works and what to do when AWS says "Oops, your "_____" isn't available".)
>> Both EBS boot and instance-store AMI ids are listed, but I recommend you start with EBS boot AMIs.
Why two opposite recommendations from alestic.com [authority on AWS] and practitioners?
Not a flame - I am planning my AWS deploy strategy and need to make a decision between these two approaches.
Thus, most of the time, performance issues with EBS bite you when you have application data on an EBS volume. If you're constantly hitting the volume to serve data to customers, you'll be seeing the hiccups in EBS performance and passing them right on to customers. Clearly this is a suboptimal experience, and can lead to other failure modes; a slow/stuck bunch of ops can get an application to completely stall.
It's important to note here that EBS has no error reporting system or timeouts; common operating systems aren't very good at handling disk IO errors, so EBS will never produce them, even when its having issues. This lack of transparency can make handling/working around stuck EBS operations nigh impossible.
So, having experienced these pains, people generally say "Don't constantly use EBS if you care about your application performance/stability," and they aren't wrong.
That said, if you use EBS as a boot volume only, you are likely not hitting that volume enough to feel the pain of slow-IOs; once booted you won't touch it much. At the same time, using EBS as a boot volume has all kinds of conveniences: the ability to start and stop the instance without loosing data, snapshots and AMI creation, persistent machine-specific configuration, etc.
I hope that's useful. EBS is not a perfect service, but it definitely has its uses. EBS as a boot volume has made using AWS quite a bit friendlier.
Note: I worked on the EBS team at AWS a few years ago. I currently use EBS in my startup's infrastructure. If you have further questions, feel free to reach out to me.
Our problem with EBS as a boot device was that occasional system logs, etc. would go it, and when EBS was down the inability to reach the disk (the block-store abstraction problem) would lock up the whole OS.
I'm yet to play with AWS on any significant level, but this is the kind of thing to bookmark.
Further to that, are there any recommendations for books/sites/entries that discuss more best practices?
FWIW, we (at DuckDuckGo) ended up in much the same place: ditched EBS, avoid anything that relies on EBS, and multi-zone and multi-region redundancy (also for latency purposes). For ephemeral storage purposes, we end up mainly using xlarge machines since they have the greatest stability and speed (with 4 drives in RAID-0).
OTOH, if you need equal or more than n servers, rolling your own dedicated servers is better and cheaper.
The only issue is to determine the current value of n.
I've also looked into ephemeral storage, but ultimately I decided to just rent a dedicated machine from elsewhere. Building a B2B site, I'm not as worried about massively scaling on a dime. The project has still used transient EC2 instances for a few odd things, though. It's nice to have that option when you need it!
I assume it runs a standard linux distro, therefor patches , firewalls, dependencies etc are still an issue surely?
Many times what I've seen is simply that position is simply 'filled' by developers whom have ideas about designing and securing scalable systems, but when it comes right down to it... they get hacked within some period of time because it turns out, they don't know these things well enough to take the place of someone whom dedicates their time to it.
AWS's security groups pretty well cover you on the firewalls front (although some folks like to still run software ones on their instances to be doubly sure), but you're otherwise correct.
For a single instance, you're not going to see significant differences. It's running large clusters of instances you're going to see things being better than the average VPS provider, with stuff like Virtual Private Cloud etc.
Does AWS provide easier tools for that stuff? Because you can't really cover it with a general firewall.
No reason you can't run stuff like mod_security, but that's not strictly a firewall, just Apache setup.