Low Cost EC2 Instances With Burstable Performance
aws.amazon.com
aws.amazon.com
ubuntu@aws:~$ dd bs=1M count=1024 if=/dev/zero of=test conv=fdatasync
1024+0 records in
1024+0 records out
1073741824 bytes (1.1 GB) copied, 23.1199 s, 46.4 MB/s
ubuntu@aws:~$ sudo hdparm -tT /dev/disk/by-label/cloudimg-rootfs
/dev/disk/by-label/cloudimg-rootfs:
Timing cached reads: 23292 MB in 1.99 seconds = 11704.62 MB/sec
Timing buffered disk reads: 232 MB in 3.02 seconds = 76.70 MB/secRelatively poor disk performance is somewhat expected. I'm not sure how fair it is to compare it to instance volumes on other platforms, given the significantly reduced flexibility that brings with it.
I think there would be a huge market for a performance-oriented VPS provider who could provide each node with its own, dedicated hard drives/SSDs. All major virtualisation tech (KVM/Xen) already supports raw disk mode.
Obviously, space and SATA ports inside servers are an expensive commodity with off-the-shelf hardware, so this project would require at least some custom hardware to offer competitive pricing. I think the tiny mPCIE SSDs sometimes found in laptops would be a good area to explore.
The thing is, most of the problem with sharing a disk simply goes away once you go SSD.
Spinning disk has the characteristic that sequential access is pretty good... and random access is terrible.
If you have two processes both streaming sequentially off the same disk at the same time, if your scheduling algorithm doesn't allow for terrible latency, the disk is going to see random access. This is why spinning disk is so terrible for multi-user setups.
SSD doesn't have that problem. SSD has problems with writes that are not cell-aligned, but this is functionally similar to the problems raid5 has, and much like raid5, it's not a problem when it comes to reads.
Anyone know what recommends EC2 over Digital Ocean, Vultr, Linode, etc.? Are they more reliable? Enterprise features? Network bandwidth? Cause right now they look hugely overpriced.
I've hosted on Digital Ocean and Vultr for some time and my uptime is great on both. I run constant ping testing and I do see little glitches from time to time between data centers, but that could be network weather on the global backbone. (I have a geo-distributed architecture so there's stuff running at five different locations.)
I'm leaning towards becoming an EC2 apologist on here, but just running quick benchmarks on a t2.micro versus both the $5 and $10 Droplets.
sysbench --test=cpu --cpu-max-prime=40000 run
$5 Droplet ("2.0Ghz", bogomips 4000)- 99.4981s
$10 Droplet ("2.4Ghz", bogomips 4800) - 88.3740s
(I can't find any actual documentation detailing the $10 option being faster, so perhaps this is just random luck on instantiation)
t2.micro ("2.5Ghz", bogomips 5000) - 69.5248s
Now of course the t2.micro won't let you run that around the clock, which for many workloads is entirely fine: as a standard blog host and the like, or the overwhelming majority of server implementations, bursty CPU is exactly what most natural workloads look like.
Add comparisons of the CPUINFO for each-
Both droplets (identical cpuinfo flags) -
flags : fpu de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pse36 clflush mmx fxsr sse sse2 syscall nx lm rep_good nopl pni vmx cx16 popcnt hypervisor lahf_lm
Amazon t2.micro-
flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx rdtscp lm constant_tsc rep_good nopl xtopology eagerfpu pni pclmulqdq ssse3 cx16 pcid sse4_1 sse4_2 x2apic popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm xsaveopt fsgsbase smep erms
Of particular relevance is that the Amazon instance (E5-2670) exposes SSE4 and AVX to your VM, which for many workloads could dramatically increase its advantage.
I guess the whole point of this is that the vague CPU terminology that the various cloud vendors use is seldom really comparable. However to your core question, Amazon becomes a value proposition when you are using all of the parts -- S3, load balances, elastic IPs, shared volumes, availability zones, security zones, VPCs, private networks...it is all multipliers to the value of the platform.
Amazon used to promote their instances via the somewhat comparable ECU metrics. Now, however, unless I'm missing something, you need to try to determine by narrative, because 1 vCPU is very much not equal to 1 vCPU on other instance types.
A) Already have a lot of other infrastructure on EC2.
B) Want to use the other services offered by AWS.
AWS is a collection of services and APIs that you can compose to build big, complex, scalable things. If you just need a VM host, AWS should be the last place you look.
On the flipside, if you have an application that could benefit from outsourcing some of your infrastructure, AWS could save you time and money. For example, we have AWS managing our DB server (RDS), Load Balancing (ELB), Redis/Memcache servers (ElastiCache), static media storage (S3), DNS (Route53), and some of our CDN capacity (CloudFront). We can manage all of these through the same API client (boto), and the inter-service latency is typically low.
I run a service that needs to run about 72 hours worth of processing each day, and it all needs to happen during a 3 hour window. That's a natural fit for spinning up a couple dozen instances then killing them when they finish.
I'd love to see a comparison of what would happen if I kept the same amount of compute power on standby 24/7 using this new instance type.
1. Laziness. Which I don't necessarily mean in a pejorative sense. Maybe someone just doesn't have time, yet, to learn/configure/maintain spinning up an instance for limited times.
2. Single instance. To spin up an instance, you need another computer. If you want that "manager" computer to be an instance at EC2, too, now you need two instances. With this approach, you can set up just one instance and get much of the same economic benefit.
EDIT: Also...
3. Predictable cost. If your manual spun-up instance turns out to need to run for 4 hours instead of 2, you get a bigger bill. With the t2 instances, you'll get a slower compute (if you run out of "credits") but not a bigger bill.
Again, this probably appeals most to small/new customers?
I think you can do this with cloudformation, having it respond to the size of a work queue, however:
> Maybe someone just doesn't have time, yet, to learn/configure/maintain spinning up an instance for limited times.
This is why I can't answer the question above for certain, I got about that far in documentation and went off to find a simpler solution (for me, tutum: https://www.tutum.co/ )
24 * $0.21 = $5.04 per hour
You could burn those for two hours, almost, to match the lowest $9.50 per month cost of what they're talking about in the blog.
The c3 approach would give you 96 vCPUs during that time. The t2 micro for $9.36 or whatever per month, gives you one vCPU. I'd have to strongly favor spinning up 24 to 48 instances of the c3 large and clocking the job in one to three hours if possible.
24 * $0.032 = $0.78 per hour
I run all my CI infrastructure from spots for dirty cheap. Sure, it could all be yanked out from under me, but it's been running non-stop for over a year now. Plus, it keeps you from making "special snowflake" instances that you shouldn't. A minute and Puppet/Chef have got you a splendid new instance. ;)
With burstable instances, you accumulate 6 CPU credits every hour so you can run at 100% load for an hour, once per 10 hours for t2.medium (once every 13.3 hours for t2.medium; once every 15 hours for t2.micro)
It would be nice to have a credit window greater than 24 hours though.
EDIT: ColinCera pointed out the math is incorrect. Updated and removed erroneous conclusion.
DO was competitive with EC2 on price but not on features (and certainly not on security), now with the price advantage gone...
EDIT: corrected calculation
The price advantage is definitely not gone.
While 1 EC2 instance may not use more then 1 GB (which is a very low quota unless your CDNing everything), if you have a couple of instances your almost certainly going over that.
It would be better to do real speed tests of each service to determine average "CPU" speed. I'm sure both services are constantly optimizing for both shared hardware usage and speed, so the stats would have to be updated regularly.
To understand your billing, you need to understand what you're consuming, which you always should. These credits add a little wrinkle, but also make the service cheaper and more deterministic. If you have credits, you'll get the CPU you bought with them.
And even then, one still needs to factor in the 'other' costs like I/O or IOPs, disk (persistent/EBS), IPs, internet and inter-region data transfer… before you understand the real cost.
And then you need to compare to other instance types (which soon will cover the full alphabet -- c, cg, cr, g, h, i, m, r, t… ) and then other providers.
You still have several unresolved issue -
1. Are your assumption on usage (cpu, I/O, internet etc) correct? Will they change? 2. How do I compare performance across providers for a given VM specification. 3. Can I get support when I need it?
And I am sure there are others
It certainly means there is room for other players who just make it simple, whether they are infrastructure folk (like DO/Linode etc) or platform plays that make the pricing understandable by the audience they are trying to target (like Heroku/Ninefold)
Just an observation. I'm not criticizing either way of doing things; obviously, lowering prices straight out is better for the customer, and keeping revenue stable while just upgrading hardware is better for the provider. Last time I lowered prices, I lowered prices directly, and just took the revenue hit. I'm planning my next upgrade now, and instead of lowering prices, I plan on giving everyone more ram/disk/ssd, while holding prices steady.
It is something I've thought about... the problem is that I'm going to have to go down by more than half, and it's way easier to lease enough hardware to more than double everyone's allocations than it is to double my customer base to make up for the lost revenue.
It reminds me of something I learned while working for Comcast years ago - never lower prices, just keep adding "value".
Yes, exactly. I'm saying that is the standard way to do it in the VPS market, in part because until D.O. most of us were self-funding, and it's way easier to pay for double the compute resources than to deal with a 50% cut in revenue.
In the "cloud" market where amazon is, the standard way to do it is to directly lower prices.
Huh. In the VPS market, from what I've seen, the rule is "treat your existing customers as well as your new customers"
while, say, the co-location market is like the real-estate market. "Subsidize your new customers, and if they are still alive when the lease is up, take profits in the form of much higher renewal rent."
I guess what you describe with pre-pays is sort of inbetween. There's a difference in most minds, I think, between raising a price and just not lowering it when you perhaps could be expected to. Most people new to the real-estate market feel pretty bent out of shape when they find out that they have to pay significantly more in rent to renew their existing contract than they will pay if they move.
I do observe that there seems to be a price floor phenomena; for any customer, any price below $x is largely equivalent; they will go for the best thing they can get for $x, so providing a better product helps, but lowering the price below $x doesn't change the equation for that customer. Of course, $x is different for each person, so lowering your price does get you customers who had a lower value for $x.
I've already lost most of the customers that had a value for $x that was greater than what they were paying me at this point; I'm not losing customers nearly as quickly as I predicted. Right now, if I screw something up, of course, I lose the effected customers; I mean, it's really dramatic. You always lose some customers when you screw something up, but I lose way more now than when my prices were lower than the credible competition. But other than that, things have largely stabilized.
(Also, it doesn't help that the wiki is crufty and out of date, and boot menu, last time I rebooted, was still on CentOS 5.)
I think similar principles govern business spending, only $x for them is usually higher. I have a couple of business co-lo customers who have been customers for like half a decade; some of them are still using the hardware they came in on. They could save a lot of money by upgrading hardware (and thus reducing their footprint) or even moving to "the cloud" at this point, because while co-locating modern hardware is cheaper than "The Cloud" - co-locating ancient hardware is not.
The idea is that it works for them, so they aren't going to fuck with it. I'd bet money, though, that if I fucked something up and caused them a serious outage, they'd be gone pretty quick.
>(Also, it doesn't help that the wiki is crufty and out of date, and boot menu, last time I rebooted, was still on CentOS 5.)
I just want to acknowledge those problems. We only have vague plans for the wiki, but we're actively working on upgrading the rescue image and the hypervisor (which, I imagine, is the part of the boot menu you are complaining about.) - these changes will probably not be implemented until our switchover to the new ganeti-based system, but... that should be soonish.
[1] http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ri-modify...
It's not a huge deal, just something to take into account when projecting out costs: the 3-year reserved instances are locking in today's prices until 2017, in return for a discount over today's non-reserved prices. Whether this produces long term gains requires some assumptions about how the market will change over the next 3 years.
And of course Amazon has transparent disclosure for outages and security issues unlike say Linode.
I've always used paravirtual AMI's, as I understood that gets the best performance for a Linux box.
Given that I try to use the same self-baked base AMI's for various purposes (and instance sizes), I would either have to mix and match or switch everything to HVM. However, I have no clue what the practical consequences of that would be.
Can anybody enlighten me?
Yes you'd have to build new AMIs with HVM. It'd be easiest if you had some kind of configuration management so you didn't need as many AMIs baked. When I build machines I use a script to handle the creation and mounting any extra volumes on a machine that I have as "nonstandard". I have only 2 custom AMIs - one for PV and the other for HVM. You'll need to have at least both, because if you wanted to use certain instances (t1.micro, m1.small come to mind) you can only use PV.
[1]: http://www.brendangregg.com/blog/2014-05-07/what-color-is-yo...
The thrashing will increase gradually until user experience is pleasant.
A quick click around dell finds that a mid-range 1U rackmount server (R320) with that much RAM costs $3,135.
So a back-of-the-envelope calculation makes it seem workable, especially for high-RAM low-CPU configurations, which is what this is.
There are other tricks that they might be employing, such as swapping out part of RAM to SSDs behind the scenes, as well as compressing RAM contents. On low-load servers like these, typical usage would imply that RAM would be mostly static.