EC2 Users Should be Cautious When Booting Ubuntu 10.04 AMIs
silverline.librato.com
silverline.librato.com
My personal hunch is that there are probably startups and developers that dumped Cassandra due to "stability issues," which were in reality symptoms of this bug. There's obviously no way of confirming or denying this, so I'll reiterate that it's just a hunch.
Due its short support cycle, Maverick isn't suited for servers.
When you have Linux on your server, you own it. You can do whatever you want. The distro is the beginning of your power, not the end of it. If you're running a Linux server and you currently have the attitude that you are boxed in by your distro I recommend that you immediately dig into the relevant packaging system and learn enough to put your own patch on top of any existing software package, and recreate the package in the relevant manner (new RPMs, new .deb, whatever).
(Yes, there's a cost/benefit tradeoff to each such patch you have to carry on, but there is still economic value merely in having the option.)
The only time I recompile a kernel is when I'm working on kernel code. If UNIX distributions are doing their jobs, a sysadmin should never have to touch it.
Alas.
That's nonsense. Most EC2 AMIs are linked to the amazon AKIs which are unrelated to whatever distro the AMI contains. Most of my debian instances run on a kernel tagged "fc8xen".
The ability to chainload a self-compiled kernel on EC2 is a relatively recent invention (mid-2010) and I have yet to see a good reason to do that for linux.
The article does unfortunately not mention which AKI(s) are affected, but it seems likely this bug was introduced because someone figured "newer is better" and went with the latest Ubuntu kernel instead of sticking to a proven amazon AKI.
You're sort of contradicting yourself here. You suggest that the distro you're running is independent of the kernel version you're running. But then you go on to claim that this bug was introduced by someone who was not running the default supported kernel. Are you saying that people should run the supported kernel, and be tied to whatever's supported upstream, or are you saying they should risk building their own? Clearly there are benefits and drawbacks either way.
[1] https://bugs.launchpad.net/ubuntu/+source/linux-ec2/+bug/708...
Running "time-tested" kernels is not really the best advise either in this case. Xen is a fairly new environment, and EC2's implementation has some quirks, so there's a pretty regular stream of bug fixes and other improvements in recent kernels that are often worth picking up. If I went to Canonical with a "time-tested" kernel bug they'd tell me to upgrade before they'd give any real support.
When we talked to Amazon about switching to their AMIs they advised us that it was probably _not_ worth switching, that switching might not fix the problem, and that the AMIs we were running were widely used and supported. They made it clear that they work closely with Canonical and other providers to get high quality AMIs into their ecosystem. Long story short, the people who you admit know the most about the EC2 environment advised us that they weren't necessarily the best option, or at least not the only option, for good AMIs (sort of like how hardware manufacturers aren't the best option for an operating system).
So the answers aren't really cut-and-dry here. Every time Amazon changes their dom0 there's a chance your "time-tested" kernel will stop working. And just because Amazon runs the infrastructure doesn't mean they're the best choice for a Linux distribution.
http://uec-images.ubuntu.com/query/lucid/server/released.txt
Looks like their judgement didn't work out so well this time. I'd be wary of running the latest untested kernel for no reason other than "because we can".
We asked the Amazon kernel team if we should try switching to one of their kernels/distros, and they said "No, just upgrade to Maverick and the accompanying kernel." It's been pointed out that Maverick has its own set of Xen bugs. I guess Amazon doesn't know everything.
The horse you're getting on about using the "proven" Amazon kernels is a bit high. Turns out this whole virtualization thing is somewhat new, and the kinks are still being worked out. Old kernel builds don't work particularly well because a lot of their assumptions are broken by virtualization; new kernels are what they are - new.
(Edit, forgot initially): Finally, we ran 10.04 - the Long Term Support release of Ubuntu from a year ago. There was no "because we can."
Frankly, I'm a bit amazed at your disdain for people sharing their findings from practical experience running into these issues in high-load production environments.
Neither is infallible. But Amazon probably knows the intricacies of their platform better than Canonical. And they likely run some of their own stuff on these kernels for a while before releasing them to the public.
Old kernel builds don't work particularly well
Don't work as in what? This is the first time I hear about a kernel problem on EC2.
disdain
I don't see where I voiced disdain. I merely responded to the guy who claimed your EC2 kernel is linked to the distro you run. That's simply not true.
The guy who claimed EC2 kernels are linked to the distro you run was simply claiming that, unless you want to go it on your own, you're tied to the kernel provided by a supported AMI. As you've suggested multiple times, there are benefits to running an environment that is supported and that other people have operational experience with. Honestly, I'm not even sure what you're arguing anymore... seems like you're just being antagonistic.
[1] https://bugs.launchpad.net/ubuntu/+source/linux-ec2/ [2] https://bugs.launchpad.net/ubuntu/+source/linux/+bugs?field....
There are plenty AMIs based on stable AKIs out there. Moreover if you manage a "very large EC2 infrastructure" then you don't rely on 3rd party AMIs, do you?
Finally, your links point to... Ubuntu bugs. If I missed one that was tracked back to an amazon AKI then a deeplink would be appreciated.
I'm in the process of upgrading all of our instances to Natty from 10.04 or younger. It's actually weird that this issue didn't get any attention whatsoever.
It is a pretty minimalist AMI, which is what we wanted. One caveat: it is in beta.
(And I say that as a longtime RH fanboi)
Ubuntu shouldn't be used for web servers to begin with. You should be using a distribution with a long and thorough stable release cycle with minimal packages, such as Debian.
Server, however, is perfectly appropriate, especially the LTS release (which is supported for five years). Ubuntu LTS will probably actually be supported longer than (for example) Debian Stable.
I've personally found Ubuntu more bare bones out of the box than CentOS.
Debian's primary concern with its release cycle is stability. Other distributions like Ubuntu or CentOS trade a little stability for newer software.
My method: Use Debian but when you must have newer versions just add an unofficial repository to your apt sources (assuming you're prepared to deal with the complexity and inconsistency this might introduce).
I believe that the debian stable release cycle is currently at about 2 years, w/ a very short support cycle afterwards. The LTS cycle is 2 years with a 3 years of additional support afterwards.