Massive RedHat Perl performance issue
blog.vipul.net
blog.vipul.net
This is why. Eventually, something dumb will happen inside your tools, and you'll need to figure it out. Blind faith in any of the software you rely on is bad, you need to know how it works. And to first approximation, all the software you rely on is written in C.
Apple totally fucked sqlite for awhile (it may still be, I compile from source now) by doing a full filesystem flush (not fsync) on every commit:
http://adiumx.com/pipermail/adium-devl_adiumx.com/2008-April...
There was a (crazy-talk) rationale for it, though: fsync wasn't thought to be "reliable" enough, so the order-of-magnitude slowdown was for our own good. Doubt bless() really "needed" the slowdown for RHEL.
Fsync() causes all modified data and attributes of fildes to be moved to
a permanent storage device. This normally results in all in-core modi-
fied copies of buffers for the associated file to be written to a disk.
Note that while fsync() will flush all data from the host to the drive
(i.e. the "permanent storage device"), the drive itself may not physi-
cally write the data to the platters for quite some time and it may be
written in an out-of-order sequence.
Specifically, if the drive loses power or the OS crashes, the application
may find that only some or none of their data was written. The disk
drive may also re-order the data so that later writes may be present,
while earlier writes are not.
This is not a theoretical edge case. This scenario is easily reproduced
with real world workloads and drive power failures.
It's still Apple's fault on some level (after all, they control everything from the fsync implementation to the hard drives they choose to ship in Apple hardware) but from the perspective of the guy configuring sqlite, the full filesystem sync makes sense.Way to go, RedHat!
When you need to replicate your environment, you now have to build all of the custom bits exactly as they are on the production system, rather than simply running "yum install perl foo bar baz". Depending on the length of your dependency chain, and the dependency chain of all of those components, this could be incredibly time consuming, even if you don't make any mistakes in the building process. Building a binary tarball of all the stuff you need is an option, but then compatibility issues with existing system libs and such are bound to happen, and that's pretty ugly from a paths and upgrades perspective.
You also make your environment less standard. A new hire is going to have to learn not only your application, but also all the crazy town details about your particular and very specific deployment (and setup their own copy of it on their own system). If everything except your app comes from OS-standard packages, you can expect someone familiar with RHEL/CentOS or Debian/Ubuntu or whatever OS you use to know where most things are right off the bat.
You'll probably do more things wrong with your build than the OS vendor did with theirs. In my business, I see a lot of custom PHP builds, for example, and almost every single one of them is broken in more than minor ways (and we end up hearing about it, and trying to figure out what they did wrong in their build). Your OS vendor version has a lot of people banging on their builds and reporting bugs. I'd pretty much always bet that their build is better than yours from a reliability perspective.
It makes it harder to replicate your deployment, if something catastrophic happens to your production box. Packages are more resilient to library changes and such than a big ball of crud tarball of your binary builds. And you won't want to spend several hours rebuilding on the new target machine while you're offline. A complete system backup could be restored...I dunno if you've ever done that on a remote system before, I assure you it is non-trivial and stressful.
What I would instead recommend is to find out which components you need custom (I'm not denying that sometimes you really do need, for example, perl 5.10 and the OS has 5.8.8--it happens, and that's fine), and build new packages int he native format and dump them into a yum or apt repository. It takes an extra day or two, if you don't already know how to do it, but it'll save you many many times that amount of time in the future--and those hours in the future might be far more stressful hours than while you're first setting things up. Rebuilding a package from SRPM or a deb source bundle is usually pretty easy...bumping revisions in dramatic ways might not be trivial, but recompiling with specific options is no problem at all. And, one can usually find a source package of the latest and greatest in the devel branch of the OS, which makes even major revision bumps easy (though, because it is the devel branch, you're probably giving up some maturity in the package...far fewer testers on the devel versions).
All vendors suck to different degrees. Nothing ever works perfectly. Open source gives you the ability to do something if the suckiness affects you, though. With closed source, you're stuck until the vendor gets around to your bug. With larger vendors, this may take forever.
While this is a pretty serious problem in a pretty darned popular and important package (and I'm a Perl developer with well over half of our customers running RHEL or CentOS 5--so I'm a little more than distressed by it, since most of our customers may be seeing our software run slower than it should be), it is not apparent to me that there is a great solution to this problem--upstream has fixed it in the 5.9 and 5.10 branches, but not in 5.8. So, the only real solution is a binary incompatible change. RHEL guarantees no changes that effect binary compatibility across the lifecycle of a RHEL release (unless absolutely necessary for security or stability--and even then, I've seen them opt not to change something, because the stability issue only effected a small number of users and the binary incompatibility would have effected everyone).
It's a hard problem to solve--the implication that Red Hat are ignoring it isn't really fair.
That said, some of the folks managing tickets in the RH bug tracker are assholes. I've had very few positive experiences when filing bugs about RHEL (they did finally deal with my two tickets about how much up2date sucked, by deprecating up2date and replacing it with something awesome, so I'm feeling pretty good).