Ksplice: Automatic Rebootless Kernel Updates (2008) [pdf]
ksplice.com
ksplice.com
On 21 July 2011, Oracle announced they acquired Ksplice, Inc. At the time the company was acquired, Ksplice, Inc. claimed to have over 700 companies using the service to protect over 100,000 servers. While the service had been available for multiple Linux distributions, it was stated at the time of acquisition that "Oracle believes it will be the only enterprise Linux provider that can offer zero downtime updates." More explicitly, "Oracle does not plan to support the use of Ksplice technology with Red Hat Enterprise Linux."[5] Existing legacy customers continue to be supported by Ksplice, but no new customers are being accepted for other platforms.[11]
http://www.oracle.com/us/technologies/linux/ksplice-datashee...
Ksplice Uptrack is available for Oracle Linux, free of charge, for Oracle Linux customers with a Premier support subscription. Additionally, anyone can use Ksplice Uptrack for free on Ubuntu Desktop and Fedora.
Disclaimer: I work there.
EDIT: so, how much does it cost?
Unfortunately, the website hasn't been fully updated since the acquisition, so there is some confusing verbiage. It's on my list (near the top, no less; sorry about the confusion).
But there's also kgraft and kpatch.
"There has been a great deal of interest in live kernel patching (see http://lwn.net/Articles/597407/) over the past few months, with several different approaches proposed, including CRIU+kexec, kGraft, and kpatch, all in addition to ksplice. This microconference will host discussions on required infrastructure (including tracing, checkpoint/restart, kexec, and live patching), along with expositions and comparisons of the various approaches. The purpose, believe it or not, is to work towards a common implementation that everyone can live with. It should be a spirited discussion! For more details, please see http://wiki.linuxplumbersconf.org/2014:live_kernel_patching"
(1) the update could not be applied to the running instance, so you have to restart to put the update into service, and
(2) to demonstrate that the update did not break the ability of the thing to start.
This address #1. It does not address #2.
It is extremely annoying if something has run for a year with a couple dozen live updates applied without restarting and then you have to restart for some reason (power failure, perhaps), and then you discover that somewhere in the last year it lost the ability to start. I'd much rather discover that an update has introduced a problem with startup immediately after applying the update.
Restarting things should be a regular part of maintenance.
I completely agree with your last comment, though: systems suffer from weird bit-rot diseases. If you never reboot, you may find things challenging when you finally have to.
Disclosure: I work on the Ksplice team.
Step 1. Boot your version 1 kernel.
time passes, and version 2 is released
Step 2. Apply Ksplice updates to bring your kernel up to version 2.
Your system reboots (let's say due to a power failure):
Step 1. Boot your version 1 kernel. This boots because it's the same kernel you booted a year ago.
Step 2. (Before bringing up the network) Apply ksplice updates to bring your kernel to version 2. This works because you booted the same version 1 kernel, and the updates haven't changed either.
When he rebooted his machine (after almost a year) and it failed to come up, I got to mock him back. "So can you remember what changes you made... in the last year".
On the other hand, even though I've never used it, ksplice must be a nifty way to apply urgent updates between regularly scheduled reboots.
Look at KPatch (included in RHEL7 as tech preview, will probably mature in subsequent minor version updates): http://rhelblog.redhat.com/2014/02/26/kpatch/
These companies are always hindered by the downtime incurred when one of these non-redundant devices needs restarting. In a couple of previous positions over the years I would have loved the ability to apply patches with no downtime.
There are environments where the time to failover to redundant infrastructure is critical (HVT, particle physics etc.) but the more common use case is probably going to be the companies that have no failover capability for certain parts of their infrastructure.
Any instance where you're running untrusted binaries on your server you'd want to update ASAP and not worry about getting pwned while waiting for scheduled downtime.
Before Oracle closed it to new customers, Ksplice was fairly cheap, ~$3-4/mo/server.
To me this is the first step, the second is doing the same with system libraries. If zlib gets updated what do I have to restart to ensure running applications are using the right version? Right now the answer is a reboot.