[1] http://metadata.ftp-master.debian.org/changelogs/main/l/linu... [2] http://metadata.ftp-master.debian.org/changelogs/main/l/linu...
However, I've often been in situations where I reboot anyhow, because rebooting means I'm 100% confident the old code is gone, whereas if I try to get clever and avoid the restart, I'm significantly less confident. Depending on how hard it is to validate the security bug, that can be a problem.
Plus, for much of the past 20 years for many computers, if you're going to restart all services, adding in the last step for rebooting doesn't add all that significantly to the downtime. Server-class hardware often have things that make that not true (stupid RAID cards), but for everything else you were often only adding, say, 25% for the actual reboot.
Also, boot time bugs are a huge issue. They can creep during the entire time your system is up, and only show up during a reboot. Thus, if your server only has unplanned restarts, you'll only discover those bugs when you have yet another pressing issue to deal with, and also, likely at 3 in the morning on a Sunday.
So, make things better for you, and restart those servers once in a while, when things are quiet.
It's rare to see a kernel-bug that can't be worked around in some way other than patching.
I've seen many windows systems with uptimes of several years that have never required any maintenance.
(I'm not advocating for HVAC/SCADA systems to be running, say, Windows XP Embedded with no updates and default passwords, world-facing, just observing that the preconditions changed.)
What ends up happening is that they'll bolt on some network connectivity for convenience or to take on some new process and not set it up appropriately, or not understand what it means to expose something to the LAN or directly to the internet.
I helped a friend at a municipal utility with something like this when they wanted to provide telemetry to a city-wide operations center. They had a dedicated LAN/WAN for the SCADA stuff, and the only interface was in this case a web browser running over XWindows that had a dashboard and access to some reports. I think they later replaced it with a Windows RDS box with a similar configuration.
Because of the isolation, and professional IT who understood how to isolate the environment, it was advisable to to not be tinkering with updates, as the consequence of failure is risk to health & safety.
Which, incidentally, have been the target of a lot of recent high-profile attacks.[0][1][2][3]
[0] https://en.wikipedia.org/wiki/Duqu#Purpose
[1] https://en.wikipedia.org/wiki/Stuxnet#PLC_infection
[2] http://www.computerworld.com/article/2475789/cybercrime-hack...
[3] http://krebsonsecurity.com/2014/02/target-hackers-broke-in-v...
We do regular security audits from a security firm who goes the extra mile to try and social engineer and gain physical access to all of our sites.
Plus we're talking about things like processing fish in a town of 2,000 people. If I was operating a nuclear reactor, I would surely adapt better security measures.. although against government sponsored attacks using undocumented vulnerabilities, windows update isn't really going to do much.
The Target thing you posted has to do with internet access, which is something that goes against what I was saying. I'm talking about closed, physically secure networks, possibly not even using tcp/ip or ethernet.
Stuxnet-like attacks can go after non-networked equipment, but they're based on exploiting the computer with the programming suite, not the industrial system itself.
My gut feeling is that it is kind of like driving a car without wearing the seat belt. So far, if I had never worn a seat belt, nothing bad would have happened, because I did not have any accidents. But when it happens, one goes through the wind shield, so to speak. Also, some stability/performance issues do not manifest until a machine has been running continuously for months or years.
(What is more disturbing, though, that the very-high-uptime systems (~4 to 8 years) I have seen also appeared to never get backed up, and there didn't seem to be any plans for replacement, or at least spare parts. Which is kind of bad if the machine happens to be responsible for getting production data from your SCADA to your ERP system which in turn orders supplies based on that data.)
[1] https://www.smartmontools.org/browser/trunk/smartmontools/sm...
While it's true that most critical bugs I've run into with the kernel can be workaround in some way, it's at the detriment to some use-case that some users, somewhere rely upon.
A vulnerability I discovered a bit of a year ago allowed local privilege escalation for any user with access to write to an XFS filesystem. The only workarounds were to modify their SELinux policies or switch to another filesystem. I'm not aware of any users that rewrote their SELinux policies for this. It had been fixed in the kernel and fairly quietly fixed in a RedHat security advisory, but I don't think any other distribution did anything at all.
I'd estimate that critical kernel bugs happen at least other month. Worse, discussion of these bugs happens in the public and takes dozens of months to fix. Again, assuming you even know of these vulnerabilities, while there are often workarounds, they're not always practical.
Did you know that via user namespaces, all non-root users on a machine can elevate to a root user? That root user is supposed to be limited, but it's allowed the mount syscall. Numerous vulnerabilities have been discovered as a result of this. The kernel team usually considers them low-impact and they get a low CVSS score, but when using certain applications this can lead to local privilege escalation. The workarounds are to disable user namespaces or disable mount for user namespaces, both of which will break some set of users.
Did you know that any user capable of creating a socket can load kernel modules? For a long time this allowed loading ANY kernel module! The only workarounds were to compile it out of the kernel or monkeypatch the kernel. Only last year was this was finally fixed upstream so only those modules patching a pattern were loading. Yet, it was also discovered that if you used busybox's modprobe, the filter would still allow non-root users to load arbitrary kernel modules from anywhere on the filesystem!
Point being, this is par for the course. Clearly one needs to understand their threat model, but if the model is at all worried about local privilege escalation, update weekly until you find another OS.
So, mount any existing xfs file systems read-only, and move rw systems to ext3? I'm not saying it would make sense - but sounds like a prime example of something for which there was a work around...
(I'll concede that for those that need(ed) xfs, there'd probably not be many alternatives at the time. Possibly JFS?)
Which is of course, a really bad idea for the general case.
Although it is actually kind of a requirement in some industrial environments, where certifications are involved - once the thing is certified, any change, hardware or software, requires a re-certification, which apparently is expensive and tedious. Which is how many industrial plants, too, end up running on ancient computers, at least by todays's standards.
Where we have seen a lot of innovation is in the engines -- fuel economy and noise regulations have pushed GE, RR and P&W to up their game substantially.
https://www.youtube.com/watch?v=Or5YEhiT_d4&feature=context-...
If the system is sufficiently critical, it may be hard to update or migrate it. And if it's still safely working there's no real incentive to do so.
In the very short term, perhaps. But I'd argue that critical systems are the ones most in need of the ability to be frequently updated and migrated. What's going to happen when a serious security problem demands an immediate change, and you're not prepared for it? Or the system catches fire, or floods?
SPOF critical systems are why so many organisations end up in legacy software hell.
Not my favorite company.
On BSD systems, kernel and userland are developed in lock step, so it's usually a good idea to reboot to be sure they are in sync.
So with that in mind, you can update your userland on FreeBSD without updating your kernel in much the same way as you can do with Linux. Though it is recommended that you update the entire stack on either platform.
Applications will usually work, but the base system will likely break in some places.
(On Windows, a lot more updates require a reboot, though, because one cannot replace/delete a file that is opened.)
Which is a good thing and I bet a model in all other OSes, besides UNIX.
UNIX flock model is just flawed.
How many times I crashed something just because I rmed a file that was being used.
Applications do crash when they make assumptions about those files.
For example, when they are composed by multiple parts and give the filename for further processing via IPC, or have another instance generating a new file with the same name, instead of appending, because the old one is now gone.
Another easy way to corrupt files is just to access them, without making use of flock, given its cooperative nature.
I surely prefer the OSes that lock files properly, even if it means more work.
Of course, the proper way to do this in POSIX is to pass the filehandle.
POSIX is actually a pretty cool standard; the sad thing is that one doesn't often see good examples of its capabilities being used to their full extent.
For this, I primarily blame C: it's so verbose and has such limited facilities for abstraction that it's often difficult to see the forest for the trees. Combine that with a generation of folks whose knowledge of C dates back to a college course or four, in which efficient POSIX usage may not have been a concern, and one finds good, easy-to-read examples of POSIX harder to find than they really should be.
I had my share of headaches when porting across UNIX systems.
Also POSIX is now stagnated around support for CLI and daemon applications. There is hardly any updates for new hardware or server architectures.
Actually, having used C compilers that only knew K&R, I would say POSIX is the part that should have been part of ANSI C, but they didn't want a large runtime tied to the language.
* Files are blobs of storage on disk referenced by inode number. * Each file can have zero or more directory entries referencing the file. (Additional directory entries are created using hard links.) * Each file can have zero or more open file descriptors. * Each blob has a reference count, and disk space is able to be reclaimed when the reference count goes to zero.
Interestingly enough, this means that 'rm' doesn't technically remove a file - what it does is unlink a directory entry. The 'removal' of the file is just what happens when there are no more directory entries to the file and nothing has it open.
https://github.com/dspinellis/unix-history-repo/blob/Researc...
In addition to letting you delete files without worrying about closing all the accessing processes, this also lets you do some useful things to help manage file lifecycle. ie: I've used it before in a system where files had to be in an 'online' directory and another directory where they were queued up to be replicated to off-site storage. The system had a directory entry for each use of the file, which avoided the need to keep a bunch of copies around, and deferred the problem of reclaiming disk storage to the file system.
This model is not unique to Windows, rather most non POSIX OSes.
I happen to know a bit of UNIX (Xenix, DG/UX, HP-UX, Aix, Tru64, GNU/Linux, *BSD).
Yes, it is all flowers and puppies when a process deals directly with a file. Then the only thing to be sorry is the lost data.
Now replace the contents file being worked on or delete it, in the context of a multi-process application, that passes the name around via IPC.
Another nice one are data races by not making use of flock() and just open a file for writing, UNIX locking is cooperative.
You could also point out that the hardlink/ref-count concept is not unique to POSIX and is present on Windows.
https://msdn.microsoft.com/en-us/library/windows/desktop/aa3...
> Now replace the contents file being worked on or delete it, in the context of a multi-process application, that passes the name around via IPC. ...
Sure... if you depend on passing filenames around, removing them is liable to cause problems. The system I mentioned before worked as well as it did for us, precisely because the filenames didn't matter that much. (We had enough design flexibility to design the system that way.)
That said, we did run into minor issues with the Windows approach to file deletion. For performance reasons, we mapped many of our larger files into memory. Unfortunately, because we were running on the JVM, we didn't have a safe way to unmap the memory mapped files when we were done with them. (You have to wait for the buffer to be GC'ed, which is, of course, a non-deterministic process.)
http://bugs.java.com/view_bug.do?bug_id=4724038
On Linux, this was fine because we didn't have to have the file unmapped to delete it. However, on our Windows workstations, this kept us from being able to reliably delete memory mapped files. This was mainly just a problem during development, so it wasn't worth finding a solution, but it was a bit frustrating.
But years as a Unix user might have made me biased. I am certain people can come up with lots of stories about how a file being opened prevented them from accidentally deleting it or something similar. I am not saying it is a complete misfeature, just a very two-edged sword.
That policy also means that replacing some critical Windows DLLs means a very special reboot, one that has to complete, otherwise the entire system is hosed.
Also we should remember the first worm was targeted at UNIX systems.
We should also remember that the 2nd worm was targeted at VMS systems (https://en.wikipedia.org/wiki/Father_Christmas_%28computer_w...) and appeared only a month or so after the RTM worm. Imitation is the sincerest form of flattery, no?
But arguments about your system aside, you could mitigate user error by installing lsof[1]. eg
$ lsof /bin/bash
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NODE NAME
startkde 1496 lau txt REG 0,18 791304 6626 /usr/bin/bash
bash 1769 lau txt REG 0,18 791304 6626 /usr/bin/bash
[1] https://www.freebsd.org/cgi/man.cgi?query=lsof&sektion=8&man...You might even be able to script it so you'll only delete a file if it's not in use. eg
saferm() {
lsof $1 || rm -v $1
}
If you do come to rely on that then you'll probably want to do some testing against edge cases; just to be safe.This means I only update the netbook on rare occasions, because when I go to bed, I want to... sleep, you know, not update my netbook. So currently, that thing has about 180 days of uptime (although it spent most of that time sleeping, of course). I have been meaning to install updates and reboot it for months, but during the day, I forget about it, and only think of it as I go to bed... A vicious cycle... ;-)
With a kernel update, one doesn't have to reboot, technically, it is merely required for the update to take effect.