“Most serious” Linux privilege-escalation bug ever is under active exploit
arstechnica.com
arstechnica.com
commit 89eeba1594ac641a30b91942961e80fae978f839 Author: Linus Torvalds <torvalds@linux-foundation.org> Date: Thu Oct 13 13:07:36 2016 -0700
mm: remove gup_flags FOLL_WRITE games from __get_user_pages()
commit 19be0eaffa3ac7d8eb6784ad9bdbc7d67ed8e619 upstream.
This is an ancient bug that was actually attempted to be fixed once
(badly) by me eleven years ago in commit 4ceb5db9757a ("Fix
get_user_pages() race for write access") but that was then undone due to
problems on s390 by commit f33ea7f404e5 ("fix get_user_pages bug").
In the meantime, the s390 situation has long been fixed, and we can now
fix it by checking the pte_dirty() bit properly (and do it better). The
s390 dirty bit was implemented in abf09bed3cce ("s390/mm: implement
software dirty bits") which made it into v3.9. Earlier kernels will
have to look at the page state itself.
Also, the VM has become more scalable, and what used a purely
theoretical race back then has become easier to trigger.
To fix it, we introduce a new internal FOLL_COW flag to mark the "yes,
we already did a COW" rather than play racy games with FOLL_WRITE that
is very fundamental, and then use the pte dirty flag to validate that
the FOLL_COW flag is still valid. 1) On the host, save the following in a file with the ".stp" extension:
probe kernel.function("mem_write").call ? {
$count = 0
}
probe syscall.ptrace { // includes compat ptrace as well
$request = 0xfff
}
2) Install the "systemtap" package and any required dependencies. Refer
to the "2. Using SystemTap" chapter in the Red Hat Enterprise Linux
"SystemTap Beginners Guide" document, available from docs.redhat.com,
for information on installing the required -debuginfo packages.
3) Run the "stap -g [filename-from-step-1].stp" command as root.
From https://bugzilla.redhat.com/show_bug.cgi?id=1384344#c13An interesting counter-example to the idea that "more architectures expose problems that would hide in a monoculture".
If you have an iPhone you can scroll horizontally even if it doesn't look like you can (though I am also annoyed by HN truncating the text)
Not doing this when you paste in a block of written material
Or just doing something like this to indicate you pasted in a block of text"Or just using quotes"
for chicken in chicken:
chicken = chicken['chicken']
chick = chicken['chicken']
print ( chicken + '; ' + chicken )
return "blockquoting is for code"Ubuntu - https://appcanary.com/vulns/45984
Debian - https://appcanary.com/vulns/45983
Amazon Linux - https://appcanary.com/vulns/45992
Centos - no patch yet
If you found this useful, please let me know!
I'll admit I spent a bit of time on your homepage thinking "the use of those birds are a bit twitter". Then I realised you're called App Canary.
edit: https://appcanary.com/vulns made my browser crawl. Please fix that page.
vulns page: yeesh that is slow. This is the first time I'm sharing our vuln pages outside of our logged-in users, and yeah, that index is definitely not ready for public consumption yet.
They really need a lot of work.
Look, the Azimuth people have forgotten more about reliable exploit development than I have ever known, but, no, as stated, this is clearly not true. Not long ago, pretty much all local privesc bugs were practically 100% reliable.
What I think they mean to say is that this is unusually reliable for a kernel race.
I still think, though, that the right mental model to have regarding Linux privesc bugs is:
1. If there's a local privesc bug with a published exploit, assume it's 100% reliable.
2. In almost all cases, whether or not there's a known local privesc bug, assume that code execution on your Linux systems equates to privesc; this is doubly true of machines in your prod deployment environment.
I think this goes for any mainstream OS, Linux is not particularly special here.
Nobody's perfect. Not even Theo.
> Theo
Do you mean OpenBSD ?
Edited to fix.
If you're worried that someone might be trying to deliberately compromise your security, you can't give that person the ability to run code on your system.
It depends. I've seen "oh well if someone has rce they probably have root anyway" used way too many times as an excuse to avoid defense-in-depth measures.
ASLR, NX, and CFI would be an example of a defense in depth stack that is meaningful.
SSH, Fail2Ban, and SPA would be an example of a defense in depth stack that basically just wastes time.
I would be more comfortable with a system where I knew I had to burn the box if I lost RCE on it than I would be with a system that somehow depended on RCE not coughing up kernel, and persistence, to an attacker.
The other thing defense in depth can provide is increased attacker cost. That's why there are economically valuable DRM systems (BluRay's BD+ is an example here). All you have to do is push attacker cost across a threshold (for instance with BD+, that's keeping titles secure past the new release window) to make a defense in depth control valuable.
But if someone has a kernel exploit, probably nothing you've done for defense in depth is going to meaningfully increase costs.
A really good example of this is Spyro 3: The developers set up a system of overlapping checksums (which could in turn bet part of the data being checksummed by other, overlapping, checksums) so that it was virtually impossible to change even a single bit without failing the test. It was eventually cracked, as the check only ran at boot time (it required 10 seconds of disk access, and adding 10 seconds to every loading screen in the game would have been unacceptable), which meant it took over two months for pirates to get a crack working (unusual for the time). And since most game sales come in the first two months...
But that's really just me using this as an excuse to share a bit of technical trivia.
I agree with you that ASLR, NX, and CFI are the most important system level defenses to employ.
This assertion confuses me.
I use fail2ban on boxes I have key-only ssh configured for.
Are you aware fail2ban works for services other than ssh?
If an attacker / script knocks unsuccessfully on my ssh door, other doors are then closed to them.
I also get much (much!) cleaner logs thanks to fail2ban.
I suspect that you're confusing fail2ban and port-knocking (or using fail2ban as a port-knocker).
The point of fail2ban is to prevent an attacker from brute-forcing your server. In a key-only config, the chances of getting brute forced is smaller (by a few orders of magnitude) than getting hit by an asteroid and having the server get hit by an asteroid, so fail2ban doesn't really help.
_In theory_, the same would be true for port-knocking.
However, in practice, sshd can have security holes which a malicious scanner could exploit. And while port-knocking doesn't help against a determined attacker (it's subject to MITM, replay-attacks), it does help with defense-in-depth.
Too bad that sshd can't enforce use of password-proctected keys on the server side..
My point is that in general it would be better to disable password auth and only use key based auth, but only if you could somehow guarantee that the users wouldn't do crazy things like use password-less keys. But as you can't do that on the server-side, what other options do you have?
It's about reaction of the staff to key leak:
>> A HPC center [...] disabled key logins IIRC due to some incident where an attacker had got hold of a password-less key.
This reaction seems just silly.
There's more to it, but that's the flavor of it.
Tell this to the container community. They would have you believe containers are as secure as VMs.
That's certainly a goal, but I've never heard the claim.
This article is very hopeful and positively worded, but at its core it acknowledges that security parity is still a work in progress.
Chances are what you want is "simply" access to a shared folder rather than root.
1. Most users won't be affected by all the exploits (you don't stuff in a VM all models of network cards, SCSI controllers, etc)
2. Many deployments of QEMU (through Xen or Libvirt) are protected by AppArmor/SELinux. This would at least forbid access to /proc/self/mem but I can't say if this is enough to prevent evasion. IMO, this is likely to make the task quite harder.
Your shell runs unconfined because your user role is unconfined. Any process you might start will therefore run unconfined, unless stated otherwise in a policy.
So this exploit will run unconfined and will be allowed writes everywhere on the system.
I once tried the staff_r role on a Fedora 23 system and it worked out of box but there were more errors and it would not be recommended for beginners.
I believe the same goes for apparmor since apparmor only defines "armor" for processes, not for users. How many use pam_apparmor today? [1]
I am actually surprised that sane and safe defaults are ignored and left to user's discretion. Most users think Linux is secure by default.
It's interesting to see Windows going into other direction and locking down more and more by default.
I only wish I had the competence to help out because I think it's a very important effort.
Sad to say that in Fedora 23 I was able to easily put my user into the staff_r role, and thereby confining it. But in fedora 24 there seem to be only three default user contexts defined. Not sure what happened but that likely means I have to define my own user context and then I can't know how well supported it is in the policy.
It's impossible for ordinary users to do any of this.
The scale is so high, that kernel security has become a major discussion subject.
http://arstechnica.com/security/2016/09/linux-kernel-securit...
Just to clarify this, any process you start from the shell. Like the PoC exploit.
But in an actual scenario, if the exploit were launched from Firefox, or Nginx, it would run under a confined context and be prevented from overwriting most critical system files.
It's invaluable in setting up new policies.
Doesn't this actually validate Andrew Tannenbaum's argument[1] over 25 years ago when he said monolithic operating systems are inherently insecure and a rethink is required.
[1] https://groups.google.com/forum/m/?fromgroups#!topic/comp.os...
While it's true that vendor drivers living in kernel space is horrible for security... that's somewhat offtopic here. This particular bug is in the memory management system, which is one of those things that kind of has to be in the kernel. A microkernel architecture seemingly would not have helped in this particular case.
I am not very good at theory of Operating Systems but since Memory Management is separated from kernel it would have been difficult for a memory bug to impact other subsystems.
Another argument is modularity which would have allowed better testing hence lesser chances of bugs.
Not really. L4 family (the post-Liedtke world) and Minix3 both have MM out of the kernel.
Firstly, as the MM daemon runs on its own process and is well-separated from other code, it is far easier to audit, debug and so on. Its interface is also entirely explicit. There's value in modular programming. It's far more reasonable to expect quality from such a MM daemon than the mess in a random monolith kernel.
Secondly, in seL4, physical pages are capabilities. There might be more than one MM daemon, owning separate sets of capabilities to physical pages. Security-critical memory might be managed by a MM daemon your vulnerable process has no capability to talk to.
Just my two cents.
Which one will a security minded person pick?
I'm not sure what you're asking.
C, due to arrays, strings, arithmetic operations and memory allocations requiring unsafe code leads to 100% unsafe code across the existing code.
A security minded person will pick those 10%.
Imply is just slightly too harsh. Writing safe C code is very possible, as proven by projects such as seL4 or engineers such as djb.
Since the early 90's I keep hearing that it is possible to write safe C code, yet outside in the real world, unless constrained by processes like MISRA-C and Frama-C, which isn't really C anymore, it never works.
The proof is the amount of CVE exploits, that get reported almost daily!
Just yesterday while reading some papers on Cyclone, I discovered this jewel:
"X El Capitan v10.11.6 and Security Update 2016-004" release notes
https://support.apple.com/en-us/HT206903
From 36 bug fixes, 31 are related C memory corruption issues!
MACH based hybrid kernel garbage.
A shame, considering Apple actually has the resources for doing a proper rebase of XNU on L4 and with actual pure microkernel multiserver architecture.
Of course one always has to validate security, but with C each line of executable line of code is a possibility exploit, which grows exponentially with the amount of developer touching the code and their respective skills and UB knowledge.
CVE-2016-5195
This flaw allows an attacker with a local system account to
modify on-disk binaries, bypassing the standard permission
mechanisms that would prevent modification without an
appropriate permission set. This is achieved by racing the
madvise(MADV_DONTNEED) system call while having the page of
the executable mmapped in memory.
Excellent example why mounting partition with system binaries (such as /usr) read-only is a good idea. CoreOS does this.[EDIT] added "read-only"
What is with that call?
Bryan, if you're reading this, it's merely because I doubt that you actually check Linux bugtrackers.
Also, GNU tail provides tail -F, which does what you want tail -f to do. There is a reason for this. I don't remember what it is, but I think the manpage talks about it.
Ironically, given that you mention M. Cantrill, GNU tail does not really handle truncation properly, and gives up for almost the very case that M. Cantrill did: when the truncation doesn't decrease the size, or is very closely followed by a write that ends up not decreasing the size.
Of course, truncation is not the best way to organize writing log files in the first place. daemontools family style log management (in cyclog, multilog, et al.) starts a fresh file whenever there is a rotation, so these problems of truncation never arise.
And as I then went on to explain, this whole idea of truncating one log file over and over is a poor one, and not the best way to do logging in the first place. So the fact that both M. Cantrill and the GNU people gave up should perhaps be viewed as stopping when an inferior mechanism is pushed beyond its limits.
The ambiguities of language sometimes make two people with the same idea think their ideas are different.
The state of the disk and mounting of the disk generally wouldn't matter because the page is already being forced from read to read/write, and that has no barring on the mounting of or data on the actual disk. It doesn't matter if this data is actually flushed to the disk as long as the kernel uses it from cache without noticing it has been changed (Which it probably doesn't check regardless of read-only status).
If you want a secure system by default, you should probably not use Linux. I would go with OSX or OpenBSD to start.
(And finally: mounting /usr read-only isn't actually a security feature, because if you can exec code you can run a privesc and remount /usr read-write; mounting as noexec could arguably be considered a security feature)
Not really surprising since it's overwhelmingly used in practice as a single-user system.
What would have helped:
* Block ptrace() syscall using seccomp.
* Don't mount /proc, or mount it read-only.
As I understand it, those steps would close all attack vectors for this bug.
FWIW, the Sandstorm.io sandbox blocks ptrace() and doesn't mount /proc at all, so I think the bug has never been exploitable by Sansdtorm apps. (Disclosure: I am the tech lead of Sandstorm.)
I think Docker now defaults to mounting /proc read-only and blocking ptrace(), so it may mitigate this vulnerability as well, but I'm not 100% sure about that.
> Please note that this mitigation disables ptrace functionality which debuggers and programs that inspect other processes (virus scanners) use and thus these programs won't be operational.
> The in the wild exploit we are aware of doesn't work on Red Hat Enterprise Linux 5 and 6 out of the box because on one side of the race it writes to /proc/self/mem, but /proc/self/mem is not writable on Red Hat Enterprise Linux 5 and 6.
Is everyone barking up the wrong tree here?
EDIT: All of the PoCs here use ptrace() or /proc/self/mem. Why would they do that if they didn't need to?
https://sandstorm.io/news/2016-10-25-cve-2016-5195-dirtycow-...
Of course, if you have evidence to the contrary, we'd all like to know about it!
(You are technically correct that the writes can come from another thread rather than another process, but the important part is that it has to go through one of those interfaces.)
(They're actually completely different products.)
"Dirty COW is a community-maintained project for the bug otherwise known as CVE-2016-5195. It is not associated with the Linux Foundation, nor with the original discoverer of this vulnerability. If you would like to contribute go to GitHub."
Seems fishy.
What's with the stupid (logo|website|twitter|github account)?
It would have been fantastic to eschew this ridiculousness,
because we all make fun of branded vulnerabilities too, but
this was not the right time to make that stand. So we
created a website, an online shop, a twitter account, and
used a logo that a professional designer created.
I think the author is just snarking about either branded vulnerabilities or the hype that this issue is getting. or both?Very irresponsible on the part of the author. There's a time and a place for humor - this isn't it.
So my question is: is simply updating and upgrading enough to protect me from this MOST DANGEROUS BUG EVER IN THE WORLD OH MY GOD YOU'RE GOING TO END UP PART OF A BOTNET AND HURT LITTLE CHILDREN!!1!!1! Which is how this reads to even a semi-technical reader, I mean I know my way around the command line but I'm at a loss as to what to do here.
Help me out HN please!
As an engineer you can argue and plead with management to not release something that you don't intend to provide timely updates with a well-communicated support time. Like a 2 year warranty that's prominently communicated, this would highlight to consumers that it's unsafe to use the device unless disconnected from the network. Just like a car that doesn't pass your local safety regulations is not allowed into public traffic.
Actually, I'm surprised modern cars do not require periodic zero-expenses-for-the-owner software updates at licensed dealerships. You can explain to a driver that tires go bad because they drove X miles and have to be paid for, but you cannot argue that software updates need to be paid for because from the time they bought it Y days have passed. Take the Samsung battery optimization that went wrong, where the separation layer was a tiny bit too shallow. It's fair to assume some regulation will follow for safety purposes. Similarly, networked devices, which are not (and cannot be?) microcontrollers with mere 500 lines of code, have to be regulated in terms of software updates.
Now you may say the industry will go broke if they're required to provide upgrades, or less devices will be made, but I think this will lead to consolidation of the software stack, which is mostly a good thing, as those who want to produce dozens of cheap IoT devices can do so without hiring kernel developers. It's like other industries where cheap toy makers source materials like plastic from vendors, knowing it's safe, or create the materials following a detailed recipe which is certified.
Concerning smartphones there are so many privacy and security issues that are far easier to exploit than something that involves kernel hacking... But anyway, isn't Google rolling out security updates for Android? I use CM and I know they don't. There are projects like Replicant which provide a mostly free distribution, but I don't think they're rolling out security updates either. If you're interested maybe contact them?
It's true that there are a high number of bugs available just in mobile browsers, which do receive google play updates, if you have google play, but viewing the underlying code as verified to be correct would be naive.
If I know that a smart phone or smart fridge will not get software updates and be substantially limited in functionality by that, I wouldn't pay more than 100 bucks for it, because I expect to buy another one in probably 14 months.
However, if the update problem would be fixed properly, I wouldn't mind paying a premium.
It seems that this isn't just laziness by the vendors but also calculated into nudging customers to buy new appliances and gadgets although the hardware is capable and perfectly fine. No vendor would admit to that, but this is being investigated and called planned obsolescence. If the price would reflect the artificially limited lifespan of a device, then the problem goes away, and it's just a matter how much of the materials gets recycled.
From looking at the example code, it seems like the general process is:
- Open some (normally un-writable) file as read-only and mmap it in to your process.
- Kick off two threads. One thread to repeatedly write to the same mmap-ed address via /proc/PID/mem and another thread to keep issuing the madvise call.
- Wait for some race condition to be (un)satisfied such that you're able to write to a cached copy of the file.
What I don’t fully understand is how the /proc/PID/mem thing works.
Here’s what I’m curious about:
1. What would happen if you tried to write to the mmap-ed region directly? Since it’s been mapped in with “PROT_READ”, does this mean that you’ll get a segmentation fault or something? From the manpage, it seems like “MAP_PRIVATE” allows it to be a COW mapping, but I don’t see how the combination of “PROT_READ” and “MAP_PRIVATE” is even valid. Unless this means that any writes to data copied from the mmap-ed region into other buffers will be COW-ed and that you can’t actually write to the mmap-ed region itself? That would make sense to me.
2.How is writing to /proc/PID/mem any different than writing through the mmap-ed region directly? Assume that you weren’t running the madvice thread. What would happen then if you tried to write to the /proc/PID/mem file? Presumably the same thing that happens if you just tried to write to the file directly…
3. Finally, how does the madvice call cause a race condition? I realize this might be a little too much to cover in a comment, but this seems like the meat of it.
If only privileged users can SSH into my server, does this really affect me? In other words, I already allow only SSH users to become root.
Interesting whether it will give new root exploits for Android as suggested in the comments.
I'm running Kubuntu 14.04 with the latest security updates, and I'm still on kernel version 3.13.0-98-generic.
~ $ lsb_release -a
No LSB modules are available.
Distributor ID: Ubuntu
Description: Ubuntu 14.04.5 LTS
Release: 14.04
Codename: trusty
~ $ uname -a
Linux anon-pc 3.13.0-98-generic #145-Ubuntu SMP Sat Oct 8 20:13:07 UTC 2016 x86_64 x86_64 x86_64 GNU/Linux
No idea why I haven't gotten an update to 4.x. Should I just switch to a rolling release distro like Arch to have the latest updates of everything?https://www.ubuntu.com/usn/usn-3105-1/ http://people.canonical.com/~ubuntu-security/cve/2016/CVE-20...
~ $ sudo apt-get upgrade
Reading package lists... Done
Building dependency tree
Reading state information... Done
Calculating upgrade... Done
The following packages have been kept back:
ffmpeg libva1 linux-generic linux-headers-generic linux-image-generic
0 upgraded, 0 newly installed, 0 to remove and 6 not upgraded.
The newest available version of linux-image-generic according to apt-cache showpkg is 3.13.0.100.108. (I'm running 3.13.0.98 right now.) Maybe 3.13.100 has the fix to this bug, but I'll have to figure out what's keeping back linux-kernel-image from being updated.What's really puzzling though is that I should have kernle 4.4.x, since I'm running Ubuntu 14.04.5, according to the Ubuntu Wiki: https://wiki.ubuntu.com/Kernel/Support#A14.04.x_Ubuntu_Kerne... It's strange that my Kubuntu installation is frozen on 3.13.x.
After things randomly broke 3 times I decided not to add the backports PPA, and to do manual updates every now and then.
See also: https://help.ubuntu.com/lts/serverguide/automatic-updates.ht...
use dist-upgrade or just explicitly install those packages.
I don't want 16.04; I want to stay on 14.04.
That is a reasonable assumption, but it is incorrect. Check the man page for apt-get.
On Ubuntu systems, the command to upgrade to a new release is "do-release-upgrade". Insanity, but there it is.
If you're always reviewing it manually it's ok to just use dist-upgrade, alternative if you want to install new packages but still not let it remove packages, you can use: sudo apt-get upgrade --with-new-pkgs
Personally I always just use dist-upgrade and it's not a problem as long as you check it before you hit go.
I might still switch to Arch Linux. It's been a hassle to get the latest releases of various packages (like python, gcc, etc). I've had to use third-party PPAs or manually install them. Ubuntu's freezing of packages makes it great as a base image for Docker containers and other reliably reproducible deployment scenarios, but that's not so great as a regular desktop user.
Chris from LAS does say I believe in User Error 6 or 7 that if you don't update Arch in a while you could have stability issues when you update.
[0] https://github.com/dirtycow/dirtycow.github.io/wiki/Vulnerab...
It's not just for debugging, but for any tool that needs some measure of process control. Probably the next most common ptrace-caller I know is "strace".
I'm not sure about this. Ideally, yes, but if you don't know what's causing an issue it can be difficult to reproduce it, and strace can be phenomenally helpful in figuring out the cause. Of course, you could leave it off until you think you might be in such a situation.
Here, it really is a difference between VM and container though.