Unraveling rm: what happens when you run it?
blog.safia.rocks
blog.safia.rocks
The name itself is the giveaway -- "unlink" because you're removing a hard link to a file.
Similarly, access permissions are also properties of the embedding in the directory, rather than the bits of the file itself.
This is a place where POSIX and Win32 diverge significantly -- in Win32, permissions and access happen at the file level, which is why Windows is testy about letting you delete a file that is in use, while POSIX doesn't care -- the process accessing the file, once the file is open, maintains a link to the inode, and all the file data is intact, just not findable in the directory where it was initially located.
A neat trick here is that you can effectively still access the file (even restore it) if a process has the file open, through the /proc filesystem.
No, in POSIX they are properties of the file. You can use fchmod() and fchown() to change mode and ownership via an fd.
> A neat trick here is that you can effectively still access the file (even restore it) if a process has the file open, through the /proc filesystem.
Yes, you can access them; but I don't believe you can link them. (But I'd love to proven wrong. A few years ago I actually needed to un-unlink a non-regular file that was still open.)
But it is probably possible to write simple kernel module that would allow you to do that through some non-standard interface.
I can do this experimentally by creating a symbolic link in /tmp (a different filesystem) to a file in /home, and then creating a hard link (with ln -L) from the symbolic link to another file in /home, and the result is a valid hardlink to the same inode as the original file.
This doesn't work through /proc for an unlinked file, but only because the underlying link call requires a path, not an inode. You can create a hard link out of proc if the file has not been deleted, though, without any cross-filesystem problems.
symlink() does not increase reference count of anything and in fact its target does not have to be meaningful filename at all (although in the practical non-POSIX sense there does not exist any string that is not valid filename). One interesting ab-use of this is that you can use symlink()/readlink() as ad-hoc key-value store with atomicity guarantees (that hold true even on NFS). For example emacs uses exactly this for it's file locking mechanism.
IIRC the files in /proc/pid/fd are not true symlinks but something that behaves as both file (you can do same IO operations as on the original FD) and symlink (ie. you can readlink() them and get some string) at once.
http://man7.org/linux/man-pages/man2/open.2.html
edit: nevermind, seems like O_TMPFILE is the one that has been special-cased here, from man 2 linkat:
> This will generally not work if the file has a link count of zero (files created with O_TMPFILE and without O_EXCL are an exception).
http://man7.org/linux/man-pages/man2/linkat.2.html
:(
linkat(AT_FDCWD, "/proc/self/fd/N", destdirfd, newname, AT_SYMLINK_FOLLOW);
Will do it.
This is the longest-available `flink` syscall method. See https://lwn.net/Articles/562488/Sadly the AT_EMPTY_PATH change was backed out between 3.11-rc7 and release. https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
$ uname -rv
4.15.0-3-amd64 #1 SMP Debian 4.15.17-1 (2018-04-19)
$ touch foo
$ exec 3<>foo
$ rm foo
$ ls -l /proc/$$/fd/3
lrwx------ 1 jwilk jwilk 64 Apr 25 10:02 /proc/324/fd/3 -> '/home/jwilk/foo (deleted)'
$ strace -e linkat ln -L /proc/$$/fd/3 foo
linkat(AT_FDCWD, "/proc/3447/fd/3", AT_FDCWD, "foo", AT_SYMLINK_FOLLOW) = -1 ENOENT (No such file or directory)
ln: failed to create hard link 'foo' => '/proc/3447/fd/3': No such file or directory
$ sudo strace -e linkat ln -L /proc/$$/fd/3 foo
linkat(AT_FDCWD, "/proc/3447/fd/3", AT_FDCWD, "foo", AT_SYMLINK_FOLLOW) = -1 ENOENT (No such file or directory)
ln: failed to create hard link 'foo' => '/proc/3447/fd/3': No such file or directory
+++ exited with 1 +++ open("/tmp", O_RDWR|O_TMPFILE, 0666) = 3
linkat(3, "", AT_FDCWD, "/tmp/bar", AT_EMPTY_PATH) = 0There's still some truth to the path-dependent notion, in that you may not be able to access a file through a hard link in a directory that you do not have access to, even if you have access to that same hard link through another path. But if you don't have access to the file itself then you're out of luck.
This does make sense from a security perspective, but I thought that the path-dependent checks in the kernel were strong enough to not require inode-associated ACLs.
You're right about restoring the file by re-linking to the hard link, but you can access the contents and cp it out of proc at least.
I think you can if you use debugfs. I wrote a post here about recovering a running binary after deleting the file on disk: http://lukechampine.com/recoverbin.html
debugfs(8) manpage says that ln "does not adjust the inode reference counts". Is there a way to increase that number?
set_inode_field foo links_count 1I've never used it, but it appears to operate in read-only mode by default[0]:
> -w
> Specifies that the file system should be opened in read-write mode. Without this option, the file system is opened in read-only mode.
also, being able to 'create' files 'owned' by another user in other locations (by linking them into place) could create quite a few bizarre and undefined corner cases, some of which might have implications for system stability and/or security.
An option to disable this behavior (/proc/sys/fs/protected_hardlinks) was addded only in Linux 3.6, and then it's still disabled by default.
Yes you can. I remember YouTube's flash viewer back in the day would put the downloaded flv video in /tmp and then delete it. I used to check the flash pid, go to /proc/{pid}/fd and see the symlink to the deleted file. Then a cp would give me the actual file.
If you modified the old file after the cp, you wouldn't see the changes in the new one.
That’s not part of POSIX though.
On NTFS permissions live in MFT, which essentially is same thing as inode.
A technique used in the most extreme unix system recovery story I've ever read, by Al Viro:
http://yarchive.net/comp/linux/extreme_system_recovery.html
Wherein:
* system libraries are recovered from init still having them mapped/opened, as you describe
* basic system utilities like 'ln' are recreated from their syscalls and writing the assembly
* ELF binaries are recreated by crafting their headers manually
vlc $(stat -c %N /proc/*/fd/\* 2>&1 | awk -F[\`\'] '/tmp\/Flash/{print$2}')
With newer versions of Flash this too went away (around ~2014). But I still keep a browswer profile around that uses an old one around just so I can access the downloaded file to play in VLC (much smoother).Only when the reference count is zero will the file data be removed.
Not even necessarily then, right? (That is, there's no guarantee that a file's data is zeroed out just because its reference count drops to 0.) It's more just that it's only when the reference count is 0 that the actual space occupied on disk can be overwritten.
There's another tool of that variety: ltrace shows calls to functions in dynamically-loaded libraries. It... works, mostly, but it doesn't know things like function type signatures, so it has to guess, and sometimes gets it wrong in weird ways. It also doesn't know defines and enums, so it can't turn numbers back into symbolic constants like strace and its close kin can:
https://en.wikipedia.org/wiki/Robert_Morris_(cryptographer)
I have no proof of this, but through oral tradition, such a tale has been relayed to me. Believe whatever you will.
Alternatively, we could make "Robert"/"Bob" a euphemism for nuking files ;) "Yeah I Bob'd the whole build directory" "Bombed?", "No, Bob'd. Like deleted... never mind"
https://en.wikipedia.org/wiki/Mark_Zbikowski
https://en.wikipedia.org/wiki/Phil_Katz
("PK" is not only the beginning of ZIP files but also of, among other things, ODT, DOCX, and JAR files, which are all in turn implemented as ZIP files.)
Specifically the original man page (http://minnie.tuhs.org/cgi-bin/utree.pl?file=V1/man/man1/rm....), dated November, 1971, shows Dennis Ritchie and Ken Thompson as the original authors.
$ touch a b c d e f g h
/tmp/x $ strace find . -delete
execve("/usr/bin/find", ["find", ".", "-delete"], [/* 32 vars */]) = 0
[...]
openat(AT_FDCWD, ".", O_RDONLY|O_NOCTTY|O_NONBLOCK|O_DIRECTORY|O_NOFOLLOW) = 4
[...]
fcntl(4, F_DUPFD, 3) = 5
fcntl(5, F_GETFD) = 0
fcntl(5, F_SETFD, FD_CLOEXEC) = 0
getdents(4, /* 10 entries */, 32768) = 240
getdents(4, /* 0 entries */, 32768) = 0
close(4) = 0
fcntl(5, F_DUPFD_CLOEXEC, 0) = 4
unlinkat(5, "d", 0) = 0
unlinkat(5, "e", 0) = 0
unlinkat(5, "f", 0) = 0
unlinkat(5, "b", 0) = 0
unlinkat(5, "c", 0) = 0
unlinkat(5, "h", 0) = 0
unlinkat(5, "a", 0) = 0
unlinkat(5, "g", 0) = 0
[...]
exit_group(0) = ?
+++ exited with 0 +++
https://linux.die.net/man/2/unlinkat int unlinkat(int dirfd, const char *pathname, int flags);
If the pathname given in pathname is relative, then it is interpreted relative to the directory referred to by the file descriptor dirfd (rather than relative to the current working directory of the calling process, as is done by unlink(2) and rmdir(2) for a relative pathname).I would expect the reason for this is in case someone is moving directories around while the find is happening ... Each time find enters a directory it doesn't actually chdir(), so it opens that directory and uses it to anchor the removal.
rm -r does something similar.
This is not correct. "sudo dtruss" makes sudo run dtruss, so it is not possible for dtruss to track what sudo is doing. You can see what "sudo" does (which is much more complicated) by running "sudo dtruss sudo". You need sudo twice because sudo is a setuid program, and obviously it would be insecure to allow anybody to trace a setuid program (for example, it has to read the shadow file, so if you could trace it then you could dump all the hashes and crack them on your own time).
On a separate note, it's probably a bad idea to go reading Linux man pages and expect those to give you accurate information about Mac system calls.
But then I'm not sure what sudo was for... Does dtruss require root privileges?
(similar situation for other unices and posix-likes)
https://www.freebsd.org/cgi/man.cgi?query=dtruss&sektion=1&m...
The dtruss utility traces system calls and (optionally) userland stack
traces for the specified programs.Does anyone know what's going on here? It looks like all syscalls are shown to have 3 arguments, even when they wouldn't need that many... Except close() which is shown to have only one.
There are special cases for 0..6 arguments (including one for close(2)) so I guess getpid(2) slipped through the cracks.
but bsd derivatives essentially convert userland libc 'syscalls' into to a call to a single lower level 'syscall' function which passes data to/from the kernel using a macro / integer list to determine which actual functionality is desired..
This is due to the fact that dtruss is actually a shell script trying to emulate truss with a d script and not quite succeeding.
See the section starting with the comment "print 0 arg output" in dtruss if you wish to understand the meaning of the 3 "arguments."
getpid() is most likely being called with 0 arguments.
To be clear this is due to SIP.
It's probably System Integrity Protection, right?
It can however be used to leak information or read info out of other processes.
I was going to say that for this reason even on linux you can't ptrace processes that are not children of the current process (e.g. you can run something under strace, but not attach to an existing process unless you twiddle a flag or do it as root). Having said that, you CAN modify data with ptrace, unlike dtrace. So that's kindof an aside. In any case the idea is that one process can't hijack another even from the same user for ptrace.
There are a few “destructive actions” that can be enabled with a DTrace flag and also require appropriate system permissions.
Never the less.
[1] https://github.com/coreutils/coreutils/blob/master/src/rm.c
[2] https://github.com/coreutils/coreutils/blob/master/src/remov...
Coreutils is not Linux, and GNU tools are notoriously heavyweight. For contrast, here is Toybox implementation of rm:
https://github.com/landley/toybox/blob/master/toys/posix/rm....
busybox:
https://git.busybox.net/busybox/tree/coreutils/rm.c https://git.busybox.net/busybox/tree/libbb/remove_file.c
And finally openbsd:
On Linux (Debian unstable in my case) you will get order of magnitude more, because rm is dynamically linked (although it seems that only with libc and nothing else), because of libc startup (did you know that linux has amd64-specific syscall arch_prctl(PRCTL_SET_FS), that does exactly what it sounds like?). And then because core utils rm cares about such things as whether stdin is terminal (probably because -f/-i behavior changes depending on that) and does the actual unlink in somewhat convoluted way that involves fstatat() (called twice, for some reason) and only then unlinkat(). Somewhat notably last thing that rm does is trying to lseek() stdin only to get ESPIPE...
E.g. writing your own shell sounds hard but you can do it in an hour (of course it'd be very bare compared to even ash), I once did it during a very casual C class while trying to impress the teacher. He even joked it was self hosting when he has seen the end result (as in - I didn't need other shells anymore and could use vim and gcc from my shell so I could use my shell to work on my shell). I wanted to put that into tutorial at some point but there already exists one[0] (that blog seems quite tinker-y actually, it has implementing own Linux sys call too, which is also surprisingly easy to do and I did it for a class at one point too[1]).
Going knee deep into this stuff also completely dispels the magic that language and system runtimes, filesystems, file formats, linkers, shells, standard commands (e.g. ls, I had to reimplement it once as a homework) or whatever have, or even better - it's still magic to most people and you're the wizard now! It's also very accomplishing to do something so unique in an hour or two (although to me due to my C and C++ bias at some point making a pastebin clone in an hour in PHP or Python became an unique experience).
Even JIT (which sounds scary due to V8, LuaJIT, etc. being so tightly made and complex) is easy to get into and understand at toy scale[2].
[0] - https://brennan.io/2015/01/16/write-a-shell-in-c/
[1] - https://brennan.io/2016/11/14/kernel-dev-ep3/
[2] - http://blog.reverberate.org/2012/12/hello-jit-world-joy-of-s...
He then had to jump through a bunch of hoops and use a bunch of strange commands to restore / because of all the missing binaries. I wish I could find it.
Here's the one I found in my bookmarks: http://lambdaops.com/rm-rf-remains/
It's not the story you want, but the origin story of the ext3grep tool is interesting too: https://web.archive.org/web/20110529114328/https://carlo17.h...
Long story short: they accidentally ran find with an exec of chmod that takes away the x bit on /.
It sounds ouch-y but it's not that bad because it turned out that find went alphabetically (or so) so first it went into /bin and then quickly made chmod unexecutable, so only the stuff that came before chmod in /bin was affected and all it took was to make it executable again normally via Windows.
It's actually even good bash happens to come before chmod in the sorting and that made cygwin completely "broken" at a glance or else it'd Murphy's law its way into "I don't have executable bit on this exotic rarely used command or other" 5 months down the line with the relevant lines of ~/.bash_history and everyone's short term memory long gone.
> I had a handy program around (doesn't everybody?) for converting ASCII hex to binary, and the output of /usr/bin/sum tallied with our original binary.
I copied busybox's nc onto it (needed for the ssh jump); had to copy to /run as / was unwritable due to the disk failure. Now, scp no longer ran, so "copy" was a Python script to turn the binary into a printf command, which is a shell built-in and can write arbitrary binary.
(If you ever get into a jam like this, busybox's utilities are very useful.)
> But hang on---how do you set execute permission without /bin/chmod? A few seconds thought (which as usual, lasted a couple of minutes) suggested that we write the binary on top of an already existing binary, owned by me...problem solved.
I think umask'ing correctly prior to a printf should work nowadays, no? (IDK about in the author's time.) Thankfully in my case, chmod was in disk cache still.
Thank you for finding that.
Close. That's one thing that csops can do, but in this case it's being used to extract the entitlements from the binary (CS_OPS_ENTITLEMENTS_BLOB == 7).
> I wasn’t sure what the memory addresses that were referenced in the mprotect call actually corresponded to or what the best way to figure it would be.
You could fire it up in LLDB and break on mprotect…