When you deleted /lib on Linux while still connected via SSH (2022)
tinyhack.com
tinyhack.com
Though I'm moreso tempted to just create a `.1111aaaa-antidumb` directory and store my caches and backups there.
This has also un-inspired me from creating a fast `rm`-esque utility.
Only if the file is too big to fit into the garbage bin, I can unalias rm, rm the thing and then reset the alias immediately after.
In Bash I believe instead of `ls` you can `\ls` to get the unaliased version.
I can’t grasp how “power users” like Linux users are stuck working in primitive environments.
And if you rm/del/Remove-Item on Windows, it will also delete without sending to the recycle bin.
When I delete something I want it deleted. If I make a mistake I grab it from a backup or just re-obtain it.
Just use btrfs, and set up btrbk to snapshot your home directory every 5 minutes, and have its gc keep every snapshot from past few hours, an hourly snapshot for past few days, and weekly for past few months.
The problem with snapshots is that they take up space. The trash has the same issue, with the added drawback that you have to think about emptying it, and if you empty it before realizing you made a mistake, it won't help you.
With a "garbage heap" model, all deleted files would automatically end up on the "heap" when deleted. The heap would be exactly as large as the amount of free space you have left, and shrink when necessary by deleting the oldest files, perhaps with a minimum size (expressed in days) that would require manual action to shrink beyond.
Btrbk can also take disk space into account for garbage collecting snapshots? Not sure.
- ~/.local/share/docker is 29G of data that continuously mutates. - Rust/Cargo produce "target" directories in $PWD whenever they run, which is gigabytes of crap littered all over the place.
Keeping snapshots of home would take up a huge amount of disk space, especially if you ever want to keep more than one.
That said, keeping TWO snapshots for 24hs for "un-deleting" files might be an interesting idea.
Stuff like .cache and podman stuff stays unsnapshotted.
I just recommend to others to snapshot home since most people don't want to bother categorizing their data.
Good idea, for safety of course. Just for completeness sake I would add one should not rely on that especially if they are using XFS and especially if using a hardware raid controller. The recursive rm may complete faster than the inodes are removed due to the nature how XFS among a few other filesystems operate in the background. I doubt many are using hardware raid controllers on their workstations but if one is on a server there is a chance there may be one and those will also perform some transactions in the background and one may get their prompt back sooner than the inodes have actually been removed. This is all edge case off course. People should have a massive .pr0n folder regardless.
Service technician did not see the . in the command
rm -rf ./bin
so he proceeded to run rm -rf /bin
When that didn't work he did what everyone who knows a little bit Linux does and added sudo in front.I was on a terminal from the other side of the globe when the server suddenly started acting weird.
We were able to use scp or rsync (one of then was in sbin or something) to get back the bits from an identical server, which saved me from three days of tedious work :-)
In hindsight of course we should have written the docs in a way that would prevent this exact situation but in the beginning it was just the output of the history command after I had done it, dumped into a document with some explanations.
My lesson from them is when a machine servers acting strangely, do not reboot and instead troubleshoot right away. Had I done that, I could have found the problem and had only minimal partial downtime (I would have had to restart dns afterward). Since I was too busy to do things the right way, my internet was out for a half hour. Your story reminded me of this, since you would have had a bigger headache if you had rebooted.
These are especially bad with services that generally grow their logs very slowly but when something like a net split happens or a server is down they generate as much log data in an hour as they typically do all week. So you get close to full and then an incident happens, and now you have two incidents because the screaming machine goes down with a full disk right after you lose your internal DNS server or what have you.
Many years ago, I've got to recover a remote server where /usr was nuked (and /bin, /sbin, and /lib were all symlinks into now-empty /usr). Ended up writing a one-liner perl to convert /bin/busybox-static from my local machine into a series of:
echo -ne "\x7f\x45..." >>~/busybox-static
...and copy-pasting that, chunk-by-chunk, into a single surviving ssh/bash connection, and then used that busybox binary to pull in from a backup.Use visudo to edit the file of course, because not doing so can blow everything up by rendering the sudoers file unparseable and then everyone is gonna have a bad time.
But also the temptation when altering sudo is to immediately log out as super user and try using sudo to do the new command. If you’ve fucked up the file you might not be able to sudo anymore. So now use your second shell window to undo whatever you just did in the first window.
Whenever one of the daemon tools scripts doing remote forwards needs to be modified, a second reverse forward script is added and the reverse forward from that is used for ssh, before changing the original script. The second script is removed only after confirming the first still works after the edit. This procedure prevents fatfingering from locking out remote access, since if something goes wrong, you just need to redo the previous step(s) until you get things working.
If anyone wants to replicate that, I suggest setting -nNT as arguments to ssh and restricting what the user login can do via sshd_config.
pwd # Then check something?
pushd bin #
rm -rf .
Probably still with pitfallsRun enough commands enough times and you will find Murphy is waiting for you.
If you’re just removing a bin directory one time, odds are low but not zero. If you’re writing a run book for people to use, odds are 100% that you will have to help someone rebuild at least once.
If the problem is defined enough to create an exact series of commands for an operator to execute, it is defined enough to create a script to do it for you
I would say the bigger failure is relying on a human typing things into a terminal rather than automating the tasks or changing the system so the task is no longer needed.
It is impossible to test for human errors in advance, though.
After a raised eyebrow he went on to explain, “because when the machine is broken it’s always because some stupid user did something.
It took me a couple more stupiduser incidents of my own before I instituted a rule of counting to five before hitting enter on any `rm -rf` command.
That's just (should be) standard practice. Another risky command is `sudo dd`/`sudo cat` for writing disk images, always chant the disk device against an fdisk -l listing like a magical spell, lest you nuke your main drive.
It can help if you press enter instinctively after finishing writing the command, and also against fat fingering an incomplete but valid command.
The reason why this keeps happening is that in regular UNIX root is omnipotent and the filesystem is ultimately unprotected. Immutable systems and restricted execution environments may make this a thing of the past.
[1] https://www.wolczko.com/rm.txt [2] https://news.ycombinator.com/item?id=7892471
A coworker telnet'd (or rsh, can't recall) into such a machine to do some maintenance and after a while fat fingered:
umount /
Would you believe it, back then being root meant this was absolutely unprotected and the minicomputer OS (some ancient AIX) dutifully complied.The chaos that ensued is but a blur.
‘Recovery’ wasn’t a strong concept in Linux os installers then, and I wasn’t sure my home directory would survive a reinstall, (not on a separate partition) and I didn’t think I’d be able to reboot successfully in any event.
My brother ran the same distribution but at a school across the country; I was able to recover by carefully pulling what I needed down with ftp to his dorm computers IP. I’m not sure scp even existed then on Linux; I guess maybe we could have used netcat in a pinch. Well, this was already a pinch. In a real pickle.
This was one of the first moments where high bandwidth connections to endpoints really saved me / impacted me, and I haven’t forgotten. For many years after university this kind of direct high bandwidth connection got much harder to achieve, first because we were back in a low bandwidth residential world, then because ipv4 was mostly denied to consumers, then because of nat.
Today this would again be achievable, but with vastly more complexity. For a home computer rescue you’d want Tailscale in both sides. And it’s extremely unlikely you’d be using the same distribution as your sib, much less have the same libc linking. And god help you if you had to restore your systemd directory by hand.
$ ls -R /{lib,usr,bin,sbin}
ls: cannot access '/sbin': No such file or directory
/bin:
sh
/lib:
ld-linux.so.2
/usr:
bin
/usr/bin:
env
Oh right... $ ls -l /usr/bin/env
lrwxrwxrwx 1 root root 65 Mar 21 23:39 /usr/bin/env ->
/nix/store/9m68vvhnsq5cpkskphgw84ikl9m6wjwp-coreutils-9.5/bin/env
$ ldd /usr/bin/env
linux-vdso.so.1 (0x00007ffff7fc4000)
libacl.so.1 => /nix/store/dyizbk50iglbibrbwbgw2mhgskwb6ham-acl-2.3.2/lib/libacl.so.1 (0x00007ffff7fb3000)
libattr.so.1 => /nix/store/vlgwyb076hkz7yv96sjnj9msb1jn1ggz-attr-2.5.2/lib/libattr.so.1 (0x00007ffff7fab000)
libgmp.so.10 => /nix/store/dsxb6qvi21bzy21c98kb71wfbdj4lmz7-gmp-with-cxx-6.3.0/lib/libgmp.so.10 (0x00007ffff7f06000)
libc.so.6 => /nix/store/maxa3xhmxggrc5v2vc0c3pjb79hjlkp9-glibc-2.40-66/lib/libc.so.6 (0x00007ffff7d0e000)
/nix/store/maxa3xhmxggrc5v2vc0c3pjb79hjlkp9-glibc-2.40-66/lib/ld-linux-x86-64.so.2 =>
/nix/store/maxa3xhmxggrc5v2vc0c3pjb79hjlkp9-glibc-2.40-66/lib64/ld-linux-x86-64.so.2 (0x00007ffff7fc6000)I also once did "rm -rf /", was deleting a dir which started with a "[" and accidentally hit "enter" instead of "\." That one taught me the dangers of absolute paths.
Edit: That last one is not quite right, would not have been an absolute path issue, that dir must have ended up in root somehow, can't quite remember the details, been too long.
In scripts things like ../../../../file are a pain to read and assume everything in the script before all those previous dirs worked as it should have and everything is where it should be. Cd to an incorrect absolute path produces an error code so we can be sure we are in the proper dir, cd ../../../ will never produce an error and always succeeds even if you are at root.
Are you a bootcamp coach having watched countless of people at the terminal? What makes you assume this is a general habit?
Some of the best early professional advice I ever received was, in moments like these, to keep my hands off the keyboard for at least a timed minute.
chattr +i /lib/ld-linux.so.2
Sounds tempting.How in 2025 do we not have these types of innovations? Are people afraid of breaking old scripts?
(macOS saves us from a lot of this stupidity, Linux should have something similar. I would love a way to mark folders as “permanent” or as “restore on reboot”)
When realizing what I had done and it was taking too long I powered off the machine. When I told the sysops he looked horrified and asked me how far it got.
The fun thing was all websites and data where mounted from the server to each workstation to make it easy to update source code.
-d, -F, --directory
allow the superuser to attempt to hard link directories (this will probably fail due to system restrictions, even for the superuser)
which implies that that's not quite an absolute limit. I don't see any comment either way on https://illumos.org/man/1/ln , but it's plausible that some version of Solaris had wiggle room; it's a terrible idea for obvious reasons, but there's really no hard technical reason why a system couldn't allow you to create hard links to directories.• Non-root users aren't allowed to make directory hard links.
• Many versions of the userspace program `ln` don't let you do it.
But the `link` system call can, at least on Solaris, if called by user 0. (Not sure about Linux: I tried it once and it didn't work, but I was doing weird things with FUSE and also trying to name the link `..`, so I don't know why it failed.)
https://docs.oracle.com/cd/E88353_01/html/E72487/link-8.html
The illumos man page is less clear:
https://illumos.org/man/8/link
The illumos ZFS driver makes it very clear that this is not allowed under POSIX in a comment and explicitly disallows it in zfs_link():
https://github.com/illumos/illumos-gate/blob/master/usr/src/...
However, It appears that the illumos UFS driver supports this:
https://github.com/illumos/illumos-gate/blob/master/usr/src/...
Presumably, the Solaris 10 UFS driver also supports it (or supported it in older versions of Solaris 10). Given that someone at Oracle likely modified the Solaris man page to differ from the older OpenSolaris man page in illumos, I would expect recent versions of Solaris to disallow this on UFS, but someone would need to check.
That said, I have to recant my previous comment. smw likely linked the directory, which is insane, but would have worked on older Solaris versions if we assume the modern illumos UFS driver is unchanged in this regard.
5 minutes that felt a lot longer.
Yum being a Python program made recovery tedious, since some of the libraries it depended on were installed via RPM and gone. Not a difficult recovery, but tedious.
Just checked Slackware and no more, I wonder if that is a casualty of the /bin /usr/bin merge ?
The article author did not try to recover libraries from the anonymous files, which is probably good considering that only a subset would have been in use and thus only that subset would be recoverable from the anonymous files (unless there are filesystem snapshots).
then only `sln` was statically linked (a variant of `ln`)
today most distros statically link nothing and you're up shit's creek in this situation
This is somewhat wrong. From ksh/bash/zsh, you can run:
(exec -a someappname /arbitraryexecutablepath args...)
This won't work on most ash derivatives (including /bin/sh on Debian, FreeBSD, or NetBSD), but does work on busybox ash.The parentheses prevent the `exec` from actually replacing your current shell, which might be less important for emergency rescues, but which otherwise is often what you want with `exec -a`.