Unix Recovery Legend (1986)
ee.ryerson.ca
ee.ryerson.ca
So true. A colleague of mine managed - on his second or third day on job - to delete every single user account in our Active Directory. After an hour, we gave up trying to restore the AD (it was an SBS2008, so no AD recycle bin) and simply restored the entire DC (at the time, our domain only had the one DC) from backup. Surprisingly, most of our users took it very well and used the time to get some paperwork done or clean up their desks or something like that. Still, it was one of the most stressful days of my life. So we kind of panicked. In restrospect, I think another hour or so of research might have saved us the eight hours of restoring that server (did I mention that our backup infrastructure really, really sucked at the time?).
In smaller desasters, I've found the ability to remain calm most valuable, though. Having your boss breathing down your neck impatiently can instill a deep desire to simply do something just to show that you are working on the problem. But if you don't understand what's wrong, at best you are wasting time, and possibly making the problem even worse.
The proactive fix was simple. Backup, restore the table in a two hour outage window. I asked for 5 hours, as the previous DBA never tested a restore. "Unacceptable", said the SVP.
Fast forward two weeks. We hit the limit, and the entire company is essentially down. Between lost revenue, SLA fines and payroll, they were losing something like $5k/minute. Recovery at this point required 30 hours of full database restore, including journal recoveries from slooooow DAT tapes.
There were three of us, everyone stayed calm, provided regular updates and handled a few hours of direct observation by the CEO and a board member for a few hours.
Funny story is that he forgot to pay the phone bill for one of the call centers, and when they went to walk him out, they found him in a "compromising position" with his secretary in the office.
That company was an unlimited source of material! Good times!
The sad thing about such events is that afterwards, you could go "Told you so", but usually, people will not only not listen, but sometimes will still find a way to blame you for what went down. (In our case, it was our mistake, but we were lucky our CEO took it very well - he has no problem with people making mistakes as long as they are open about it and try to learn from their mistakes. What he cannot stand, though, is people trying to cover their butt and/or shift the blame onto others...)
Bonus: a year later, at a different company, I was faced with having to undelete source files on FreeBSD (UFS). I documented that as well here: http://tech.bluesmoon.info/2004/08/undelete-in-freebsd.html
After learning a valuable lesson in exactly how dynamic library work and the recommended process for live libc upgrade (don't do it if ABI changes) I fixed it by using my IRC client which was already running so unaffected to get a statically linked copy of /bin and /sbin from another machine, via DCC Send...
Recovery then consisted of restoring libc5 from slackware 3.2 install media.
I can't remember how I got root, either su was statically linked (believable since it's setuid) or I had a logged in root session. I did have to used the tcsh "echo *" trick for file listing and the shell built-in cd...
python -c 'import os; os.chdir("/")'
Notice that the working directory of your shell is unaffected. :-)
Or:
bash -c 'export foo=bar'; echo $foo
That being said, aside from a kernel patch there is definitely no official way of doing it. This is the hacky way:
http://stackoverflow.com/questions/2375003/how-do-i-set-the-...
Sample:
/tmp$ exec /tmp/cd /bin
/bin$ exec /tmp/cd /home
/home$ exec /tmp/cd /usr
/usr$
Sample code: #include <err.h>
#include <stdlib.h>
#include <unistd.h>
int main(int argc, char *argv[])
{
char *shell;
if (argc != 2)
errx(1, "Usage: exec cd /some/path");
if (chdir(argv[1]) == -1)
err(1, "chdir");
shell = getenv("SHELL");
if (!shell || !*shell)
shell = "/bin/sh";
execl(shell, shell, NULL);
err(1, "exec");
}This seems a little awkward and I'm not sure what is being proven. :P
When they introduced something like the fork/exec model we now have they discovered cd didn't work any more and had to write it as a built in.
Now, for bonus points, work out how goto worked when it wasn't a built in.
Makes sense, since those systems might not have enough memory to run the shell and another program at the same time.
The hoops needed to run it on semi-modern Linux was a real PITA (think chroot jail and some ld shenanigans to patch shared libs). Microsoft's backward compatibility is reallly hard to do.
He was copying directories to the new server and deleting them when he was done. Well...
I don't remember the sub-directory he had just finished copying but when he went to delete it he typed in "rm -rf / SOMEDIR/SOMESUBDIR" and hit enter.
He almost immediately realized what he had done and hit CTRL-C but by that point the damage was done. Our boss had access to the previous week's backups and since the server wasn't critical, he just rebooted it and restored but we had a good laugh about it.
The next day, I made a paper airplane and on the side I wrote in red Sharpie "Linux Air" and then "rm -rf" inside of a red circle with a line through it and taped it to the top of his monitor.
He was a good sport about it and left it up for a couple of months.
"If either of the files dot or dot-dot are specified as the basename portion of an operand (that is, the final pathname component) or if an operand resolves to the root directory, rm shall write a diagnostic message to standard error and do nothing more with such operands."
http://pubs.opengroup.org/onlinepubs/9699919799/utilities/rm...
Warning: not yet widely implemented outside the GNU world.
1) that rm had a (configurable) number of files / directories above which it'll ask you to double-check.
2) that files had a "pleasedontdelete" flag that rm would check and ask about.
For some programs (rm isn't one of them), removing the write permission is sufficient to get an interactive prompt when you try to overwrite/truncate/delete the file.
I really like OSX's current metaphor, where files that haven't been touched in a while get "locked", and must then be "unlocked" to modify them further. Phrasing it in terms of stability rather than permission makes a lot of sense to me. It's too bad the metaphor isn't echoed in the CLI.
$ touch foo
$ chmod -w foo
$ rm foo
rm: remove write-protected regular empty file ‘foo’?
There's also the immutable attribute many modern filesystems support, which prevents the file from being modified unless the attribute is removed. $ sudo chattr +i foo
$ rm foo
rm: remove write-protected regular empty file ‘foo’? y
rm: cannot remove ‘foo’: Operation not permittedIf I'm tapping a file that's been sitting there since the OS was installed and hasn't been touched since... Probably fair to ask me to confirm!
Install safe-rm; it does exactly that.
I'd love to know whoever it was that thought tilde was a good character to use in backup filenames.
[peter@bamboo-sb bin]$ sudo mv /bin/ /usr/local/bin/go1.3 [peter@bamboo-sb bin]$ sudo mv /tmp/go/bin/ /usr/local/bin/go1.3 sudo: mv: command not found [peter@bamboo-sb bin]$ mv -bash: mv: command not found [peter@bamboo-sb bin]$ ls -bash: /bin/ls: No such file or directory [peter@bamboo-sb bin]$ ls -bash: /bin/ls: No such file or directory [peter@bamboo-sb bin]$ ls -bash: /bin/ls: No such file or directory
All during a skype call. It is very recoverable, but I felt quite foolish.
Luckily anotehr client had the same RS/6000 (think 3rd in the country outside IBM) and was able to borrow there install DAT to bring AIX back to life.
Odd as had problem with RT/6150 in which (nobody admitted it) had similiar problem and that involved to get it limping along copying files from a working system onto this holed system to fill the gaps. Which given the eventual reinstall that weekend took most of the weekend only to find that floppy disk 70 odd was corrupt, much fun.
But *nix is great as always more than one way to get things done and on many systems can also be true.
Still good education in not only backups, but backup integrity as you never know when you want to read them back.
If you enjoyed this one, you'll probably enjoy the others in there as well.
http://wiretap.area.com/Gopher/Library/Techdoc/Lore/rumor.ne...
Hard tools for hard admins; the Gods of BSD, McKusick and Joy and the rest, were and are wise. We live in more refined times....
As root, typed in chown -R nobody:nobody / some/dir/somewhere/
Ended up having to re-image the system. We still give him a hard time about it.
Now I'm wondering if he could have tried to find a more creative solution as displayed in TFA.
Thoughts?
Edit: He hard killed the box right after he typed that in.
> ARPA: miw%uk.ac.man.cs.ux@cs.ucl.ac.uk > USENET: mcvax!ukc!man.cs.ux!miw > JANET: miw@uk.ac.man.cs.ux
miw is clearly Mario Wolczko's username.
Those will be his email address in different formats. The first would route email to him at miw@uk.ac.man.cs.ux using cs.ucl.ac.uk first as a gateway. The second address format is what's called a bang path, for UUCP email. The final one is a modern IETF standard email address. JANET is a network provider for UK academic institutions, similar to an ISP but structured differently.
You might find these references interesting:
https://en.wikipedia.org/wiki/Email_address http://www.faqs.org/docs/linux_network/x-087-2-mail.address.... http://www.livinginternet.com/e/ew_addr.htm
The first email address (the ARPA one) is presumably using a machine in UCL as the gateway between internet routing and JANET routing, since the part before the % is the JANET NRS component ordering. The obvious conclusion would be that the mail server he was using didn't have an ARPAnet connection at all.
Well, as a kid I clobbered my windows install while trying to destructively check the validity of a disk. Checks to see if a filesystem is mounted won't protect you if it isn't.
Since then I've been -extremely- careful for any destructive operation on block devices, and generally careful otherwise.
Also, I don't know if blkid existed at the time, but it's /very/ useful in avoiding those types of mistakes.