Man accidentally 'deletes his entire company' with one line of bad code
independent.co.uk
independent.co.uk
And today, over two decades later, a person just accidentally destroyed his entire company with one line without warnings on a UNIX. History repeats when its lessons aren't learned. This problem, like setuid, should've been eliminated by design fairly quickly after it was discovered.
EDIT: Added link to ESR's review of UNIX Hater's Handbook which links to UHH itself. Nicely covers what was rant, what was fixed, and what remains true. Linking in case people want to work on the latter plus my sour relationship with UNIX. :)
Now, we can call this a failure of design, but really, people who rely on technology they don't understand can't be saved by good design. Sure, this particular case could be fixed by disallowing the recursive flag on the file system root, but safety is never going to be able to be the primary design concern of any technological system.
Imagine if a sword were made safety as first-class concern. You can't design a sword that can be used safely by the untrained. No weapon can be, training with a weapon is a prerequisite for safely using it. Similarly, every technology has to be understood by those using it. If you don't you're just inviting trouble.
For a business using technology, the needs are actually fairly straightforward. You need an understanding of what needs to be backed up, and a process for performing the backups. If you've picked the former right, (backing up human-readable information rather than data only readable by software programs that might go away in a crash) then risk is minimized.
By the same logic, you could strip away all the airbags, seatbelts, comfy seats and assistance systems of modern cars: After all, accidents still happen and safety mechanisms might even lure drivers into more reckless behavior. (This is in fact happening with seatbelts)
I think it's less useful to think about absolute safety than about which failures are likely to occur and how effective our measures against those specific failures are.
The --force flag is obviously not an effective measure against root deletions, otherwise we wouldn't have so many stories about it. My theory is that there are three reasons for it:
- As other people wrote, if you frequently batch-delete files, you get trained very quickly to always use -f as plain rm is very annoying to use for large sets of files. Unlike other flags, -f won't make you stop and think. This could be fixed by making rm-without-f actually useable - for example by only asking once and not for every file like, oh I don't know... Windows.
- rm can interact with shell parsing in very intransparent and fatal ways. My guess is that most root deletions happen similar to this post: not a literal rm -rf / but some unfortunate variable interpolations where the author didn't realize that they can evaluate to "/". That's a very unobvious point of failure that takes a lot longer to learn than just using rm. Therefore rm should absolutely warn about it.
- there is actually an expectation that rm could be safe as most deletes you do on a modern system are reversible - either because you have a " recycle bin" or a backup. So a warning would make sense to counter that expectation.
It's important to consider when designing software that safety should be about gracefully handling mistakes, and not something that should lure the user into false sense of not having to know what they're doing. Unfortunately, the latter attitude is what drives todays' UX patterns and software design in general, which is a big part of why tech-illiterate people remain tech-illiterate, and modern programs and devices are mostly shiny toys, not actual tools.
So, this isn't some theoretical, abstract, extreme thing some are making it out to be. It's a situation where there's a number of ways to handle a few, routine tasks with inherent risk. Some OS's chose safer methods w/ unsafe methods available where absolutely necessary. UNIX decided on unsafe all around. Many UNIX boxes were lost as a result whereas alternates rarely were. It wasn't a necessity: merely an avoidable, design decision.
It's certainly possible to make a product "safer" then necessary and hinder utility (though I think "safety" is the wrong concept to look at here - see below) but if the common opinion of your product from tech-illiterate people is "complicated and scary", I think you can be pretty sure that you are still a long way away from that point.
In fact, some versions of rm do add additional protection against root deletions, e.g. the --no-preserve-root flag. What utility did that flag destroy?
I believe if you really want to make people more tech-literate (which today's apps are doing a horrible job of, I agree), you have to give them a honest and consistent view of their system, yes. But you also have to design the system such that they can learn and experiment as safely as possible and can quickly deduce what a certain action would do before they do it.
Cryptic commands, which are only understandable after extensive study of documentation, and which oh by the way become deadly in very specific circumstances don't help at all here.
Exactly. That's another problem that was repeatedly mentioned in UNIX Hater's Handbook. It still exists. Fortunately, there's distro's improving on aspects of organization, configuration, command shells, and so on. I'm particularly impressed with NixOS doing simple things that should've been done a long time ago.
Not at all. We should absolutely work to make things safer. But we need to be realistic and temper our sense of idealism. Nothing was going to save this guy from disaster, if it wasn't 'rm' it just would have been something else.
My point is that you can't expect safety features to obviate the need to know what you're doing.
If someone wants safety features, let them pay for them. If someone wants to add one, sure, so long as I can remove it if it gets in my way. Who knows, maybe they'll actually be worth having. But I'm not going to lose sleep over every idiot who ruins his life over something he didn't or couldn't learn about. There's absolutely nothing you can do to save stupid people from making stupid decisions.
Maybe I remove a safety feature I don't need and hurt myself with it. Now I'm the moron. Hopefully I learn from it. Nothing you could have done about that either.
Show me something foolproof, and I'll show you a greater fool.
“A rite of passage”? In no other industry could a manufacturer take such a cavalier attitude toward a faulty product. “But your honor, the exploding gas tank was just a rite of passage.” “Ladies and gentlemen of the jury, we will prove that the damage caused by the failure of the safety catch on our chainsaw was just a rite of passage for its users.” “May it please the court, we will show that getting bilked of their life savings by Mr. Keating was just a rite of passage for those retirees.” Right.
I'm surprised how relevant parts of this book are 22 years later.
http://www.vbcf.ac.at/fileadmin/user_upload/BioComp/training...
Modern OS's will warn you if you try to delete stuff, but you can still ultimately do it anyway, I don't see it as something particular to UNIX.
The only similar problem I had was on windows, 98 I guess, I deleted all my files that weren't readonly by fiddling with a .bat script.
Another option (though last time I tried it, it didn't work..) is something like libtrash: http://pages.stern.nyu.edu/~marriaga/software/libtrash/ Deletes become moves and you can really delete when you like.
Practically speaking, if you're quick an 'rm' isn't totally destructive even without backups. There's a good chance your data is still there on the disk, it's just not associated with anything so it could be overridden at any point. Best to mount the disk read only and crawl through the raw bits to find your lost data (I recovered a week's worth of code this way several years ago).
And then you "actual delete" is where the data loss occurs :D
At the very least when you rm important-file.txt instead of importanr-file.txt you have a chance.
That's very simple and powerful, I can't tell why it is still not implemented today.
You can redesign rm so people don't find themselves typing -rf as force of habit.
You can have multiple delete commands: mark as overwriteable if need and remove from 'ls -a', actually delete, overwrite sectors with zeros; like we have in guis.
You can have a permission system that isn't just "this account can do literally anything" or "this account can't do anything"
Sure it's still possible to --nuke something due to a bug or negligence but I bet it'd cut down on fatal errors. Plus the user could have their own ~/.rmblacklist.conf to guard against particularly persistent dyslexias.
You could even erase the contents entirely if you really think being able to delete root on a whim keeps your kung fu strong.
Then remove the immutable flag if needed
Use something like that with scripting. It really is that easy to protect critical stuff with otherwise easy-to-use rm. UNIX has always been resistant to do stuff like this. However, it's architecture makes it fairly easy for developers to do it themselves. That's to its credit.
Edit: i.e. it isn't so nice. Not much progress.
I was tank crewman in the US military, and there is a high chance of death of dismemberment from regular tank operation. Drivers have very limited visibility, tank turrets can pin people as they traverse, the breech of the main gun violently recoils into the crew compartment with limited guards, the rounds can be exploded by static electricity (or just plain lit on fire), I've saw one guy smash all teeth out when riding inside and the tank hit a hole, and so on.
We had a saying: "tanks are designed to kill, they don't care who."
http://superuser.com/questions/542978/is-it-possible-to-remo...
If a disaster recurs repeatedly -- and if fixing it costs essentially nothing -- then it should be fixed.
http://www.nytimes.com/interactive/2015/03/26/world/europe/g...
This seems like a silly example. A weapon, meant to dismember and maim attackers of its owner, is one of those things that's impossible to completely make safe. Granted I could think of plenty of ideas that maybe would make it safer at first that wouldn't compromise its ability to be effective as a weapon but it's simply not an apt example to be used here.
A computer is a general purpose device. It can be used to help image cancer, launch a nuclear weapon or play games. Considering that it's meant to be used by everyone, without discrimination, it seems to make sense that you need to do the best you can to protect the user from themselves.
I worked in Apple Care support for about a year. The majority of your users are not going to know all of the consequences to their actions, even ones doing system administration (because let's face it almost every company in the world needs at least a little of that now and not all of them are going to hire someone who knows what they're ding).
You can't protect a user from everything. But when you can protect a user from doing something that would have screwed up their whole system, lost a project, etc? That's helpful. Correcting input is what computers are essentially there for.
> It's impossible to have 100% safety, so let's not bother placing any importance on safety.
http://www.smecc.org/The%20Architecture%20%20of%20the%20Burr...
Notice that's good UI design for the time, hardware elimination of worst problems, interface checks on functions, limits of what apps can do to system, and plenty recovery. Systems like NonStop, KeyKOS, OpenVMS, XTS-400, JX, and so on added to these ideas. You can certainly bake a strong foundation of safety into a system while allowing plenty flexibility.
So, for example, critical files should be write-protected except for use by specific software involved in updates or administrative action. Many above systems did that. Then, one can use VMS-style, versioned filesystem that leaves originals in there in case a rollback is needed so long as there's free space for that. Such a system handling backups and restores with modern-sized HD's wouldn't have nuked everything. Might have even left everything if using lean setup but can't say for this specific case.
"You can't design a sword that can be used safely by the untrained."
A sword is designed to do damage. A better example would be a saw that's designed to be constructive but with risk of cutting your hand off. Even that can be designed to minimize risk to user.
https://www.youtube.com/watch?v=esnQwVZOrUU
"If you've picked the former right, (backing up human-readable information rather than data only readable by software programs that might go away in a crash) then risk is minimized."
That's orthogonal. A machine-readable format just needs a program to read it. The risk is whether the data is actually there in whole or part. This leads to mechanisms like append-only storage or periodic, read-only backups that ensure it's there. Or these clustered, replicated filesystems on machines with RAID arrays that lots of HPC or cloud applications use. Also, multiple, geographical locations for the data.
People doing the above with proven protocols/tools rarely loose their data. Then there's this guy.
> ...or administrative action.
You mean like, "sudo rm -rf {$undefined_value}/{$other_undefined_value}"? D'oh!
Anyway, pertaining to RM, here you go:
Also, never having all backup disk volumes mounted at the same time is good practice.
When traction control and antilock brakes became mainstream, one result was that some people started driving faster on snowy roads, up until their risk tolerance was the same as before.
If you understand that a typo can destroy your business, you'll be careful to not log in as "root" on a routine basis and double check everything you do and keep good backups. On the other hand if you expect the system to prevent you from doing anything really damaging, you might be more careless about your approach.
This is a really weird argument, since weapons are used in combat, which is not 'safe' by definition.
But if you do want a weapon that the untrained can use without much chance of hurting themselves, look to a spear. It was the go-to weapon for untrained militias from the time history began up to gunpowder taking over - and even then, bayonets are stuck on rifles to turn them into spears.
path=$foo/$bar
if [[ $path =~ [[:space:]]*/[[:space:]]* ]]; then
echo NOPE
else
rm -rf $path
fi $ sudo rm -rf /
rm: it is dangerous to operate recursively on ‘/’
rm: use --no-preserve-root to override this failsafe
Relatively modern distro, but this has been in coreutils for awhile (it was fixed in Ubuntu's coreutils in 2008 for instance) $ egrep '^(NAME|VERSION)=' /etc/os-release
NAME="Red Hat Enterprise Linux Workstation"
VERSION="7.2 (Maipo)"
Yes, I just did this on the workstation I'm typing on. I'm somewhat curious if he did that on absolutely ancient distros (>8 year old), or didn't actually run rm -rf, but the ansible python equivalent.Also tangentially, if you don't have a sensible backup in place that would protect you from (or at least mitigate) a complete wipe of a single machine (or even all primary ones), you are doing something wrong.
Really, the -f flag should just mean "don't ask for confirmation" and a separate flag should be required to mean something like "yes I do want to nuke my computer". And maybe there should be a flag that means "cross device boundaries", and by default it could refuse to delete anything that has a device number different than the argument it started with. That would at least prevent you from nuking your network-attached storage.
And then someone will come along and bitch that rm isn't safe enough yet again.
Come to think of it, I don't think I've ever used rm to delete across device boundaries. It just doesn't seem like an action you usually want to take.
I don't understand this attitude. Of course software isn't perfect; it's not even close, it's pretty awful. But the best thing about it: it's malleable. When things don't work, you change them to work better.
I'd submit that adding layer upon layer of complexity to prevent all the myriad stupid things people might do using a particular piece of software isn't axiomatically "better".
Maybe if lives depend on its correct function, it's worth it, but that kind of strict requirements gathering and execution is well-understood by the people who live in that world.
Making sure that J. Random DevOps Dude doesn't foot-gun himself when he's paid to know better isn't that.
At my last job, the senior DevOps dude foot-gunned the entire company, by running a read-write disk benchmark (using fio(1)) against the block device (instead of a partition, which, while still stupid, would at least not have been actively destructive) on both my master and all of my slave PostgreSQL hosts. At the same time. And, of course, without telling anyone what he was doing, so the first inkling I had that there was a problem was about 20 minutes later, when I started getting a steadily increasing number of errors suggesting disk corruption.
How does one make such a tool drool-proof enough to prevent that kind of idiocy? Please, help me figure that out. And then give me a time machine, because that was a 16-hour day I'd really rather not have experienced.
And, no, the right move is generally not to fire the jackass who makes that kind of mistake. In my case, above, the company spent about three quarters of a million dollars (just in revenue, never mind how much time was burned in meetings about the incident, my efforts to fix the problem, as well as his and the rest of his team's efforts, and so on) teaching him never to do that again. You don't buy lessons that expensively and then let someone else benefit from them.
(That said, he did get fired several months later for telling the entire engineering lead team to fuck off, in so many words, for their having made a perfectly reasonable request, which was entirely within his responsibilities, and his skills, to satisfy.)
"--one-file-system
when removing a hierarchy recursively, skip any directory that is on a file system different from that of the corresponding command line argument"
When I was in seventh grade, I took an Industrial Arts class ("woodshop"). The first few weeks of the course were spent going over safety of the machinery. In particular, I remember a heavy-handed message that Mr. Hopfer gave at the drill press:
This is a piece of industrial machinery. It is not a toy. If you put your hand on the stage and lower the bit, the machine will not jam up and make funny noises because it is too difficult. Instead, it will drill a hole through your hand. That is what makes it useful. If it didn't do that, you wouldn't be able to cut through wood.
http://www.sawstop.com/why-sawstop/the-technology
Just because people should use something responsibly doesn't mean one shouldn't try to improve its inherent safety.
http://www.amazon.com/SawStop-TSBC-10R2-Cartridge-10-Inch-Bl...
http://www.fairwarning.org/2013/05/after-more-than-a-decade-...
Are you against aircraft collision avoidance too? If your pilot wants to fly into another plane, then the guidance system shouldn't try to stop him, right?
I agree, that's what's so good Unix and GNU-Linux the freedom of root to do anything even major mistakes.
It says "ignore nonexistent files and arguments, never prompt".
Seems to me that the "never prompt" behaviour is an important requirement for a backup-script, though? Cause backup-scripts should work unattended, and definitely not pause and wait for input under any circumstances, right?
Also, perhaps they shouldn't be running as root. There's no reason why any script should have permission to write everywhere. It just needs write access on the backup device.
Like what happened with the Linux version of Steam: https://github.com/valvesoftware/steam-for-linux/issues/3671
https://en.wikipedia.org/wiki/Poka-yoke
Poka-yoke (ポカヨケ?) is a Japanese term that means "mistake-proofing". A poka-yoke is any mechanism in a lean manufacturing process that helps an equipment operator avoid (yokeru) mistakes (poka). Its purpose is to eliminate product defects by preventing, correcting, or drawing attention to human errors as they occur. The concept was formalised, and the term adopted, by Shigeo Shingo as part of the Toyota Production System. It was originally described as baka-yoke, but as this means "fool-proofing" (or "idiot-proofing") the name was changed to the milder poka-yoke.
https://web.archive.org/web/20120213211126/http://m.simson.n...
1. -f specifically forces the change - it's a "I know what in doing, don't warn me" option
2. He was running this in a script, and automating a process so he didn't want a warning
https://news.ycombinator.com/item?id=11499679
Here's an example of a simple alternative that lets you do what this person is doing while avoiding unnecessary hits to critical files:
Any idea why the command actually ran? If $foo and $bar were both undefined, rm -rf / should have errored out with the --no-preserve-root message.
The only way I can think of that this would have actually worked on a CentOS7 machine is if $bar evaluated to , so what was run was rm -rf /.
As the above notes, I'm pretty sure recent versions of Redhat/CentOS actually protect against this sort of thing.
On the offchance you're not running a recent server, however, this could also be avoided by using `set -u` in the bash script, as it would cause undefined variables to error out.
i.e. the variables where happening in a Jinja template and because undefined, rm -rf {foo}/{bar} was transformed by the template engine into rm -rf /
http://docs.ansible.com/ansible/playbooks_filters.html#defau...
foo = ""
Or were set via some function that could return a null bar = getValueOrNull()It is a tragic story but rm -rf has been almost a joke in the industry for a very long while now. Even really old systems should have received an update of some form, to such an extent that the story in the op would be ridiculous rather than a discussion topic.
When I use the command I need to block out all distractions. I check my surroundings for things which might fall on my keyboard. I borderline make sure my phone is turned off before I carefully begin typing that.
I feel uncomfortable typing it into hacker news anywhere but the middle of a sentence. I can't imagine the bullets I would be sweating while deploying a bash script to all servers that included it. There is a problem that needs to be addressed.
Really you just want to use a real programming language instead.
rm: remove write-protected regular file?If a hosting company had deleted 1535 client accounts, we would have heard other stories about it from angry clients?
It's still better than a lot of gaming news sites though, where 'some guy mentioned something on Twitter/Reddit/4chan' is suddenly front page news within ten minutes.
He inverted the `if` and `of` arguments. You'd expect him to pay attention, after what happens. This doesn't pass the smell test for some.
Then again, you'd also expect him to be quite stressed out. That does make that mistake a bit more likely.
He could have swapped arguments like `/dev/sdc` and `/dev/sde` though…
* http://serverfault.com/questions/769357/recovering-from-a-rm...
* http://serverfault.com/questions/769357/recovering-from-a-rm...
Both of those people are ServerFault diamond moderators.
The poster's SO profile is here: http://serverfault.com/users/251721/bleemboy
That lists twitter and github accounts. The github account lists the website of: http://www.thenetworksolution.it/
Which is an Italian provider of web design and hosting services. There is a phone number on the page.
* "more or less 1535 customers"
* ansible would fail if variables are unset (though they could be empty)
* no mention of --no-preserve-root
* who mounts an off-site backup, instead of pushing with scp/curl/whatever
* googling his name does not dig up any company
However, under Linux rm -rf / needs --no-preserve-root to work, right?
"I swapped if and of while doing dd. What to do now?" – Marco Marsala Apr 11 at 7:02
https://serverfault.com/questions/769357/recovering-from-a-r...
(The user who asked that question uses now a nick, but had the real-soundish name mentioned in the article when I read that serverfault question first time earlier this week.)
edit. ...I really hope there isn't a real Marco Marsala someone pretended to be. Search engine results for that name are not great ATM.
https://github.com/samalba/acdcontrol/blob/master/Makefile
While typing quickly I tab-completed to 'Makefile' and hit enter. Although it was a Makefile, it was executed as a bash script. bash ignored the incorrect syntax and executed line 10:
rm -rf $(DIRNAME)/*
If make parsed the file, $(DIRNAME) would have been nonempty. But it was empty under bash.
--no-preserve-root did not protect against this, because the target of the command was '/*'
In my case that's bash, debian based systems use dash.
* http://pubs.opengroup.org/onlinepubs/007908799/xcu/chap2.htm...
* http://pubs.opengroup.org/onlinepubs/9699919799/utilities/V3...
rm -rf foo /
It was supposed to be: rm -rf foo/
It didn't run as root, but still managed to wipe out all the business data files. What saved us was that the servers were configured with RAID 1 and before the start of the nightly batch cycle, the mirror was "split" and only one copy mounted.So we just had to restore the missing files from the other half of the mirror to revert to the start of the batch window and rerun the entire night's jobs.
In the minicomputer era, it was common for a programmer to be required to run it on this one poor donkey of a machine to make it caught nothing on fire before moving to the big machine.
Yes, there were test systems and programers were supposed to test all their changes, but they also were the ones who deployed their own changes to production so there were ways for this to happen pretty easily.
https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/commi...
...and its fix:
https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/commi...
We have code reviews and change controls not only to reduce the number of defects, but also to provide cover when mistakes inevitably slip through.
http://serverfault.com/questions/769357/recovering-from-a-rm...
"the r deletes everything within a given directory"
which is what it does. Non technical readers aren't going to understand "recursively remove directories".
Those two letters are eerily too close to eachother.
My personal crontab is in a separate file in a source control system. I don't use `crontab -e`; I edit that file and feed it to the `crontab` command.
(It would be nice if HN handled backticks the way they're done in Markdown.)
I remember using crontab -r assuming that -r is to open it in read only mode. like vim -R
Bad assumption!
Some platforms it's -r, others it's -d. I suspect it's down to which cron daemon you run but never really cared enough to investigate. In any case, both are next to the 'e' key so either are just as dangerous in terms of typos.
So when someone on their end did something catastrophic to their data and it took them an hour to notice, they were incredulous that we couldn't help them restore their data even though it was "backed up offsite!" because their "backup" solution had already caught up and duplicated the broken data.
>>the code had even deleted all of the backups that he had taken in case of catastrophe. Because the drives that were backing up the computers were mounted to it, the computer managed to wipe all of those, too.
Backups are expected to protect against data loss for a number of different failure cases (eg. disk failure, hardware fault leading to slow filesystem corruption, fire/theft, failed upgrade, "undo" for accidental change or deletion). There is a point where something addresses so few of these failure cases that you can't reasonably call it a backup.
Redundancy is there for fast recovery times (even zero downtime depending on how redundancy is implemented). It's not intended to run as a backup as redundancy devices are live and can fail from many of the same causes that will take your primary devices offline (fire, sysadmin fail, etc)
Likewise, if your "backups" are always online then it works better for business continuity than it does as a backup. So realistically it's more of a redundancy share.
In fact Time Machine on OS X looks like it does backups in this manner...
Do you have your backup servers in the same configuration management software (ansible, puppet, ssh-for-loop etc) as the rest of the servers? One grave error (however unlikely) in your base configuration really can take down everything together in one fell swoop.
How "cold" are your backups? If the backup media are not physically disconnected and secured, you can most likely construct a scenario where the above, malware, a hacker or a rouge admin could destroy both the backups and the live data.
I will certainly suggest some additional safeguards for our backups.
We have backups off-site on disconnected media, so that alone prevents the kind of accident we're talking about.
We use btrfs send / receive to send OS images from the primary container host to the backup container host. The snapshots are read-only, so I'm fairly sure I can't just 'rm -rf' them, I'd have to actually 'btrfs subvolume delete foobar' them.
I should try that though on one of the test servers...
I wish distros would migrate to making those settings the default, over the years. Even if it would take a while, I think it would be priceless
It made me nervous to type rm -rf in this comment form. Those letters are dark magic.
That sounds more like an accident waiting to happen than a single line of bad code.
Maybe things have changed, but rm doesn't zero out the drive. And with the backup that was rm too it should all be recoverable. Or am I missing something?
So yeah, it could technically be recovered, but it's going to be a very big chore.
[0] http://pubs.opengroup.org/onlinepubs/9699919799/utilities/rm...
http://thenextweb.com/media/2012/05/21/how-pixars-toy-story-...
Now that I'm graduating, we've started the process of refining Shill into a product that we can offer to administrators and developers to make their lives simpler. If this sounds like a tool you wish you had (or if you wish a similar tool existed for your platform of choice), we'd love to hear from you.
1 : http://serverfault.com/questions/769357/recovering-from-a-rm...
2 : http://serverfault.com/questions/769357/recovering-from-a-rm...
(Original thread has been deleted.)
A company I left a while back recently had two servers accidentally rebooted through sort of automated task (probably puppet). The fine, I'm told, was one billion dollars.
Someway, somehow, he still works there. :)
With the way they treat their employees - fucking good riddance. There was a giant mess when Disney forced their NOC to train their replacements, but yet these guys did the exact same thing, plus some, and there was no public awareness during or after it. The best part was their push to move everyone to Montreal. Lower pay, not a guaranteed extension and you're forced to move? Okay.
The AMRS CTO actually left about a month after he got the position and took me along with one other person over to a new company. Goldman's head of tech actually just left to go to the same place. Not gonna lie, it sounds incredibly suspicious, especially considering the the kinds of shenanigans that went on there... thankfully I'm no longer working there.
It's a very, very strange place in finance.
Luckily recent backups were available, so the damage was rather small, but it was interesting to see someone just pasting & executing commands without knowing what they actually do, especially when logged in as root.
[0] http://serverfault.com/questions/769357/recovering-from-a-rm...
... Especially when your RAID is busy rebuilding for N hours every other week.
So you could still rm -rf / all you want, delete everything but still have /home or /var/www content untouched.
We run certain programs with limited privileges to mitigates risks (bugs, exploits, etc.), why shouldn't we also limit the privileges of root to mitigates the risk of buggy system administration ?
Obviously having actual backups and testing your code before applying it to production is good practice but I feel like doing system administration with root while having potential bugs in your sysadmin code (as in any other software) leaves the door open to the next catastrophic failure.
"luckily we recovered almost all data!"
And he has no backups? Including rolling backups in unconnected storage?
>Mr Marsala confirmed that the code had even deleted all of the backups that he had taken in case of catastrophe. Because the drives that were backing up the computers were mounted to it, the computer managed to wipe all of those, too.
Then the probably probably deserved to die. Sorry for the customers though...
Perhaps they are unfamiliar with extundelete? http://extundelete.sourceforge.net
"root" is the name of the default administrative account on Unix and Unix like systems.
When i discovered what i had done and stooped it /var/www was already gone.
Luckily we had backups, but that sure did teach me a lesson about rm.
These days i look very carefully before using rm -R and also i type the entire path.
The upside is that we knew we had issues, and with everything broken the impetus is on the right people to ensure they're fixed before we get distracted by the next shiny feature.
Sometimes, setting your servers on fire is the solution to technical debt.
More seriously, this isn't the first I've heard of rm -rf backfiring - one of my friends said at one place he worked at, an IT guy walked out one day & never came back after trying to fix a co-worker's computer. He found out after by investigating on his co-worker's computer that the IT guy must have ran rm -rf while root & wiped out everything.
[0]: https://github.com/sindresorhus/guides/blob/master/how-not-t...
[1] http://waxy.org/2008/04/exclusive_google_app_engine_ported_t...
I never bothered to count exact numbers, but from my experience, close to two thirds of all people, when presented with root shell and no consequences, will run rm -rf in some way.
Humble ones issue "rm -rf /usr" or "rm -rf /lib", others go straight to "/bin/rm -rf /". I've seen one person do "rm -rf /* ", immediately followed by "find / -delete". I'd really like to take a peek on his/her thought process at that moment, looked like the desire of destruction was really strong in that one particular brain ;-)
So yeah, while its not particularly useful one, there's indeed a situation where one definitely want to run it.
disclaimer: I run SELinux playbox with free root access and session recording, and peeking into what others do is also fun.
Edit: rm minus rf slash asterisk formatting.
I'd be super interested, because I cannot think of one at all :D
Sure, there should have been aliases for rm -i and I shouldn't have used -f etc etc etc. But sometimes this stuff is going to happen.
Tape/blu-ray disk backups can come really handy in these cases, not being easy to wipe them.
I guess the best course of action to prevent this would be to alias rm to a custom script, then parse the arguments to make sure the root directory is never recursively deleted, then calling rm from within your script.
https://serverfault.com/questions/769357/recovering-from-a-r...
- No developers have a local copy of code on their machines?
- No backups at all?
Worse case scenario, couldn't you attempt to retrieve the data from the hard-drive? Though, the database(s) would likely not be retrievable.
Maybe I misread the article and he runs a niche hosting company that has different requirements, but it seems strange to me to be able to completely remove your online body of work in a matter of minutes.
This story smells a wee bit fishy to me.