Bug 1202858 – Restarting squid results in deleting all files in hard-drive
bugzilla.redhat.com
bugzilla.redhat.com
Actual results:
All files are deleted on the machine.
Expected results:
Squid is restarted.
Not many details yet but it sounds similar to the Steam bug [0] from last year.[0] https://github.com/valvesoftware/steam-for-linux/issues/3671
rm -rf "$STEAMROOT/"*[1] https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/issue...
`echo /* `: /bin /dev /etc /lib ...
`echo /*/`: /bin/ /dev/ /etc/ /lib/ ...
`echo "/*"`: /*
`echo "/*/"`: /*/
If you try it with `ls`, you'll find that `ls "/* "` results in `ls: "/* ": no such file or directory`.Edit: Formatting.
set -eu
on top of your bash scripts -- execution will stop on errors (non-zero retvals) and on undefined variables.[[ "$VAR" ]] && rm -rf "$VAR/*"
I think most of these issues stem from the fact that most developers that write shell scripts don't actually understand what they're doing, treating the script as a necessary annoyance rather than a component of the software.
Anyways, that is not anything like other programming languages. Checking in that way is error prone and not really an improvement (nor equivalent to set -o).
[[ "$DAEMON_PATH" ]] && rm -rf "$DEAMON_PATH/*"
See what I did there? It's an rm -rf /* bug because "checking variables" is not the answer.In other programming languages, if an identifier is mis-typed things will blow up. E.g., in ruby if I write:
daemon_path=1; if daemon_path; puts deamon_path; end
I get "NameError: undefined local variable or method `deamon_path`"These issues do not always stem from bad developers. Bash's defaults are not safe in many ways and saying "people should just check the variable" isn't helpful here.
set -u
Man page quote: "Treat unset variables and parameters other than the special parameters "@" and "*" as an error when performing parameter expansion. If expansion is attempted on an unset variable or parameter, the shell prints an error message, and, if not interactive, exits with a non-zero status."
some real-world programming languages don't have undefined variables :)
Being the emptys string "" would work just as well.
These bugs are indicative of Bash's design problems. Why is it used for init scripts? And don't even get me started on how Bash interprets filenames as part of the arguments list when using * (e.g. file named "-rf").
Say what you will about Powershell, but having a typed language that can throw a null exception is useful for bugs like these. The filename isn't relevant, and a null name on a delete won't try to clear out of the OS (just throw).
That's not Bash. That's just... programs in Unix. Such is life when everything is stringly typed.
Not just scot free - during the Great systemd War of 2014 is was a talking point for the antis that using anything other than the pure, reliable simplicity of shell for service management was MADNESS!
rm -r "${VAR:-var_is_not_set_so_please_fix_this_script}"
which substitutes the var_is_... if VAR is not set.BTW, I hate hate hate -f. It has two meanings: 1. 'force' the removal 2. ignore any error
I've seen an instance of this sort of bug in my sysadmin career that I remember. It was a Solaris patch which wiped a chunk of the system.
If you're suggesting using parameter expansion, at least suggest the correct one (i.e. one that will give a meaningful error message):
${parameter:?word}
http://pubs.opengroup.org/onlinepubs/9699919799/utilities/V3...set -e and set -o pipefail really should have been the default, rather than an opt-in.
Consider /tmp/test.sh:
set -o pipefail
yes foo | head
$ bash /tmp/test.sh >/dev/null
$ echo $?
141I've collated other mishandling of closed pipes at: http://www.pixelbeat.org/programming/sigpipe_handling.html
2. set -e
3. type an invalid command or run one that returns non-zero
4. "crap, where did my shell go?"
The real answer is that this has not been the default in the time between shells being invented and this comment being posted, and so the squillions of lines of shell script out there in the wild keeping the world turning have not been written with this in mind. Making it the default now would break a lot of things.
With the benefit of hindsight, though, i would say that yes, this should have been the default in scripts. Oh well.
http://mywiki.wooledge.org/BashFAQ/105
disagrees and refers to GreyCat's preference not to use -e at the bottom of the list of 'complications'.
You can use set -e, and turn it off (set +e) for code blocks and things that are problematic. He could also add '|| true', and you may be able to use colon to avoid point problems without turning everything off. These are edge cases and you can easily work around them if you an advanced user.
If you are not an advanced user then you should certainly use -e.
$ diff -u /tmp/a /tmp/b
--- /tmp/a 2015-03-24 08:33:00.021919797 -0400
+++ /tmp/b 2015-03-24 08:33:05.629963015 -0400
@@ -1,5 +1,5 @@
#!/usr/bin/env bash
set -e
i=0
-let i++
+let i++ || true
echo "i is $i"
$ /tmp/a
$ /tmp/b
i is 1
$There were some circumstances where if there was no cache_dir line configured, or if the cache_dir was a link or something, the details are very sketchy in my mind after so much time, but it would end up destroying /.
I'm guessing this is of that same nature.
Every time I've seen such a bug (honestly, not many), it was created when cleaning a temporary dir.
It looks like a bug in the init script; runnign it as squid's user wouldn't have triggered destroying the whole filesystem; likely just squid's config and anything under its /var.
It looks like a bug in the init script...
Ha ha ha ha ha ha ha.
This because what it provides it rapid spin up of containers and VMs, while everything talks to each other via APIs and DBUS.
But this rapidity also leads to issues with field repairs and debugging.
"Everyone" is adopting it because the Linux money is in web servers/services.
Which is what happens when you have every daemon writing their own PID handling code, running as root, in a language whose interpolation rules nobody really understands.
It is quite possible to have the script for PID handling be written once, and imported as needed.
Stopping squid: ................[ OK ]
rm: cannot remove `/boot': Device or resource busy At the time of this writing, RHEL 6.7 is still pre-beta
and this bug was found in an *UNRELEASED* update to squid.
Bad enough, but not like it's out in production. touch /-@
I also always do it in my home directory: touch ~/-@
That's the first thing I do on a new host.When accidentally running rm -f *, the command expands to -@ first, which is not a valid option and makes the command fail before doing any harm
rm: illegal option -- @
usage: rm [-f | -i] [-dPRrvW] file ...if a directory has 1.txt, 2.txt, and 3.txt then << rm * >> expands to "rm 1.txt 2.txt 3.txt" and is then executed.
if you have -@, 1.txt, 2.txt, and 3.txt, that expands to << rm -@ 1.txt 2.txt 3.txt >>, and that can't execute.
(if you really wanted to remove your -@ you'd do << rm -- * >> because a double-dash signals the end of command-line options.)
Can someone knowledgeable about the shell expand on this? I don't dare test it on my machine.
$ touch important
$ chmod 400 important
$ rm *
override r-------- vbezhenar/staff for important? n
$ touch -- -rf
$ ls -l
total 0
-rw-r--r-- 1 vbezhenar staff 0 Mar 24 14:35 -rf
-r-------- 1 vbezhenar staff 0 Mar 24 14:35 important
$ rm *
$ ls -l
total 0
-rw-r--r-- 1 vbezhenar staff 0 Mar 24 14:35 -rf
$ rm -- -rfYou can spin up a droplet and use the online shell tool or ssh in (very easy when you've set up a cert as the droplet can have the cert setup automatically).
Then you can mess about with a droplet as much as you like, virtually speaking. Once you're done then use the control panel to destroy the droplet - it costs a few ¢ a day and if you don't have a droplet in use (which means active or paused; preserving images is cheaper but non-zero) then you don't pay anything.
Basically sign up and have a year of uptime to mess with a full install of various OS with no charge.
Make sure you don't write "rm -rf /*" in the wrong terminal!
Defending against this being the use of -- to signal an end of command line arguments.
That sounds like a bug to me, or at least depending on suboptimal behavior.
Example:
$ echo /*
/bin /boot /dev /etc /home (...)
It's also excellent for scripting, and has far more features than bash.
My 'zsh' has this one too, when I 'rm -rf /some/dir/' always asks if I'm "sure". Truth to be told, I'm not even expecting the text in "stdout" anymore, my finger goes to the 'y' automatically, which means that if I make something stupid it won't be able to protect me :-P
The last couple of years I stopped doing 'stupid things' by stop working on the shell when I'm very tired. That was the cause of my rm-related-incidents in the past :-)
Ooops.
touch /--
rm -rf *
(admittedly, this would be a malicious attempt rather than a careless bug)That way the pain of having to type AsteRISKdeleteALL instead of * for rm events offsets any anxiety by far.
You can also catch the rm and mv to a difectory with quota's you can call a recycle bin, some low end attached storage can be fine as well as not many situations when your wildcard deleting with a time factor. Can accommodate this in your own skulker to clean up in a more organised way overall in a timely manner. and and scripts you can path to the real rm command if needs be, last time I called it P45Generator, but not the finest for readability in any such scripts.
If you uninstalled our software it deleted a major chunk of your Windows registry, crippling your computer. It was a one character error in our script. The first ticket read "Uninstalling [Product] destroys your computer".
I was responsible for customer support. Good times! Was a rough week. We managed to not get sued.
This is a bug from 15 years ago, much much before the sandbox feature was introduced.
A sandboxed iTunes would have prevented that.
[1] See "Powerbox and File System Access Outside of Your Container" at https://developer.apple.com/library/mac/documentation/Securi...
But I'm neither a redhat user nor an OS dev, so I might completely wrong.
In stop():
rm -rf $SQUID_PIDFILE_DIR/*
and in restart(): rm -rf $SQUID_PIDFILE_DIR/*
SQUID_PIDFILE_DIR is hardcoded to "/var/run/squid" at the top of my copy of the init script. But, neither of those rm commands check first to make sure that SQUID_PIDFILE_DIR isn't empty (or, better yet, is in /var and doesn't contain ".."), and either the submitter's copy of the script is mangled or something else somewhere is stomping on SQUID_PIDFILE_DIR in the shell environment....I should grep my init scripts for "rm".
https://github.com/mozilla-services/squid-rpm/blob/master/SO... ... however ... whatever that is, does.
Did they take this upstream init script somehow?
https://bugzilla.redhat.com/show_bug.cgi?id=1102343
I guess they applied that change which was obviously written against a very different init script where the variable is actually defined, got QA to test it and immediately backed it out.
"Thanks Swapna and Red Hat QE for catching this issue before the package was released. Great work!"
Looks like this wasn't released into production.For more details on why this issue doesn't affect squid users: https://rwmj.wordpress.com/2015/03/24/restarting-squid-resul...
Listen to yourself. We've all written dumb bugs. We've all had that one line of code that was an obvious mistake. I still trust people who write a bug here and there because if I didn't I would have to forgo trusting everyone for everyone makes dumb mistakes sometimes.
It very well could have been an instance of the bumblebee bug (https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/issue...) where it's "rm -f / some/file" instead of "rm -f /some/file"
Link to the code if you want to make that claim.
Does it really matter if no one commented "Oops, we screwed up"? It's kinda self-evident that there was a mistake and there's not really much to say; it's clear from the description how bad it is and marking it "Fixed in version" already says it all pretty much.
To me it does. You just deleted someone's hard drive, an apology wouldn't be out of order.
The context matters a lot. This title is linkbait. It omits mentioning that this was not publicly released and the reporter is a QA for Red Hat. Given that context I doubt you'll still agree an apology is so necessary or that this was not handled well.
this is a work tool used by developers and QE
Because as we all know, developers don't benefit from good UIs. One wonders why they would use a webpage at all, rather than simply connect to an ncurses-based bugtracker that only supports an 80x24 terminal.
Know and understand the tools you are working with, and simply click the Modified (History) link at the top to determine the bug timeline: https://bugzilla.redhat.com/show_activity.cgi?id=1202858
I realise my comment was snarky, but both of these responses were in a patronising tone themselves. Does bugzilla have some sort of rabid fanbase akin to the vim/emacs wars?
I don't think anyone else has had a tone other than short-spoken
In return I got that it was 'exceptionally clear' (which it clearly isn't, given there's a few people in this thread that missed it); that UI doesn't matter for developers; a possible insinuation that I don't know the kind of tools devs use; and a follow-up comment that tells me I need to know my tools but then proceeds to completely ignore the use-case it explicitly quoted when telling me what I should do. I don't know how that second one can be seen as anything but patronising.
Continuing on with the theme, your point still doesn't answer any of the issues I had in my downmodded comment regarding what is 'exceptionally clear'. Can you tell at a glance from a meter away? My point isn't about the existence of a field, but its presentation.
If you don't think the UI is iffy, that's fine, we disagree. But let's not make up nonsense about things like devs not benefiting from good UI or offering solutions that don't match the quoted use-case.
I feel this is more due to many people wanting to look at the bug itself, instead of its metadata (due to the title of the link)
> Can you tell at a glance from a meter away?
I seriously can't tell anything that isn't colored, I do not have any idea what is the concept of distanced vision, (short sighted, glasses still make things fuzzy)
I would wonder if the status in all bugzilla implementations would warrant viewable-at-a-meter, I personally find the presentation fine though: it is the first thing I see other than the title, unless I'm specifically not looking for it.
Or, even if you don't want to declare a threshold, publish the current stats on its front page: "Over the last 30 days, our 99-percentile triage wait time was: XX hours."
Similarly, open tickets with priority=urgent should never go 24 hours without a new comment from the owner.
Now there's a good idea. Major open source projects should have software quality dashboards tracking things like that.
We're now seeing hospital emergency rooms displaying their current wait time in minutes on billboards.
Fresh steaming proof as requested:
https://bugzilla.redhat.com/buglist.cgi?bug_severity=urgent&...
All high severity bugs against 7.1 which was relased 16 days ago. Check the dates on half of them. They're before the release date and half of them haven't even been assigned or triaged.
When 7.0 came out, datetimectl and systemd didn't even work properly. Enabling ntp threw dbus errors galore. On some kit it didn't even boot. Total lemon.
RHEL doesn't generally work properly until the .2 releases. I've been using it for 10 years so I've got plenty of experience on the matter.
I would go into detail about the CIFS/smb kernel hangs I've had on 6.x but I've had enough of it by now.
Update: I think if you wanted to find out which critical bugs affected RHEL 7.0 on release, you'd probably want to look at the list of z-stream packages (RHEL 7.0.z) which subscribers have access to. These are bugs which didn't affect the installer or first boot, but were important enough to need fixing in RHEL 7.0 after it went out. (If a bug was critical enough to affect installation or first boot, it would have delayed the release).
"Customers would like to be able to use their IdM users to log on to Window clients that a part of the trusted domain."
"Doc Type: Enhancement"
Couple that with what rwmj said, you've effectively debunked yourself.
"Why is grepping /etc taking so long? Binary files in /etc?!? WTF?!?"
Coming from Debian, Redhat seems to make a lot of irk-worthy choices.
-d, --directories=ACTION how to handle directories;
ACTION is 'read', 'recurse', or 'skip'
-D, --devices=ACTION how to handle devices, FIFOs and sockets;
ACTION is 'read' or 'skip'
-r, --recursive like --directories=recurse
-R, --dereference-recursive likewise, but follow all symlinksYou'll find most distros follow some FHS standard, although there's some differences in interpretation.
(That said, making the package self contained is the most sensible way for the developer to release it. It's just not a good option for a distro package.)