If you use GNU grep on text files, use the -a (--text) option
utcc.utoronto.ca
utcc.utoronto.ca
The fun part was that the way the data got chunked through the pipes was not deterministic, so sometimes you got the desired output from grep, but other times just "binary file matches", even when the raw output from the first tool was identical in both runs. That took quite some head scratching to figure out.
Actually I recently found that coreutils and ls behave fairly well with funny filenames:
Here is an invalid utf-8 byte and then a valid utf-8 sequence
$ x=$'\xce\xce\xbc'
$ touch "$x"
You can list it: $ ls
?μ
And here 'ls' does better than other tools that display filenames. It shows the invalid byte and then keeps decoding with error recovery: $ ls --escape
\316μ
However GNU stat (which I think is also in coreutils) does something similar, but weirdly messed up: $ stat *
File: ''$'\316''μ'
(it looks like it's outputting a valid shell string, except with extra quotes)-----
Most command line tools are not aware of stuff like this. For example you can touch "x$ANSI_TERMINAL_CODES" and if you do "bash x??" or "python x??", then your terminal will change color because of the escape codes printed back to the terminal.
I just changed Oil to use a well-defined format I called QSN (quoted string notation):
http://www.oilshell.org/blog/2020/04/release-0.8.pre4.html#t...
It adapts Rust's string literal syntax to express arbitrary byte strings precisely and losslessly. (JSON can't express arbitrary byte strings.)
The QSN encoder does UTF-8 decoding with a specific error recovery mechanism. So it's basically like what ls and stat do, but it's more precise.
(If anyone is interested in QSN, please contact me. I think it's more generally useful in a lot of places. It's something we already do but it's precise like JSON.)
https://www.gnu.org/software/coreutils/quotes.html
At least with Gnu, you can recompile your own, non-broken version, which is the only saving grace of these stupid, trendy changes.
I deal with a lot of filenames with spaces and think this change is a great improvement for listing such files. With this change it’s much easier to see where one filename ends and the other begins. Before this change, I had to use the `-1` option to ensure that each filename was listed on a line by itself. Now the filename listings are much more readable and it takes less cognitive effort to take it all in.
The way it handles filenames with ASCII apostrophes/single quotes works particularly well (wraps the filename in double quotes instead of single quotes) and makes it very easy to copy and paste filenames to and from the terminal.
Best of all, this change only applies when standard output is a TTY device so this does not break any shell scripts (even though parsing `ls` is a bad idea in any case) and is still compliant with the POSIX specification[1] which states that “If the output is to a terminal, the format is implementation-defined”.
1. https://pubs.opengroup.org/onlinepubs/9699919799/utilities/l...
I guess I should have figured that oblique references to "ls quoting fiasco" is shorthand for "I don't understand what's wrong but I'm angry about it..."
(On the other hand I would say the grep -a issue is bad both before and after because either way it relies on autodetection. The fundamental issue there is that there is too much variance in encodings, which isn't easy to fix. Luckily UTF-8 is growing in popularity, and it doesn't have this issue because it doesn't require metadata for extremely common operations like "find ascii substring".)
If you’re a human, yes. If you’re a script, it breaks you in half. If you’re a script that has to run on various versions, then maybe it’s time fix yourself and use find. You’re a sophisticated script after all, not one of these who require a human with a debugger. Modern culture may not appreciate that little ‘compat’ thing, but it is essential if you want something to continue to work and not just stop and wait for someone’s educated guesses. Good software doesn’t point fingers at you, it just works. I remember how recently I wanted to check network interfaces on some machine and commanded ‘ifconfig’. Now it’s called ‘ip a’, and there is no ifconfig. I can guess the reason – ifconfig was bad and ip is good. There is also an eternal “FAT” label issue in unetbootin app, which resurrects every time Apple changes its fdisk output format (in every release, as it seems). The workaround is to run it with a cli option – a very thing that unetbootin was created for to skip. This is what makes systems so much fun. Without all these cool things, we would just sit there and cry over our uselessness.
ed: I read below that ls does that in interactive mode only, maybe it’s not that bad then.
Take it to Reddit. I know perfectly well what's going on, and also know better then to discuss it with people who act like that.
the point is that this breaks an unbelievable amount of already deployed scripts. The new functionality should be optional, and accessed via switches and aliases if you like it.
this was very much a change only a few people liked, that they decided to force down literally everyone else's throats. it's very poor stewardship.
% touch foo
% touch 'bar baz'
% ls
'bar baz' fooI thought I proof-read that sentence when I wrote it...
Can you be more specific? I don't see any commmit or release on that date.
I guess you had this commit in mind:
https://git.savannah.gnu.org/cgit/grep.git/commit/?id=cd36ab...
Maybe the fix would be to only activate the "detect binary files" code if stdout isatty?
Because it is a nice feature when I do a big grep to find something among my home directory or the entire filesystem. It is certainly annoying to get binary garbage in my terminal. Or maybe the binary detection could get smarter, maybe making the determination on a match-by-match basis ("This line I'm about to output is a kilobyte and half of it is non-printable", say).
Though, ack-grep doesn't seem to avoid putting binary garbage on my terminal, so maybe reasonable to switch to something that isn't so clever? Most of my terminal greping is done with ack these days, so I'd probably be happy with gnu-grep disabling this cleverness.
IIRC the grep heuristic only considers a short prefix of the file. If the garbage comes later, you lose. Unfortunately, this makes things seem a bit unpredictable.
This was changed about five years ago to just keep looking. Which makes things a bit unpredictable in a different way.
At a glance, I couldn't find a reference on the grep page:
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/g...
> The input files shall be text files.
"Text file" is defined in https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1... :
> A file that contains characters organized into zero or more lines. The lines do not contain NUL characters and none can exceed {LINE_MAX} bytes in length, including the <newline> character.
Interesting, especially the part about the LINE_MAX. Even though it kinda makes sense, I would never have thought that having a very long line makes a file a non-text file when all characters are 'normal' characters.
And I also think that this would not be a discussion if it was just about NUL. There is something else that grep does not like (escape characters? Invalid utf8? Wrong codepage? I don't know), but this makes the binary detection annoying.
$ printf 'a\222b' | LC_ALL=en_US.UTF-8 grep a
Binary file (standard input) matches
This is POSIX-compliant, because "character" is defined as:> A sequence of one or more bytes representing a single graphic symbol or control code.
https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
Logs need a lot of attention.
'CWE-117: Improper Output Neutralization for Logs'
That is something probably often forgotten when simply dumping some requests into a log, but at least it should be obvious that the source of the content is untrusted. On the other hand, a log file is a file on your server, so you would probably think of it as nothing dangerous, as everybody has cared about CWE-117, right? ;-)
And I was wondering all the time why mysql reported this strange error "SQL error in Binary file" when the .sql file was clearly a text file...
The speed benefit isn't really a huge factor; working how I'd expect, omitting ignored files, being able to specify file extensions to search, plus simple editor integration. Amazing.
strings file | grep search_pattern
https://sourceware.org/git/gitweb.cgi?p=binutils-gdb.git;a=c...
edit: I would say that if you are doing forensics on an untrusted binary and you are not using a dedicated VM for it then you are not careful enough. objdump, nm are still attack vectors, not to mention debuggers and disassemblers.
This is a serious anti-feature as far as I can see.. can someone clarify otherwise for me?
Wuff, reminds me of the completely incompatible difference between BSD sed and GNU sed
alias rg="rg --no-heading"
It will still show file paths and line numbers. Those can be disabled too, with `rg --no-heading --no-line-number --no-filename`.Maybe it's just me, but that sounds like a bad default. I can definitely imagine people being confused by that.
You can disable all smart filtering (gitignore, hidden, binary) with `rg -uuu foo`. That will search the same stuff that `grep -r foo ./` will.
GNU grep doesn't have look-arounds either. It does have back-references, but that doesn't impact searches that don't use back-references.
I don't think there is really a concise way to describe why ripgrep is faster when comparing apples-to-apples. It depends on the queries and the corpus. The primary reasons are that it makes more efficient use of the hardware with algorithms that utilize SIMD.
If you do an apples-to-oranges comparison (i.e., "why is `rg foo ./` so much faster than `grep -r foo ./`), then the answer is pretty easily "because ripgrep uses parallelism and employs smart filtering by default."
Why don't most grep utilities just do the obvious right thing and support searches on .c* or .h or .txt or whatever? (HN is mangling the text, but you can read the above as "star dot c star," "star dot h," or "star dot txt.")
I've done a lot of work to make ripgrep work well on Windows. But haven't done this. You can usually work around it pretty easily with the -g flag. e.g.,
rg foo -g '*.{c,h}'
or even shorter rg foo -tcI'm basically looking to maintain the same functionality that was available under DOS with Borland's Turbo Grep in the late 1980s. Those old DOS function calls are still emulated in Win32 as far as I know, but a current implementation would normally use FindFirstFile / FindNextFile like so:
C8 name_buffer[MAX_PATH];
strcpy(name_buffer,dir_buffer); // directory to scan, if not CWD
strcat(name_buffer,filespec); // MS-DOS style filespec with optional * and/or ? wildcards (obviously use snprintf, etc. for this, not strcpy/strcat)
HANDLE search_handle;
WIN32_FIND_DATA found;
search_handle = FindFirstFile(name_buffer, &found);
if (search_handle != INVALID_HANDLE_VALUE)
{
do
{
if (found.dwFileAttributes & FILE_ATTRIBUTE_DIRECTORY)
{
continue;
}
strcpy(name_buffer, dir_buffer);
strcat(name_buffer, found.cFileName);
// name_buffer now indicates a single file
// which can be opened with fopen or
// CreateFile or whatever
}
while (FindNextFile(search_handle, &found));
FindClose(search_handle);
}
There's probably already some code like this in the area of the program that's used to implement the -g mechanism for types. The latter seems a bit overengineered when I'm just looking for wildcard expansion, but I can see it being useful in a lot of cases. It just shouldn't be the only way to constrain the file set, IMO. A new syntax isn't expected or desired, at least not by me, as long as the old one works.To put it another way, do you think you could've written ripgrep in another language? In C? In C++? In Go?
As for Go, I don't know. The garbage collector seems likely to be a problem. Ben Boyter wrote about working on a similarish tool in Go and problems with the GC: https://boyter.org/posts/sloc-cloc-code-performance/
At the very least, such a tool in Go would probably require writing your own regex engine if you want to get comparable performance. (You can see how similar tools in Go, such as pt and sift, fall off a performance cliff as soon as you lean too heavily on the regex engine.) Or at the very least, contribute back to Go's standard library and make `regexp` faster. Whether it can match the speed of a Rust/C/C++ regex engine, I don't know. It's an open question I think.
My own experience w/ Rust is limited to just kicking the tires implementing my favorite toy problem, but it was incredibly positive -- my solution ended up faster than my previous best (in C++) while also having extremely-straightforward, clean code.
(Something that particularly blew my mind was crossbeam_channel being (arguably) nicer than Go's built-in channels, and being able to spread work over threads w/ zero possibility of races -- that's some fuckin' cool shit.)
That experience, combined w/ observation of tools like fd and ripgrep, have convinced me that rust is actually unlocking a higher level of quality for software, at least in practice, if not in theory.
Your answer has not disabused me of the perception! :-)
> Although of course there are alternatives to ripgrep written in C or C++, so I'm not sure how compelling of an argument that is.
But they're not as good as ripgrep ;-)
I would naturally agree, but my users make that argument far better than I can. :-)
I said out loud, "Wow" and browsed through ripgrep. Nice work!
Was torn on whether you were tailing the HN api with ripgrep to alert you or if it was just the right place and time. I had a good time going through how that would work, so, thank you haha!
Occam's razor, right?
I do occasionally manually search HN for mentions of ripgrep. But that wasn't how I found this thread.