https://www.inmotionhosting.com/support/website/speed-up-gre...
https://www.inmotionhosting.com/support/website/speed-up-gre...
How are these tools specified to act? What were the requirements for them?
An interesting clue is how these tools are used at all. Less computer savvy users don't touch the command line anyway, so these are used by power users, developers, system operators and the like. HOW are they used? These tools cannot search through what is considered a text file by regular people, i.e. a Word document, but only what a developer would consider a "plain text" file with pure ASCII, possibly containing in-band markup. In particular, logs, configuration files, and code.
There is a strong tendency to treat these with the "C" locale, making the locale-sensititivy of the tools unnecessary and even harmful. However, another aspect comes into play: These files are sometimes NOT using the ASCII encoding but UTF-8, e.g. Java source files. The next obvious problem with these tools is that they treat character encoding as part of the locale; it should be derived from the file type, with the next problem obviously being that Unix doesn't have a solid concept of file types and cannot distinguish ASCII plain-text files from UTF8 plain-text files, resorting to a crude workaround using an unrelated setting (locale) to make the user solve the problem.
OTOH, maybe I'm just "afraid of the command line" /rant
Most tasks assuming ASCII work fine on UTF-8 too (eg sorting). We do need to get rid of those byte order marks though, it breaks eg concatenation
cliché < clichñ < clich́e
NFC vs. NFD.
Combining marks like accents are placed after the base character.
I bet they have unit tests asserting that the bug is still present.
Locales and encodings (at least ones that could be applied using LC vars) generally all behave the same as long as (1) the characters in question are in ASCII (0-127) range and (2) searches are cases insensitive or occasional false match is acceptable. In my experience, this applies to the vast majority of the grep/find invocation I have seen.
In other words: don't put LC_ALL=C in the script which searches your music or document collections. Most other cases are fine with it.
Same applies to things like Java source code: as long as you are searching for FactoryConstructorIndirectorSingleton, you can treat UTF-8 as ASCII, because ASCII is a subset of it. { is code 123, no matter if its latin-1, "C locale", UTF-8 or iso-8859-3.
Another response mentioned decimal points. In a locale that uses a decimal comma, how does (for example) grep whether an ASCII 46 symbol is a decimal point or full-stop?
And yes, sometimes you have no choice but to care about locales, maybe you are pulling data from badly designed API or parsing files generated by someone who didn't know about "stat -c" and did "ls -l" instead. But I'd argue that in this case you should be explicit about setting locale env, and you shouldn't rely on system settings. It's not like an API will suddenly start returning different data if you move your fetching script to a machine with a different locale.
Why is this a problem? In 99.9999% of Java code, the UTF-8 characters aren't going to trip up an ASCII search.
That doesn't even display on my browser[1]; tried it in Goland[2], doesn't display there either, so that's the rare case 0.0001% that I wouldn't really worry about, because if the code has undisplayable unicode sequences, there's bigger problems than searching.
[1] Chrome, on Mac
[2] Also on Mac
To demonstrate this OP explicitly used "DOTTED CIRCLE" (◌) then added the "COMBINING ACUTE ACCENT" to that. Normally there would be no dotted circle.
I wrote an article on this a while ago: https://richardjharris.github.io/unicode-in-five-minutes/
In a script you want predictability, on the command line you want convenience.
In general, you'd want these tools to behave as with the C.UTF-8 locale. Support unicode but forego the other locale shenanigans such as alternative characters for the decimal point.
It never makes sense to treat files/documents differently depending on a system-wide or application-wide "locale" setting. The same document may be worked on by multiple people in different parts of the world.
These tools are definitely supposed to act locale-specific (e.g. '.' regular expression should match one unicode codepoint even if it is multibyte), but in many cases one uses it on ASCII-only data (e.g. logs from system daemons that runs with LC=C anyways), so using this 'trick' is fine.
Mainly, i use this trick not for speedup, but because grep (with UTF-8 locale) has issues with invalid UTF-8 sequences in data, while grep with C locale accepts them.
Does anyone have a an easy recipe to define my own?