Setting the TZ environment variable avoids thousands of system calls (2017)
blog.packagecloud.io
blog.packagecloud.io
https://unicode-org.atlassian.net/browse/ICU-13694
https://github.com/unicode-org/icu/pull/2213
This affects all packages that have icu as a dependency, one of them being Node.js.
https://github.com/nodejs/node/issues/37271
I discovered this the hard way when some code malfunctioned shortly after daylight savings time kicked in.
https://www.inmotionhosting.com/support/website/speed-up-gre...
How are these tools specified to act? What were the requirements for them?
An interesting clue is how these tools are used at all. Less computer savvy users don't touch the command line anyway, so these are used by power users, developers, system operators and the like. HOW are they used? These tools cannot search through what is considered a text file by regular people, i.e. a Word document, but only what a developer would consider a "plain text" file with pure ASCII, possibly containing in-band markup. In particular, logs, configuration files, and code.
There is a strong tendency to treat these with the "C" locale, making the locale-sensititivy of the tools unnecessary and even harmful. However, another aspect comes into play: These files are sometimes NOT using the ASCII encoding but UTF-8, e.g. Java source files. The next obvious problem with these tools is that they treat character encoding as part of the locale; it should be derived from the file type, with the next problem obviously being that Unix doesn't have a solid concept of file types and cannot distinguish ASCII plain-text files from UTF8 plain-text files, resorting to a crude workaround using an unrelated setting (locale) to make the user solve the problem.
OTOH, maybe I'm just "afraid of the command line" /rant
Most tasks assuming ASCII work fine on UTF-8 too (eg sorting). We do need to get rid of those byte order marks though, it breaks eg concatenation
cliché < clichñ < clich́e
NFC vs. NFD.
Combining marks like accents are placed after the base character.
I bet they have unit tests asserting that the bug is still present.
Locales and encodings (at least ones that could be applied using LC vars) generally all behave the same as long as (1) the characters in question are in ASCII (0-127) range and (2) searches are cases insensitive or occasional false match is acceptable. In my experience, this applies to the vast majority of the grep/find invocation I have seen.
In other words: don't put LC_ALL=C in the script which searches your music or document collections. Most other cases are fine with it.
Same applies to things like Java source code: as long as you are searching for FactoryConstructorIndirectorSingleton, you can treat UTF-8 as ASCII, because ASCII is a subset of it. { is code 123, no matter if its latin-1, "C locale", UTF-8 or iso-8859-3.
Another response mentioned decimal points. In a locale that uses a decimal comma, how does (for example) grep whether an ASCII 46 symbol is a decimal point or full-stop?
And yes, sometimes you have no choice but to care about locales, maybe you are pulling data from badly designed API or parsing files generated by someone who didn't know about "stat -c" and did "ls -l" instead. But I'd argue that in this case you should be explicit about setting locale env, and you shouldn't rely on system settings. It's not like an API will suddenly start returning different data if you move your fetching script to a machine with a different locale.
Why is this a problem? In 99.9999% of Java code, the UTF-8 characters aren't going to trip up an ASCII search.
That doesn't even display on my browser[1]; tried it in Goland[2], doesn't display there either, so that's the rare case 0.0001% that I wouldn't really worry about, because if the code has undisplayable unicode sequences, there's bigger problems than searching.
[1] Chrome, on Mac
[2] Also on Mac
To demonstrate this OP explicitly used "DOTTED CIRCLE" (◌) then added the "COMBINING ACUTE ACCENT" to that. Normally there would be no dotted circle.
I wrote an article on this a while ago: https://richardjharris.github.io/unicode-in-five-minutes/
In a script you want predictability, on the command line you want convenience.
In general, you'd want these tools to behave as with the C.UTF-8 locale. Support unicode but forego the other locale shenanigans such as alternative characters for the decimal point.
It never makes sense to treat files/documents differently depending on a system-wide or application-wide "locale" setting. The same document may be worked on by multiple people in different parts of the world.
These tools are definitely supposed to act locale-specific (e.g. '.' regular expression should match one unicode codepoint even if it is multibyte), but in many cases one uses it on ASCII-only data (e.g. logs from system daemons that runs with LC=C anyways), so using this 'trick' is fine.
Mainly, i use this trick not for speedup, but because grep (with UTF-8 locale) has issues with invalid UTF-8 sequences in data, while grep with C locale accepts them.
Does anyone have a an easy recipe to define my own?
How setting the TZ environment variable avoids thousands of system calls - https://news.ycombinator.com/item?id=13697555 - Feb 2017 (143 comments)
No, the behavior is not correct, because changing /etc/localtime is an asynchronous operation. You're changing a file in the file system and expecting to observe a change in a running process. It's, by design, racy — and that's OK.
Taking this into account, the correct behavior is to keep a timestamp of when the last stat() call was, and only stat() again if it was longer than X ago. Even a few seonds would clamp down the stat syscall rate to ambient noise.
(NB: changing /etc/localtime, and then starting a new process is not asynchronous. But that's not the issue here. Two processes are doing things here — changing /etc/localtime and invoking localtime() without synchronization. You shouldn't, can't and mustn't rely on ordering of such unsynchronized events.)
Well, technically you can do both from the same process (i.e. localtime(), unlink() + symlink() to change /etc/localtime, another localtime()), in that case it would be synchronous.
But in general i agree, if the timeout was small enough (say < 100 ms) i guess nobody would object.
One reply comment of the above link mentions user vs. admin, but that’s not a good argument, because users specifying TZ may still want a long-running background process to observe updates of the timezone file, and also TZ may have been set system-wide rather than by the user.
(Now, when you use TZ=:/etc/localtime, I suppose you could expect that glibc would see that it's a symlink [maybe], and then decide to re-check it each time, but I think the current behavior of not doing that is consistent and sane.)
From what I'm seeing.. it doesn't seem to actually assume that. It does check to see if the TZ name changed between calls and will reload a new tzfile if it has. This same mechanism is what causes the file to be loaded whenever getenv("TZ") returns NULL.
See 'tzset_internal' in time/tzset.c in glibc.
If it needs to load state from a file to do its thing, it should be split into initialization, the actual operation(s), and cleanup. (You know, like creating, using, and destroying an object, but this is C so this will be explicit calls.)
You load the file into memory when I tell you to, you stringify timestamps when I tell you to, using the memory pointer I give you, and you release the memory when I tell you to.
That's what a sane API looks like.
To personalize PHP execution I used to putenv("TZ=America/Los_Angeles") for example at the top of a script based on the user's desired timezone. It wound up making all the other time based calls localized which was great.
date_default_timezone_set() is how I do it now, I wonder if it's as efficient (I should strace it)
Has glibc been updated to alleviate this since?
Default behavior - use current system timezone
Explicit TZ - use specific timezone defined by TZ variable, defined either directly (e.g. TZ="NZST-12:00:00NZDT-13:00:00,M10.1.0,M3.3.0"), or indirectly from file (e.g. TZ=":/usr/share/zoneinfo/Europe/Brussels"). As this defines specific timezone, it is not supposed to change.
TL;DR: no, not correct, put a rate limit on it.
> With TZ set during the same time period results in 8 calls to stat over a 30 second period.
This is interesting but what would be even more interesting is what that means to wall time. My gut feeling is probably not that much.
Disclaimer: I'm working on ClickHouse[1], and it is used by thousands of companies in unimaginable environments. It has to work in every possible condition... That's why we set the TZ variable at startup and also embed the timezones into the binary. And we don't use the glibc functions for timezone operations because they are astonishingly slow.
Also what about alpine musl?
https://stackoverflow.com/questions/4554271/how-to-avoid-exc...