I'll own up to misleading-and-possibly-clickbait title, I just gave up trying to think of anything better without revealing the conclusion.
I'll own up to misleading-and-possibly-clickbait title, I just gave up trying to think of anything better without revealing the conclusion.
I admit that I'm not a C++ expert, but naively I would never have expected clang-format to have any effect on my code. If you stopped and suggested header re-ordering to me my first thought would have been "I guess it must not do that. Maybe in practice header re-ordering doesn't actually matter with reasonable code?".
The title here doesn't describe the contents of the article, but it describes a very important consequence of it that C++ programmers should be aware of. It gives a specific example of the consequence. I'm ok with that.
1) clang-format is not really at fault here at all - the same effect could have happened by adding a new header, reordering the headers manually, switching to a system with a different libc version or a different transitive header include tree, etc.
2) The title doesn't tell you wtf is up. You have to get at least half way through the post to figure that out. It's in the vein of "This one food could wreck your health" posts.
That said, other aspects of clickbait are advertising or monetization to profit from the clicks. Except for the crypto mining bot I load into your browser, and which subsequently writes itself to your MBR or UEFI partition, I have no ads, affiliate links, patreon-alikes or other monetization vectors.
This is a _bad_ clang-format bug.
The article describes false complexity caused by looking atr the issue in the wrong way: a "mysterious" performance problem, to be diagnosed, rather than the obvious consequence of a broken workflow that includes an extremely hazardous operation.
In a voluntary header reordering experiment, all kinds of breakage could be detected with a trivial before and after comparison of compiled code.
I agree that a formatting tool changing the behavior or performance is a serious issue, but I don't think it's a bug in clang-format in that I suspect it is deliberate design choice, and users can opt out if they want. Maybe the default could be changed or more visible warnings added to the documentation, but the real fix here needs to be in glibc.
I’d be pissed if the issue were introduced by a formatting pass.
Given that headers can have order dependency, it is sadly part of a C++ programmer's responsibility to know which do, just as it is their responsibility to know which functions have side-effects, etc. If they are using clang-format, it is then their responsibility to communicate this knowledge to it.
SortIncludes: falseC/C++ headers are literal code-injection mechanisms; they aren't modifying a runtime environment by importing symbols, they inject code into the compilation unit.
Those are some pretty dumb coding standards if that's true.
(I mean, yeah, the original fault is with the C preprocessor and the fact that header order matters, but since we're stuck with that, having those kinds of coding standards is pretty darn silly!)
Guess I won't be clicking on the link now.
If that's the case, the clickbait-y headline's bordering on useless, since the performance differences here have nothing to do with clang itself-- you'd see the same results if you had ordered your includes the same way by hand.
That said, the title is still a lie in that clang-format is not really at a fault here at all: clang-format triggered the issue on my codebase, by sorting header files, and this in turn triggered a libc performance issue.
So as I initially say it, a commit which purely clang-format 'd my codebase caused a massive regression - but the full story is doesn't implicate clang-format at all, really.
Fixed now [1], although the HN title will remain hyphen-free free I guess.
---
[1] https://github.com/travisdowns/travisdowns.github.io/commit/...
It is quite common to have an auto-generated configuration header, for instance; or a precompiled header; or optional headers that, when present, mutate the behaviour of other headers.
Every time you run a configure script for a C project there's a good chance you're interacting with code in this way.
It's not like clang-format was written by idiots who never used the language before. Maybe read up on what the tool actually does before you get all mad at them?
I think it's a harmful anti-feature that's worthy of criticism.
This helps ensure that your specific project level header files have all the necessary forward declares and includes to be a fully functioning and complete header.
In the end though you'll have to make a decision on how to sort the projects and which one is more or less specific than another. Kind of like deciding to just sort by pointer address if all else is equal.
There is no gcc equivalent. In any case there no 1:1 relationship between formatters anyways: you might compiler you codebase with N compilers, but you'll only have one formatter unless you are insane. So my codebases are not "gcc" or "clang" codebases, but basically "linux" codebases - but the formatter is always clang-format.
The crucial __NO_CTYPE define is in <bits/os_defines.h>, which is part of the GCC paraphernalia. I don't know if clang's stdlib would also have to end up defining it.
This really has nothing to do with clang-format: I only mention that once (twice with an update: where clang format actually helps because it reorders stdlib.h after ctype.h) - it's all about header include order.
I was wondering about this only yesterday - tidying up some C++ and C code. I don't remember any standard that says it's OK to re-order headers, and clearly it can matter a lot.
My takeway would be anything that changes in semantics or performance due to header includes should be considered a bug by the (possibly joint) writers of the involved headers.
clang-format is an innocent victim here. Maybe I should make that clearer.
I don't think there is a good single-sentence title that would encapsulate the whole thing without being an already well-known fact, like say "Header include order matters".
Also, I wrote this one as a narrative, not revealing the conclusion at the start - basically a reconstructed version of how I encountered it myself. I don't write all my posts like that and it's not for everyone: but once I've decided on that, it makes no sense to reveal the conclusion in the title, I think?
The HN headline could be changed though if it places a priority on information-full titles.
I don't have a similar principled defense of clickbait titles, either in general or mine specifically! I just thought it was a funny one and didn't expend more effort on coming up with a non-clickait one.
That's an independent issue from revealing more info in the title.
I'm just explaining the HN title rules, you can see them here (do a find title):
https://news.ycombinator.com/newsguidelines.html
There's nothing in there about titles having to be 'information-full'. Lots of perfectly HN-good titles aren't.
I didn't submit it to HN so I can't comment on the title choice, other than say it at least reflects the literal title of the page which I guess would be the default when submitting.
I tend to outsource my comments to HN, since my blog is a github pages static site and introducing comments without tracking and ads is a pain, but I definitely don't choose my post titles to satisfy HN rules.
This part lost me because I can't see how you can use std::transform (which is in <algorithm>) without hitting this flaw. Unless you edit the headers? I guess you mean the algorithm isn't flawed.
I also didn't understand the title because I don't use clang-format and I skimmed past the single mention of it (my fault).
Regardless, interesting post.
It reflects how I encountered the issue, but the investigation doesn't have much of that, because I was starting from a point where I suddenly had a slow and fast algorithm (the actual scenario was more complicated than shown).
The investigation doesn't have much to do with clang-format because I'm just trying to see why a raw loop is faster, and ultimately has nothing to do with "raw" vs "std" at all, really. Only when I understood that includes and include order matter was I able to map it back to a header order swap made by clang-format.
> This part lost me because I can't see how you can use std::transform (which is in <algorithm>) without hitting this flaw.
No, two ways.
You can get the fast performance with std::transform if you happen to include <ctype.h> before <algorithm>. The order matters.
Similarly, you can get the slow performance with a raw loop if you happen to include <algorithm> before <ctype.h> in the file for the raw loop.
In fact, I'd say that the raw loop and std::transform will have the same performance almost all the time. If they are in the same file, they will have the same performance. If they are in a C++ file, you are almost certainly including some C++ header that triggers the issue. Only in special cases, like a coding convention that separates C and C++ headers with whitespace with C first (clang-format doesn't sort headers separated by whitespace) are you likely to come into it "natively".
Even if you do, it's absolutely no fault of std::transform or <algorithm> - it's a weird effect of interaction between C, C++ and OS headers, nothing inherent to the C++ algorithms.
In addition to more instructions and a function call, the slow version has many more memory dereferences, so if everything is very cold it is likely to suffer more misses.
IMO this is a library bug. <algorithm> shouldn't be limping toupper or anything else.
That's not really because performance doesn't matter, but because toupper() shouldn't really matter in most modern C and C++ programs: it just can't support Unicode. So either you are using a different way of supporting strings entirely (if you need any kind of Unicode or multibyte support), or this is some internal ASCII-only text processing in which case you are better off using your own methods (e.g., the lookup table I mention at the end) since they'll provide an easy speedup and you aren't accidentally introducing locale-dependent code where you don't want it.
This is really a three part answer, since there were three quite-different effects I glossed over.
The first one is interesting, but it is also well-covered elsewhere, e.g., [1], [2] and [3]. That said, I have a ton of stuff to say on branch prediction, so I probably will at some point. Those posts are slow going, however. It may never emerge.
For the second and third points, I am partly blocked by the fact the effect disappeared or is only intermittently reproducible.
I am more interested in the third uninteresting effect than the second. The second thing could be a lot of things, since it shows up when running a python script with results piped to disk. What I've seen in the past is that the "driver" program (driver.py) undergoes some kind of phase change, e.g., suddenly causing a bunch of context switches or allocating/freeing a munch of memory which affects the program under test (despite CPU pinning, since there are shared resources). They are phantoms (but still interesting!).
The third one is very interesting to me, because it involves a CPU-bound loop with only L1 hits. It goes to the core of the Skylake uarch and the kind of fine-grained uarch details I'm interested in. However, it disappeared as I was tracking it down. I was getting a mix of 2.07 cycles and 1.57 cycles (slow and fast) running the benchmark repeatedly, trying differnet perf counters and doing the needful to track it down and then suddenly I only got 2.07. A new plot showed "blue" only at ~2.07. So I don't know. I have seen similar things though, and I've tracked down an similar "mode shift" and will post on it soon. Stay tuned.
---
[1] https://lemire.me/blog/2019/11/12/unrolling-your-loops-can-i...
[2] https://lemire.me/blog/2019/10/16/benchmarking-is-hard-proce...
[3] https://lemire.me/blog/2019/10/15/mispredicted-branches-can-...