Forcing code out of line in GCC and C++11
xania.org
xania.org
; some code that sets CY on error
jc some_error
...
; more code that sets CY on error
some_error:
jc some_error1
...
; etc.
some_error1:
jc some_error2
...
some_error2:
jc some_error3
...
some_errorN:
; error handling code goes here
The reasoning behind this being that conditional jumps to nearby (within +/-127 bytes) targets are far shorter (2 bytes vs 6) and the error path is rarely taken. When the error-handling code is ahead, they are predicted to be not-taken with the common "always not taken" and "forward not taken" decisions, which agrees with their use of being taken only on error. A bit hard to express this pattern with C, however - and chained gotos, which would be somewhat equivalent, aren't so "high level" anymore.But different languages are appropriate for different problems and there is no shame in dropping down into assembly code for pathological cases like the lua example you cite. In the end a good programmer understands the semantics and execution case far better (i.e. at a higher level) than the computer is able to.
In fact this article is a good example of how you can provide intentionality to a modern compiler so that you don't have to drop into assembly and provide the semantics by extension.
Here it is on godbolt: http://goo.gl/iPCiSy
In practice, I'd be surprised if you can come up with a case where this makes a significant difference. If you are on a hot enough path for this to matter, on a modern processor you probably are running out of the even lower level decoded µop cache, which doesn't cache µops for branches that are not taken. If you aren't in this cache, your efforts are probably better spent making this happen.
Edit: ot's comment about how this affects the size of the parent function and whether it will be inlined is a good point, and might well make a measurable difference in the cases where it is true.
Whether or not this will impact performance depends, and will need some careful profiling. But I can imagine some situations where keeping the "hot" instructions in the cache and the "cold" (error-handling) instructions out of the cache could be beneficial.
Surely that's a decision for the optimizer to make, in the case of __builtin_expect?
If the fast path is small enough, you probably want to inline the function, but the slow path makes it large enough that the compiler doesn't inline it (for good reasons).
If you force the slow path into a separate function, then your function becomes fast path + a call instruction, and it can be inlined.
I've often done this manually, and this is a very cool trick that I'm going to adopt immediately.
EDIT: plus, it's also beneficial for instruction cache/iTLB, as cornstalks pointed out.
The next level of mistake is gathering the profile data from the unit tests. (Hey! Free Input!) Ugh.
Hardcoding "Case 5: is most often" in a library also bit me. If you're giving me the source, let me guide the paths, thanks. Maybe you use Case 5, maybe my code doesn't. (And I just stepped through Clang doing the exact right thing on a switch statement with inlining.)
Not for error cases like the ones described in the article.
The compiler can optimize the layout of the basic blocks according to its beliefs (or hints) on the branch probabilities, but it won't move a code path to a separate function. I don't know if this could be done in compliance with the standard, but I haven't seen any compiler do it.
EDIT: I stand corrected, looks like GCC 4.9 can do it (http://goo.gl/amM4Et). Still, it didn't do it when I needed it :)
(and FWIW, ICC has done this for at least 8 years)
In effect, it will just be a bunch of gotos.