I'm trying to grok the performance improvement here.
The diff points to:
+ /* tmp = rc.code < rc_bound ? rc.range = rc_bound, ~0 : 0; */ \
+ __asm__ ( \
+ "cmpl %3, %2\n\t" \
+ "cmovbl %3, %1\n\t" \
+ "sbbl %0, %0" \
+ : "=&r"(tmp), "+&r"(rc.range) \
+ : "r"(rc.code), "r"(rc_bound) \
+ ); \
This is the only use of assembly in this entire diff. That "tmp = rc.code < rc_bound ?..." statement looks like it'd probably be a cmov, at least I'd expect it to compile to cmov without much issue.
However, this is some advanced "carry-flag manipulation" stuff going on here. I'm not sure if its the "cmov" per se that was advanced, as much as the sbbl statement (subtract borrow, the subtraction-analog to adc). sub %0, %0 is obviously "zero", but sbbl %0, %0 is "0xFFFFFFFF" if carry is 1.
Its certainly a very well thought out bit of manual assembly language. The sbb is probably more important than the cmov (in that the compiler probably emits the cmov, but may not see the sbb????)
----------
The original code seems to be: https://github.com/COMBINE-lab/xz/blob/master/src/liblzma/ra...
#define rc_direct(dest, seq) \
do { \
rc_normalize(seq); \
rc.range >>= 1; \
rc.code -= rc.range; \
rc_bound = UINT32_C(0) - (rc.code >> 31); \
rc.code += rc.range & rc_bound; \
dest = (dest << 1) + (rc_bound + 1); \
} while (0)
I don't know what its doing, but the cmov stuff is very, very different entirely. There doesn't seem to be a branch involved at all in this "rc_direct" inner-loop.
Its not very clear to me how they saw this sequence of C, and the decided upon a cmov / sbb approach to do this equivalent work. Its clearly some kind of advanced thinking that no compiler would have gotten.