(Also, since this is CloudFlare, <insert rant and dream about SIMD happening in LuaJIT>... Thanks CloudFlare!)
(Also, since this is CloudFlare, <insert rant and dream about SIMD happening in LuaJIT>... Thanks CloudFlare!)
-Rpass-analysis=loop-vectorize -Rpass-missed=loop-vectorized
* They say "no": you go on with your life, no biggie.
* They say "yes": you move to SF/London and start learning :)
A lot of people really want to avoid the 2bii tree and will shrink away from any action that seems to have a chance of leading from the bearable status quo to there.
Based on the N-Body portion of the Bench Mark game it only seems like the ICC does this.
See a related bug in GCC: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=56309
"conditional moves instead of compare and branch result in almost 2x slower code"
idx = 2*idx + (key > table[idx])
(to move a key down an implicitly-stored binary tree)It generally lines up with what I've observed. Surprisingly, I've found that even arithmetic comparisons like you mentioned ran a bit faster with branches instead of conditional moves. The one case where I have seen comparisons benefit is with SSE code where the comparison results go straight to a register instead of a flag.