Doubling the speed of jpegtran with SIMD
blog.cloudflare.com
blog.cloudflare.com
(Also, since this is CloudFlare, <insert rant and dream about SIMD happening in LuaJIT>... Thanks CloudFlare!)
-Rpass-analysis=loop-vectorize -Rpass-missed=loop-vectorized
* They say "no": you go on with your life, no biggie.
* They say "yes": you move to SF/London and start learning :)
A lot of people really want to avoid the 2bii tree and will shrink away from any action that seems to have a chance of leading from the bearable status quo to there.
Based on the N-Body portion of the Bench Mark game it only seems like the ICC does this.
See a related bug in GCC: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=56309
"conditional moves instead of compare and branch result in almost 2x slower code"
idx = 2*idx + (key > table[idx])
(to move a key down an implicitly-stored binary tree)It generally lines up with what I've observed. Surprisingly, I've found that even arithmetic comparisons like you mentioned ran a bit faster with branches instead of conditional moves. The one case where I have seen comparisons benefit is with SSE code where the comparison results go straight to a register instead of a flag.
As far as I understand, these intrinsics map to SSE instructions (they are 128-bit). One could use the later AVX (256-bit, found in server chips since 2011). Probably they decided to start with the lowest common denomintor of SSE, because they don't have AVX in all of their servers? The term SIMD is generic.
Note: we could also choose to use 256-bit wide YMM registers here, but that would only work on the newest CPUs, while gaining little performance.
I'm not saying they're right and you're wrong, but you make it sound as if they never even considered AVX.
Should that be an #ifdef ?
More useful, I think, would be a runtime check far enough up the stack (i.e., away from tight loops) that it doesn't affect performance. Probably doesn't even need to be that high, it'll branch-predict correctly every time but the first iteration, and the binary bloat probably won't hurt your icache because the number of instructions you ever actually execute stays more or less the same.
parallel jpegtran ::: *.jpg
Combine it with these SIMD improvements, and your batch will be done before you hit enter.We at CloudFlare make sure that our servers run at top notch performance, so our customers' websites do as well!.
Correct me if I'm wrong, but it does seem to me that this is of greater value to CloudFlare in terms of electricity savings than to CloudFlare's customers. Less computing time for CloudFlare at the expense of larger output files for CloudFlare's customers to serve, files which will potentially be served again and again and count against the customer's transfer quota. It might not seem like much, but it adds up when you consider how many times a single image in the wild can be downloaded. Consider the following example:
# Original size is 21796912 bytes
# jpegtran with SIMD
jpegtran -copy none -optimize -progressive bigjpeg.jpg > bigjpeg_jpegtran.jpg
# bigjpeg_jpegtran.jpg resulting size is 19771224 bytes
# mozjpeg 3.1
mozjpeg -copy none -optimize -progressive bigjpeg.jpg > bigjpeg_mozjpeg.jpg
# bigjpeg_mozjpeg.jpg size is 19328832 bytes
That's a difference of 442KB! I'm not emphasizing the computing time here. I do not dispute that this is orders of magnitude faster than Mozjpeg, but isn't it worth doing the extra work to get a smaller file? That file could be served billions of times. My argument is simply that the extra computing effort in the beginning leads to much grander savings in the long run.In my (quick) testing Jpegtran's squishing performance is consistently less than that of Mozjpeg. So what, actually, is right thing to do in the end?