EDIT: Let me hedge that a bit, to advanced AVX instructions, as LLVM can do simple loops and such.
EDIT: Let me hedge that a bit, to advanced AVX instructions, as LLVM can do simple loops and such.
I was going to say maybe it's a new thing, but the following post also talks about SSE4 and is from 2011: http://blog.llvm.org/2011/12/llvm-31-vector-changes.html
Maybe it only supports a subset of SSE4? Do you know the details, compared to other compilers?
EDIT: Sorry, to reply to your question, my concern is not GCC vs. clang; If you want max out your vector ops, I would suggest you should compare to ICC as the "standard", at least on Intel CPUs.
This was on an AVX2 machine, testing the (auto)-vectorization performance of expression templates. Anything with a compile-time unknown stride or a random gather failed horribly with both. Using (semi-)explicit vectorization turned out to be much faster still.
Disclaimer: by "vectorization" I'm referring to SSE4, AVX and AVX2, I haven't had a chance to try out AVX512 yet.
http://releases.llvm.org/4.0.0/docs/ReleaseNotes.html
Really impressive how many new things came to LLVM this year!
Edit: Oh, sorry you meant that other guy's link to LLVM's vectorization tutorial. Ignore my reply ...