LLVM OpenMP Support
blog.llvm.org
blog.llvm.org
The auto-vectorization that I know of is the ability of a compiler to batch several scalar operations of a sequential loop into one single vector operation (SSE). There is still only one processor/core at work, but it processes several items at once.
OpenMP is a standard to ease multithreading, to use several parallel threads (cores) without the trouble of creating threads by hand.
Edit: SIMD auto-vectorization directives are part of OpenMP 4, released 2 years ago.
https://software.intel.com/en-us/blogs/2012/11/05/openmp-40-...
The main reason that auto-vectorization is much more "advanced" is many-fold: 1. Most cores have more resources for auto-vectorization than parallelization. 2. Auto-parallelization has more communication overhead, and as you scale, communication overhead dominates, so doing both at once doesn't help as much as you'd think.
I'm just guessing because I have only vague familiarity with this topic - but implementing that and the rest of OpenMP constructs efficiently does not sound trivial.
pragmas are just a way to consume OpenMP in C++, that part could be done trough library - but leveraging OpenMP implementation, stuff like schedulers, etc. could be useful.