You can get automatic vectorization (with -O3) like this:
bool is_sorted(const int32_t* input, size_t n) {
int32_t sorted = true;
for (size_t i = 1; i < n; ++i) {
sorted &= input[i - 1] <= input[i];
}
return sorted;
}
And the performance is similar to the AVX version (benchmarked on a MacBook Air early 2015): $ ./benchmark_avx2 1048576
input size 1048576, iterations 10
scalar : 6379 us
SSE (generic) : 3544 us
SSE : 3704 us
My example : 2769 us
AVX2 (generic) : 2679 us
AVX2 : 3360 us
So I'm getting 2769us with the above 5 simple lines of code. It's just 3% slower (that might be noise).