I think that is possible.
I've been thinking a lot about vectorizing loops recently, especially on AVX512 systems.
I've mostly been doing microbenchmarks, and I realize that microbenchmarks might not give a realistic full-program view.
LLVM (through Julia and Clang) have a striking performance pattern as sizes vary:
https://discourse.julialang.org/t/ann-loopvectorization/3284...
They are fast at multiples of 32, but then performance degrades. This is because it vectorizes loops by creating two loops:
1. 4x unrolled and vectorized (with double precision and AVX512, that translates to 4 * 8 = 32 loop iterations)
2. Scalar loop.
By avoiding that pattern, it was easy to get much better performance at most sizes in a lot of simple cases, like dot products.
In wondering about why LLVM's decisions made sense for them, I'm currently leaning towards AVX transition penalties being a big factor.
Recently shared on HN:
https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html
The thing that struck me is that there is a 9 microsecond period where avx instructions (AVX2 and AVX512) operate at a small fraction of normal speed before the CPU decides to transition to a slower state.
If most of your code is running in L0 (max clock speed) license, then any vectorized code you run into will run at 1/4 speed for about 9 microseconds. If it does run for that amount of time, it'll transition with an 11 microseconds break. It'll have to keep running for a long time to amortize this penalty.
Then, once the function returns to the rest of your scalar code, it'll eventually have to speed up. Basically, large programs are probably fastest if they stay in relatively the same state.
By having a large scalar window, like LLVM does, it's less likely to change. Most loops are probably fairly short, and most code is also scalar, therefore you'll want the CPU to generally stick to scalar mode. Only if loops are very long and likely to take milliseconds would you want them to be vectorized.
Or if they're surrounding by other SIMD code, but that's a sort of global/whole program state you cannot infer while optimizing a single function.
It is likely best to go lean very heavily to one side in your preference of scalar vs vector, but which side is better varies by program. LLVM is essentially leaning heavily toward scalar in their loop behavior, which is probably best for most C and C++ programs. Many Fortran programs might prefer vector.
My own (Julia) code does. But I have a hard time talking about Julia or Fortran programs in the abstract. I tend to ensure vectorization. Most programmers don't, so even in these languages, they're likely to prefer and benefit from different defaults than I do.