What makes this not real AVX? Because there is not one dedicated 256-bit unit?
What makes this not real AVX? Because there is not one dedicated 256-bit unit?
Except literally running at half capacity. That doesn't qualify as a "penalty" to you?
The point was that AMD advertised support for AVX in a context where you'd expect it to be performance-comparable, and what they shipped was "support" for AVX in the sense that the code wouldn't crash, but wouldn't provide any performance benefit over SSE either.
Intel reduces clockspeeds when running AVX2 and drastically decreases clockspeeds when running AVX512. AMD's slightly smaller unit doesn't need to slow down, so the actual performance difference is smaller than it would appear.
Changing frequency takes time. Also, if you are putting through an AVX instruction and a few integer instructions at the same time, the integer instructions will downclock the whole time the AVX pipeline is in use (plus the time before/and after while the clockspeed is being adjusted) decreasing performance for more than just the AVX.
There are a couple articles written on the topic. Here's one.
https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
https://software.intel.com/en-us/forums/intel-isa-extensions...
I’ve been writing SIMD code for some time now, and I disagree.
You can read the documentation I’ve generated https://github.com/Const-me/IntelIntrinsics AVX intrinsics begin with _mm256_.
You’ll find that for some AVX operations, such as _mm256_fmadd_ps, _mm256_adds_epi8, _mm256_blendv_ps, both latency and throughput of Ryzen is equivalent to Skylake. For some others, e.g. _mm256_add_ps, _mm256_mul_ps, Ryzen is close to some older Intel, Haswell or Broadwell. Only for very rarely used stuff, e.g. _mm256_broadcast_ps, _mm256_madd_epi16, Ryzen is significantly slower than Intel.
My understanding is that without any penalty was not true, but it is possible the new gen cpus have changed that, I have not firsthand benchmarked it.