I never touched AVX or SSE in the past, so this was a great learning experience. In 30 minutes you can get 90% of it, but I think that to really do great stuff you need to also understand the relative cost of every AVX operation. There is an incredible user at Stack Overflow that replies to most AVX / SSE questions, if you check the AVX / SSE tag you'll find it easily.
However I noticed that when there were many load/store operations to do, there was no particular gain. See for example this code:
#ifdef USE_AVX
__m256 es = _mm256_set1_ps(error_signal);
int psteps = prevunits/8;
for (int x = 0; x < psteps; x++) {
__m256 outputs = _mm256_loadu_ps(o);
__m256 gradients = _mm256_mul_ps(es,outputs);
_mm256_storeu_ps(g,gradients);
o += 8;
g += 8;
}
k += 8*psteps;
#endif
What I do here is to calculate the gradient after I computed the error signal (error * derivative of the activation function). The code is equivalent to: for (; k < prevunits; k++) *g++ = error_signal*(*o++);
Any hint about exploiting AVX at its max in this use case? Thanks. Ok probably this was more a thing for Stack Overflow, but too late, hitting enter.