The most useful GCC options and extensions
antoarts.com
antoarts.com
Like this:
typedef float vec4 __attribute__((vector_size(16)));
vec4 a = { 1, 2, 3, 4 }, b = { 5, 6, 7, 8 };
vec4 c = a + b;
What GCC does not allow (IIRC) is GLSL/OpenCL -style shuffle syntax like position.xyzw or (vec4)(a.xx, b.yy). Clang has __builtin_shufflevector but GCC doesn't have anything like it and you have to revert to SIMD intrinsics for SSE, NEON, etc.Interesting thing about shuffles: Clang's implementation of ARM NEON intrinsics actually defines NEON shuffle intrinsic functions using a __builtin_shufflevector. I've noticed that Clang emits a lot better NEON code than GCC does. GCC's NEON backend is a lot worse than the SSE backend. And shuffle instructions are a major difference between NEON and SSE.
If you are not working on high-performance bit-twiddling algorithms then you will probably have limited use for shuffle intrinsics. For those applications, it can save a few clock cycles for each call relative to more naive methods.
Basic example, doing addition of 4 values in 2 ALU operations:
vec4 sum(vec4 v) // return v.x+v.y+v.z+v.w repeated 4 times
{
vec4 temp = v + v.yxzw;
return temp + temp.zwxy; // did I get this right?
}
Practical examples: https://github.com/rikusalminen/threedee-simd (work in progress)
Requires this: http://gruntthepeon.free.fr/ssemath/Other nice extensions `x ? : y` => `x ? x : y`, vector extensions, C++ support for `__restrict__`, all those attributes (deprecated, format, warning, pure, etc.)
Regarding Assembler the `-Wa,-ahl` flag to get the C code interleaved.
And regarding warnings (especially for C++) the world doesn't stop with -Wall -Wextra. There are a bunch of nice warning options like `-Weffc++ -Wfloat-equal -Wdouble-promotion` and so on.
The idea is that in these days, cache misses are king, so smaller code means fewer misses.
Of course, you need a real-world load, test on correct CPU architecture for your problem etc.