My most recent misadventure: trying to write hand-coded ARM Neon assembler, and/or C++ code with neon intrinsics to optimize a piece of real-time audio effect code. The clear performance winner: plain C++ with no intrinsics, but tweaked to allow auto-vectorization (plus judicious use of the __restrict modifier for a small but significant boost). GCC produced code that had better instruction scheduling than I could (but not for the NEON intrinsics, oddly). And as an added bonus, the same plain-old C++ code generates AVX vectorization on MSVC without modifications! (MSVC also supports __restrict until the C++ standards committee gets their act together to adopt the eminently necessary C99 restrict keyword).