My most recent misadventure: trying to write hand-coded ARM Neon assembler, and/or C++ code with neon intrinsics to optimize a piece of real-time audio effect code. The clear performance winner: plain C++ with no intrinsics, but tweaked to allow auto-vectorization (plus judicious use of the __restrict modifier for a small but significant boost). GCC produced code that had better instruction scheduling than I could (but not for the NEON intrinsics, oddly). And as an added bonus, the same plain-old C++ code generates AVX vectorization on MSVC without modifications! (MSVC also supports __restrict until the C++ standards committee gets their act together to adopt the eminently necessary C99 restrict keyword).
That is wrong. Just because you didn’t succeed in writing faster code in one case, doesn’t mean it’s impossible. See e.g. [1] on why the Lua vm is written in assembly. It’s from 2011 but not much has changed in the meantime.