> At that point I figured that the people who told me that "you can't beat the compiler" were probably right and called it a day
On that subject...my understanding [0] is:
These days, the best bet for beating the compiler is to use vendor intrinsics (for SIMD, encryption, bit-twiddling, etc). Shaving an instruction off the inner loop might give you a few percent; using SIMD lets you operate on 256 or 512 bits per instruction instead of 8, 16, 32, or 64. You might be able to show your inner loop is memory-bound (and thus prove further improvements have to come from algorithmic improvements / better cache locality, rather than continuing to fiddle with instructions).
The compiler automatically uses SIMD sometimes, but it can't do so reliably:
* The transformations require things the compiler isn't allowed to do, like increasing alignment of key variables or altering the larger algorithm.
* code that might run on older processor revisions needs multiple implementations selected at runtime. I think gcc has some magic extension ("target_clones"?) to do this relatively easily; otherwise you might need to write your own logic to decide which function pointer to use.
Note that each "vendor intrinsic" matches one assembly instruction, and it's valuable to understand assembly while writing them, but the actual code you check in can end in .cc (C++) or .rs (Rust) or whatever. Doing so means it can be inlined into functions written in the higher-level language, you don't have to encode knowledge about the platform's calling convention into your code, etc.
[0] Not from personal experience. Corrections welcome.