I agree it's a trade-off, it always is :) But if the goal is to write performant code, there's a few more variables to consider.
An important one, IMO, is experience and practice.
Optimization can be found in the weirdest places. My latest experience is not with C or assembly instruction crunching, but with Processing/Java. If I weren't minded to tinker with and optimize my inner loops, I might have never discovered that the java.lang.Math routines are a lot faster than the equivalent Processing functions (functions like abs, max, etc).
Back when I was messing with 80x86 assembly code, I did the same. It's all about having a somewhat accurate model of what goes on inside. With Processing it's the model that the functions are probably wrapped in something. With assembly code it's (among other things) the model that the order of independent instructions matters for purposes of pipelining and preventing AGI-stalls (I dunno if those are still relevant, this was the 386 era).
When you get the hang of it, it's not really hard to experiment, measure, experiment, a few times and end up with spectacularly more performant code.
And really, if the optimized result took a few days to calculate, and an hour of tinkering might save you another day or two, isn't that (often) worth it?
Then there's the fact that I think it's fun :) Getting the code to work properly first, is a whole different process than spot-optimizing certain critical parts. I think it's refreshing to switch between these two different ways of looking at the same code.