Benchmarking is the only thing you can really do in some cases.
Benchmarking is the only thing you can really do in some cases.
And once you're done, you can do another round with tools which can do more detailed analysis like https://uica.uops.info
I do this when writing performance-sensitive C# all the time (it's very easy to quickly get compiler disassembly) and it saves me a lot of time.
I think the complete opposite! Benchmarking is very difficult to get right and in some situations you couldn't really make a benchmark to test what you're interested in improving. Reading assembly can be objective if you know or look up the instruction latency/throughput of each instruction (or you have a tool provide it) and you can also use loop throughput analyzers (as the other commenter mentioned) that will try to predict the typical throughput for a given loop.
IMO if you get used to looking at assembly it becomes obvious in the majority of cases whether there's performance left on the table.
And "simplicity" does not equal "faster". Although for cold code, reducing its impact on i$ and d$ of the rest of the system is probably smart, so sometimes speed is not the only factor.
If only metric you look for is number of instructions then yes and this is a wrong way of looking at this type of problems. One can write a 50 LoC heavily optimized SIMD code that outperforms the compiler generated 10 LoC by 5x.