The only thing you can do is run your code and get an empirical answer to the question 'Is it fast enough'.
The only thing you can do is run your code and get an empirical answer to the question 'Is it fast enough'.
(I too execute the code in my head, but these days usually a couple layers above x86 ASM, in my own mental "bytecode" that tracks how expensive are some of programming language's operations and stdlib functions.)
An example - some code I was working with gave subtly different results on two processors. It turned out this was because one of them implemented FMA - Fused Multiply Add.
https://en.wikipedia.org/wiki/Multiply%E2%80%93accumulate_op...
It wasn't particularly painful to find this, but if you were attempting to debug this kind of problem without considering lower level issues, you'd spend a whole lot of time banging your head against a wall.
I think if you're talking micro-code fusing then Intel does guarantee it has exactly the same semantics as a separate multiply and add.
Interesting fact - I believe modern Intel architectures actually only have fused multiply add. If you do just a multiply it'll do a fused multiply and add zero.
Most common instructions in modern CPU cores are decoded directly into the corresponding micro-operation(s), without involving the microcode ROM. Only the most complex instructions, involving many micro-operations, will use the microcode decoder. Also, there are often several copies of the simpler instruction decoder, so the CPU can decode several instructions in a single cycle, while there's normally a single copy of the complex (microcode-using) instruction decoder.
A good resource if you want to read more is part 3 of https://www.agner.org/optimize/ which describes the decoder (and other parts) of several families of x86 processors.