Don't let the small clean assembly for this simple function fool you, writing optimized assembly on modern architectures is getting harder and harder, you have to take the cache into account for instance (both instruction and data) as well the particular implementation of the instruction on the CPU (microcode etc...). Fewer instructions does not always mean faster code, maybe TFA should have been more explicit about that instead of just saying it's "a crude measure of code quality". Look at the assembly generated by a compiler for a modern X86-64 architecture, you will see stuff that seem to make no sense such as NOPs in the middle of functions.
Also, in general, it's interesting to know what the compiler can and cannot optimize to decide when it's interesting to handcraft some assembly. It's always better to benchmark first instead of doing some premature optimization.
For instance, C does not have a bitwise rotation operator (while many architectures support a rotate instruction) but gcc easily recognizes rotation patterns and optimizes them without trouble.
>Don't let the small clean assembly for this simple function fool you, writing optimized assembly on modern architectures is getting harder and harder
Maybe in the general case, but this is literally a single instruction over unchecked arithmetic. You just add and then branch on overflow. Kinda hard to screw that up.
It's really sad that people are so afraid of assembly these days that they can't even bother to write, say, a couple lines of assembly for their language VM's overflow-checking add instruction, and rely on hacks like this instead. Who cares if pure C is more portable if the assembly is shorter and it takes two minutes at worst to look up what the opcode for "branch on overflow" is for any other platform you want to port to? At some point, you're not playing it safe or even saving much time, you're just wasting CPU for the sake of laziness.
So at that point it's probably valuable to see if you can't get the compiler to generate the correct code by itself in a portable way.
I'm not afraid of writing assembly when I have to but I always consider it a last resort scenario when I really can't get the same result with some good old C.
Better yet, imagine that the compiler has performed value range propagation [1] to determine that the range checks, after inlining, were redundant or simply elidable. This would be infeasible if it were written in assembly.
[1] http://llvm.org/devmtg/2007-05/05-Lewycky-Predsimplify.pdf
Better yet, imagine that the compiler has performed value range propagation [...] This would be infeasible if it were written in assembly.
VRP doesn't depend on the instruction set; it could be LLVM IR or it could be x86 or ARM or 6502 or anything else. It's just looking at values and operations on them.
To the compiler's own IR, whatever that is.