is HFT really measured by assembly op-level performance? I'd assume the algorithms efficiency is a magnitude more crucial than branch-stuff optimization and the likes. better hardware, better parallelism. lots of ground to cover before LIKELY/UNLIKELY macros are likely (lol) to contribute. at least to my educated (but unfamiliar with HFT) experience.