(Compilers obviously do this transformation, including GCC, but it is not always beneficial, especially on x86-64.)
(Compilers obviously do this transformation, including GCC, but it is not always beneficial, especially on x86-64.)
[0] https://clang.llvm.org/docs/LanguageExtensions.html#builtin-...
Also that isn't actually equivalent since `x` needs to be all 1s or all 0s surely? Neither GCC nor Clang use that method, but they do use Zicond.
Indeed you may need to negate `x` if you have only the LSB set in it; hence "3-4 instrs ... depending on the format you have the condition in" in my original message.
I assume gcc & clang just haven't bothered considering the branchless baseline impl, rather than it being particularly bad.
Note that there's another way some RISC-V hardware supports doing branchless conditional stores - a jump over a move instr (or in some cases, even some arithmetic instructions), which they internally convert to a branchless update.