That's not true. If you write the byte swap in ANSI C using the gigantic mask+shift expression it'll optimize down to the bswap instruction under both GCC and Clang, as the blog post points out.
Assuming the macros or your giant expression are correct. But you might as well use the compiler intrinsics which you know are both correct and the most efficient possible, and get on with your life.
Sorry I'd rather place my faith in arithmetic rather than someone's API provided the compiler is smart enough to understand the arithmetic and optimize accordingly.
"Someone" here is the same compiler you're trusting to optimize your giant arithmetic expression of the same idea. Your statement is internally inconsistent.
There is a value to keeping it completely clear in your head the difference between a value with arithmetic semantics vs a value with octets in a stream semantics. That thinking will work in all contexts, while the compiler knowledge is limited. The thinking will help you write correct ways to encode data in the URL or into a file being uploaded that your code generates for discord or whatever, in Python, without knowledge of the true endianness of the system the code is running on.