I haven't benchmarked it, but surprisingly that seems to be enough to get the gcc assembly almost identical to what clang is producing https://gcc.godbolt.org/z/hfcYdP5rv
Throw it in godbolt and see if the assembly is different. I would doubt not.
The optimizer will do this anyway.
The whole point of the article is that the optimizer in GCC did not do this.