Edit: Yeah, looking at the godbolt link they provided, at least the execute_loop_basic function does essentially no work.
Edit: Yeah, looking at the godbolt link they provided, at least the execute_loop_basic function does essentially no work.
The question posed early in the blog was whether Duff's Device is still relevant in 2021. The fact that the compiler can completely elide the loop without it but cannot with it is a pretty good indicator that it's likely to cause more harm than good.
It might prevent inlining of a function call in the loop body, which would make a loop much, much slower to execute. It will make the loop body larger (even without inlining) and likely cause an increase in cache misses. It will be unpredictable based on build: an unrelated change somewhere else might misalign it with a cache line. Different CPUs might have different cache behaviours and perform differently.
The right view on this sort of thing hasn't changed for decades: don't optimize prematurely. Trust your compiler, check it with profiling, look first for algorithmic improvements because you can gain FAR more converting an O(n^2) algorithm to an O(nlogn) than you can by unrolling a loop. Duff's Device is obsolete in all but a vanishingly tiny number of edge cases.
The compiler removing the loop is due to a bad test setup in which the function in the loop does nothing and compiler can figure that out and skip the whole calculation. In a real scenario the function in the loop would (presumably) actually do some work so it could not be elided. Also, note that the compiler does this for BOTH the Duff's device version and the non Duff's device version.
Also, I'm neither attacking nor defending Duff's device. I'm just saying that we can't draw any useful conclusions from this article because of poor methodology.
For trivial loop, Both GCC and clang optimize away the entire BasicLoop, but clang failed to optimize away the entire DuffDevice, while GCC did. Thus, the only thing doing the work is clang's Duff Device.
But even then, both GCC and clang optimize the BasicLoop function itself into just single add operation (data += loop_size) while it actually loops on the duff device function (though as mentioned, these functions weren't called from the benchmark)
We have long known that in general compilers are better at optimization than humans so you should write readable code until a profiler shows otherwise. However there are cases where the compiler doesn't make an optimization. There are sometimes corner cases that the compiler can't figure out doesn't apply to you, so manually doing the optimization might work. These cases need to be re-evaluated with every CPU change, and every compiler upgrade though.
Nothing above says if duff's device applies to not though. Duff's device is very likely to hit a corner cases where the compiler isn't sure if it can safely apply optimizations so it won't. As such I wouldn't be surprised if it is still relevant. (though mostly on embedded systems where we don't have nearly as fast of CPUs as more common computers)