One of the comments on that article links to this https://llvm.org/bugs/show_bug.cgi?id=1488 which rather suggests that this surprisingly specific optimization was added in order to speed up some calculations in one of the SPEC benchmarks.
(The benchmark in question is code from a chess program called Crafty, which represents a chess position as a bunch of 64-bit bitmaps. Unsurprisingly, 64-bit popcount is quite often useful for this.)