Hi, author here. You're perfectly right, microbenchmarks have flaws. Procedures should be run for different data sizes (as I did years ago) and code should be compiled with different compilers. And of course the method described in the text is designed to deal with large data. Using it for counting bits in 64-bit value would be... not wise. :)
AVX2's speedup over POPCNT is not not big, but seems it's such due to my indolence/stupidity. I've just merged pull request by Simon Lindholm (https://github.com/WojciechMula/sse-popcount/pull/2) with manually unrolled loops and it made the AVX2 code faster 40% than POPCNT.