It should be kept in mind that these are microbenchmarks; I'd guess the much bigger AVX2 instruction sequence will have visible cache effects in macrobenchmarks (i.e. combined with a mix of other instructions) and lose its small lead --- on Haswell AVX2 is ~17% faster, on Skylake only 6%.
POPCNT is a single instruction and should definitely be inlined; I doubt that would be such a good idea for these longer sequences, but then putting it in a function means the call+return overhead could also become significant. It's hard to quantify without doing actual measurements, but with the figures given I'd be inclined to stick with POPCNT.