The execution of movzx should be negligible. The big difference between the code generated by GCC and the code generated by clang is that the clang code does everything using 8-bit registers, and zero-extends at the end, whereas the GCC code does everything with full 32-bit registers.
I actually had never seen the dil register before this (it's the low 8-bits of the [er]?di register). It's pretty cool that clang uses this when it knows the value in the first argument is byte-sized. I think it's also neat seeing how the compilers use the commutativity of binary-and differently, and end up anding in a different order. Overall I like the clang code better, as it seems closer to the asm someone might write by hand.