Strlen(buf) – Not as simple as you’d think
medium.com
medium.com
And IIRC repe cmpsb isn't bad for longer strings on new x86. It's the initial setup that adds overhead. Still, you want to make a fast general solution.
http://llvm.org/docs/doxygen/html/SimplifyLibCalls_8cpp_sour...
Also you need to calculate size from the byte in the word that is zero, not the word itself.
It would be really complicated for a compiler or human programmer to work around the quirks of the Z80 instruction set, where only certain registers can be used for certain operations. Even an optimal implementation of the word solution would probably be quite a bit slower than a straightforward naive implementation.
The Z80 instruction set only supports word reads from constant addresses (but bytes can be read from an address specified in any register pair). The Z80 has very few registers, so the more complicated algorithm will probably face register pressure. Also, word comparison is only implemented for the HL, DE register pair unless you want to use an index register which requires an instruction prefix, which will make your code even slower.
But memory reads from an address in a register pair other than HL or an index register can only go to the accumulator register A, thus you'd need HL for both the comparison and the address. So you need more instructions to save and restore HL. (You can't even use the fastest option EX DE,HL because DE is also needed for the comparison.)