What is the improvement from that change? I thought they were essentially the same.
What is the improvement from that change? I thought they were essentially the same.
bcmp only has to figure out whether the buffers differ. It doesn’t have to figure out where they differ.
For example, for s=1024, bcmp can do (at most) 128 64-bit compares (using vector registers, it could even use larger steps)
memcmp could do the same, with an additional “figure out where in the last 8-byte parts compared the difference lies”.
That’s an amount of work that’s independent of the buffer sizes, so it wouldn’t add much, relatively, _if_ the buffers compared are large and the difference (if any) most of the time is near the end of the buffers. I doubt those ifs often hold, though.
”The optimizer will now convert calls to memcmp into a calls to bcmp in some circumstances. Users who are building freestanding code (not depending on the platform’s libc) without specifying -ffreestanding may need to either pass -fno-builtin-bcmp, or provide a bcmp* function.”*
So, he’s, you may have to provide it yourself. I expect you can fairly easily copy-paste it from various BSD-licensed libraries, though (example: https://github.com/freebsd/freebsd/blob/master/sys/libkern/b..., but be warned about that “I don't believe this is a problem since AFAIK, objects are not protected at smaller than longword boundaries”. That likely is, but may not be true on your platform)
MSVC is even more notorious for this kind of issue, silently adding a number of function calls it assumes the standard library will provide even when your code doesn’t use the standard library at all.
I do wonder why this is a LLVM feature, though, and not a clang one. The system should only make this change if it came from a C(++) compiler and the source included the system header (<string.h> or <cstring>) to get memcmp.
How can LLVM know that the symbol memcmp it sees came from that header, and not from user C code that may do something different, or even from, say, a Cobol compiler?
I don't really know, but I suppose it might be possible to make bcmp faster since you have less strict requirements on the return value.