Memcpy vs. memmove
tedunangst.com
tedunangst.com
https://bugzilla.redhat.com/show_bug.cgi?id=638477#c132
I think he's probably right that there's no reason not to alias memcpy to memmove and make the latter smart enough to do the fastest possible thing given its input. Every implementation of memcpy and memmove that I've seen has been sufficiently complicated in order to optimize for big copies that an added comparison or two for bounds checking would seem a drop in the bucket.
But I'm open to practical examples where the performance difference would become significant.
> Replace "hot" bcopy() calls in ether_output() with memcpy(). This tells the compiler that source and destination are not overlapping, allowing for more aggressive optimization, leading to a significant performance improvement on busy firewalls.
int *p, *q, x;
...
memmove(p, q, 50 * sizeof(int));
p[2] = 3;
q[3] = 4;
x = p[2];
In the example above, it is illegal (with only the information available) to optimize the last assignment to x = 3. Strict aliasing does not apply because everything is int.If the call to memmove were a call to memcpy instead, a sufficiently creative compiler author could argue that the optimization of the last statement to x = 3 is legal, because the call to memcpy asserts (among other things) that the int pointed by p+2 does not have any bit in common with the int pointed by q+3.
I don't know how often that would trigger or how useful that would be. But I work on a system in which C programs can be annotated with properties (including non-aliasing properties: see http://blog.frama-c.com/index.php?post/2012/07/25/On-the-red... and http://blog.frama-c.com/index.php?post/2012/07/25/The-restri... ), and that makes me pretty confident that this would be sound and, unlike a joke I made two years ago about the redundancy of restrict in C99, workable.
I think, unfortunately, Linus' argument amounts to demanding that API's conform to how they are used rather than how they are specified, which leads to O_PONIES type problems.
Linus was butting in because a kernel sound driver regression was also suspected earlier...
It turns out not to be quite as dramatic as that description makes it sound: https://en.wikipedia.org/wiki/Direction_flag
You can put bcopy as a function call right before memmove, and then you don't need one function to call the other, which would cause a stack push. Maybe instructions need to be at addresses = 0 mod 16, so that's the "closest" you can get it. And spinning over ~12 NOPs might be faster than incrementing the PC by ~9.
I haven't tested, but I'd bet good money that 12 NOPs would be faster than a jmp.
Smart toolchains will turn those 12 bytes into 2 multi-byte nops, e.g., a 9-byte one and a 3-byte one.
nop word ptr [eax+eax+0] ; 66 0f 1f 84 00 00 00 00 000x66 0x0f 0x1f 0x84 0x00 0x00 0x00 0x00 0x00
That's a size override prefix, followed by the dedicated NOP instruction (0x0f 0x1f), and finally 6 bytes to encode an effective address with offset.
Believe it or not, memcpy() is now the only portable and properly defined way to cast one type to another in C. You are expected to use it with the explicit intent that no bytes actually be copied, simply as a (completely voluntary) offering to the type gods. If they accept your offering, the call will just disappear from your code, and never be made. If you fail to make this offering, they feel entitled to excise an equal quantity of other code to punish you.
Sure, you the programmer happen to know that bits are bits, and that on your system they are already in the register you want to use, but the compiler stopped playing that game years ago. If you wanted to work close to the metal, you should have chosen a language more appropriate for the task. There's a great discussion of the issues in this epic comp.arch thread: http://compgroups.net/comp.arch/if-it-were-easy/2993157
Search for the first occurence of 'memcpy', where you'll find a polite but beleaguered Terje Mathisen asking for the best way to portably cast a float to an integer in C. Then keep searching forward for further occurrences of memcpy as the situation becomes surreal, with a GCC maintainer Mike Stump eventually clearing things up:
>>> So what is the blessed method?
>>
>> Just memcpy it, simple sweet, fast, standard.
>
>So what you are saying is that memcpy() isn't just magic but high magic:
No. What I think I'm saying is that it is standard. See the quoted
text above. The word magic [ checking ] doesn't not occur in c99-tc3.
What is defined in that standard is memcpy, it is as standard as if,
which is also defined.
>It looks like a function call but can be whatever the compiler wants it
>to be, as long as the results behave as if the data was actually copied.
>:-)
No, you fail to grasp the totality of the standard. The implementation
is free to do _anything_ it wants. The only constraint is that the user
can't figure that it deviated from the required semantics by using a standards
defined mechanism for figuring it out. We call this the as-if rule, and its
power is awesome; we could destroy the universe, and repiece it back together
one subatomic particle at a time over a billion years, and still be compliant,
if we wanted.
So there you have it: if you want to reuse the exact contents of a register as a different type in a way that you know will work, the correct approach is to use memcpy() to copy the data from one variable to another, and then hope without confirmation that GCC has optimized out the call to make it equivalent to a simple cast.If you want a better understanding of the GCC mindest, the rest of Mike's posts are excellent: cutting, extremely techically accurate, and (in my opinion) completely missing the point of why some people are unhappy with the direction that C is evolving.
>>> So what about using a union?
>>
>> Most people screw it up, the rules for it working are slightly odd.
>> Those rules are not standard[1], rather they are gcc specific, so in very
>
>Ouch, so what you are saying is that this is one of those (according to
>Nick M) barely defined areas of the language?
What I am saying is that it is defined to not work in the language
standard gcc implements by default. There is no barely, there is no
defined. As an extension to the language standard, gcc implements
(defines) a few things so that users can make some non-standard things
always work. I say slightly odd, as the rules are just a tad harder
than trivial.
I prefer the union approach to memcpy() because I can more easily reason about its behaviour, but veiled warnings from GCC maintainers scare me away from it. But perhaps I misunderstand the warning.For that matter, I prefer simple casts to unions, and while I agree the spec makes it undefined, I don't yet see why all common implementations couldn't simply make it work in all reasonable cases. I'm currently trying to get used to using memcpy() for type annotation, but it still feels unnatural to write a function I don't want executed (yes, I have trouble with setters/getters too).
What I'd like is to have a way to specify to the compiler that I want it to compile the code I give it as written as best as it can, rather than optimizing it out as undefined. As it is, I often resort to inline assembly if I actually want an operation to happen. It seems like there should be an intermediate approach, a hypothetical '#pragma "dwim"' that could avoid this.
It is a bit bizarre that modern CPUs don't have an instruction for something as basic as "copy N bytes as fast as you can", and instead we're in a situation where library and compiler writers have to tune their assembly for different micro-architectures. (I'm not saying that it's a wrong decision, there are probably good reasons for doing things this way, but it's not something you'd expect.)
It's still not quite "as fast as you can"; it's possible to beat rep movsb when buffers are small due to edging effects, but it's about as close as it's possible to come to that today.
Actually, rep movsb/movsw/movsd has been at the top at least since Nehalem (confirmed with benchmarks), and "fast string mode" which does cacheline-sized copies has been around since the P6; you may be able to squeeze out a few % more with SSE (or MMX), but the much larger code of the SSE-based copy functions (especially for alignment) is often not a win overall. Intel only really started advertising that they made it even faster with Ivybridge.
It was the fastest way to do memory copies on the 8086/8088, and might've been the fastest until around the time of the 486 and Pentium when it could be beaten by other techniques; but now it seems that it's coming back in favour.
There's an interesting discussion about this instruction on the Intel forums here: https://software.intel.com/en-us/forums/topic/275765
It did perform competitively for large all-aligned copies, which are what most people tend to look at when they post "memcpy benchmarks", but those turn out to be a relatively small portion of actual usage in most workloads.
Apparently, libc hasn't caught up to those micro-architecture changes yet :/
That used to be the case. As is mentioned, modern processors have fixed the 'rep' prefix to be much faster. Under linux, if your /proc/cpuinfo shows 'rep_good' in the flags section, then your CPU is of this newer class.
Edit: Ohh, I forgot to mention that you will have to update/invalidate any resident cache lines on the CPU also, which requires communicating with the CPU.
I also saw it used to save memory. Onscreen sprites had to be placed into a small memory-mapped address range, so instead of putting every character animation frame into this space and using it all up, I saw just one block designated for the character, and it was animated by DMAing the frames sequentially into that space.
Linux (2.4Ghz Xeon X3430):
./memtest 10000 1000000
memcpy took 0.575571 seconds
memmove took 1.082038 seconds
FreeBSD (2.0Ghz AMD Athlon64 3000):
./memtest 10000 1000000
memcpy took 1.487334 seconds
memmove took 1.442741 seconds
https://github.com/UK-MAC/WMTools/blob/master/tests/memtest....
Additional findings: Unsurprisingly using clang instead of gcc does not affect performance on FreeBSD or Linux.