My preferred approach when optimising is to go for size first; then if more speed is desired, carefully apply size-increasing optimisations to areas which may benefit, testing with a macro benchmark to judge any improvement.
Here's some very interesting discussion about it: https://stackoverflow.com/questions/43343231/enhanced-rep-mo...
There's also the fact that `rep` startup cost was higher in the past than it is now. I think it started to get really fast around Ivy Bridge.
This is all discussed in Intel's optimization manual, by the way.
https://software.intel.com/content/www/us/en/develop/downloa...
And if you want ALL THE THINGS:
https://software.intel.com/content/www/us/en/develop/article...
I remember it was actually quite common in early PC programs, but I don't remember well enough whether it was also generated by compilers, or just existed in hand-optimized assembly (which was of course extremely common back then).
The GCC version is just bananas. https://sourceware.org/git/?p=glibc.git;a=blob_plain;f=sysde...
Compare the newer ERMS implementation: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/x86...
Edit: it's rolled into the erms version.
Do you have a source for this? Seems surprising GCC would leave such low-hanging fruit. G++ makes the effort to reduce std::copy to a memmove call when it can, or at least some of the time when it can (or at least, it did so in 2011). [0]
Related to this: does GCC treat memcpy differently when it can determine at compile-time that it's just a small copy?
A programmer should be careful about second guessing the compiler. And a compiler should be careful about second guessing the processor.
It's a performance-sensitive standard-library function, the kind of thing that deserves optimisation in assembly. It's also the kind of problem that can be accelerated with SIMD, but that necessarily means more complex code. That's why the standard library implementations aren't always dead simple.
Here's a pretty in-depth discussion [0]. They discuss CPU throttling, caches, and being memory-bound.
The trick is to optimize the right macro benchmark- one that matches your customer's key use case I suppose.
Agreed, in my experience, more often than not, size is speed. Small memory usage mean less cache misses, smaller memcpys, more chance of having data fit a single word, etc...
If you don't know what you are doing unrolling is just as likely to hurt performance because you don't fit into uOP cache and get less decode bandwidth as a result. Or you increase ICache pressure on macro benchmarks and hurt real world performance. Modern cores are really good at hiding loop accounting overhead.
But there are plenty of architectures out there where space-time tradeoffs (by optimized compilers) are still a thing (my PoV is from the embedded industry).
I watched a great talk about this, and how to avoid problems that result from it. The presenter built a framework that made random changes to alignment of various blocks of code with two different versions of the same microbenchmark until it had high statistical confidence that either one was faster than the other, or it would eject and say that the difference between the two was statistically insignificant. Or something along those lines.
I'm looking for the video now, but I can't find it. Pretty sure it was CppCon but I'm not positive.
With that said, it sounds like at least with GCC, a global constant would go in the TEXT section of the binary and I can't think of why that would affect performance, especially since the variable was seemingly unused. Neat, I hope someone finds the thread to pull here :)
Wouldn't it go into the .rodata section, not .text?
I don't know why you think it sounds like that, but GCC (on Linux) puts global constants into the .rodata section (read-only data) and other global variables into the .data section.
> I can't think of why that would affect performance, especially since the variable was seemingly unused.
I'm guessing data alignment and corresponding changes in cache hits and misses. But I wonder if there's a scenario where this change changes the offsets of other variables from the base of the data section, which changes the encodings and therefore the sizes of the instructions accessing those variables, which overall changes code size and might cause code alignment issues. Probably too far-fetched.
You can run the program I built here to see for your self: https://github.com/vivekseth/blog-posts/tree/master/Jump-Add...
Since the string in the TEXT section, we can actually execute it as if it were code!
After you build the program you can run `otool -t ./a.out` to verify that the string `execString` is indeed in the TEXT section.
Here's the output as-is: https://gist.github.com/vivekseth/20f319d2a9978af57d926b649a...
Here's the output with a null byte in the middle: https://gist.github.com/vivekseth/fc50319aaac24588bcf568209b...
From what I can tell, it seems like both strings are in the TEXT section now. Maybe something changed, or I'm remembering incorrectly.
I guess I don't understand enough about how this variable sitting in memory might affect hit/miss/alignment especially if it's not used in the program as it's running.
Would it be that some variable that is used is pushed off of a page in memory or something by the unused variable, so access it is slower? I guess that makes sense.
I guess, maybe once another global constant is introduced somewhere else, this might have exactly the same effect.
But also, once there are more unrelated changes, everything will change anyway again.
It seems strange then to keep this in because right now this seems to help.
What is "mystery" in the article is actually bread and butter of anybody who tasked to seriously optimize a piece of non-trivial code.
https://people.cs.umass.edu/~emery/pubs/stabilizer-asplos13....