A cache miss is not a cache miss
larshagencpp.github.io
larshagencpp.github.io
When data isn't in a register and isn't in the L1 cache, it takes a lot of time to fetch it from other caches or about 200 clock cycles if it comes all the way from main memory. We measure that event as a cache miss. But modern x86 processors will go to great lengths to execute the rest of the program while waiting for the data to arrive. A cache miss only really slows the program down if there aren't enough nearby instructions that can be executed while waiting for the missing value to arrive.
You could likely write a program that triggers a cache miss every 30 clock cycles but runs at the same speed as a program without cache misses. In a different program, a cache miss every 30 clock cycles can mean a slowdown by two orders of magnitude. Cache misses are only a useful metric to give us an idea where to look, not to show actual problems.
What is your suggested alternative metric?
As long as we don't have that, cache misses are a useful metric on their own, as long as one is aware of its caveats. As most things, cache misses come in various shades of grey.
EvilOfCacheMiss = TimeOfMemoryFetchCycles - CyclesSpendDoingInstructions /TimeOfMemoryFetchCycles
But yeah, sounds like a useful metric to me.
In particular, see Equation 4: http://researcher.ibm.com/files/us-viji/miss-cluster.pdf
...unless they're instruction cache misses, which can literally cause the CPU to run out of instructions to execute.
never realized 'or' is a keyword... :)
An aside, because instructions are accessed sequentially (branches aside), I would intuit that data misses are much more common than instruction misses. If this is incorrect I'd be interested in an explanation or further reading.
Without more details on data layout, you could be seeing the side-effect of aggressive pointer prefetching, which may only be possible with one layout but not another. Or there could be a fixed stride for some of the data allowing prefetchers to kick in. It's hard to tell and deserves some more experiments to isolate what is going on.
#include "stdio.h"
#include <utility>
printf("%zu\n", sizeof(std::pair<int, int>));
8
So it's the obvious layout you would see in C with an array of struct { int x; int y }.The cling interpreter is good for this kind of thing.
The library provides a template for heterogeneous pairs of values. The library also provides a matching
function template to simplify their construction.
template <class T1, class T2>
struct pair {
typedef T1 first_type;
typedef T2 second_type;
T1 first;
T2 second;
pair();
pair(const T1& x, const T2& y);
template<class U, class V> pair(const pair<U, V> &p);
};
pair();
tl;dr: It looks like the 03 standard requires it to be equivalent to the plant C version in terms of layout, although that could involve padding between first and second if e.g. T1 = short and T2 = int, just like in the equivalent C code.One nit: Please, please, stop perpetuating the use of the non-word "performant." I cringe every time I hear it (now even in person at work!) Using it just makes you sound dumb, and clearly the author is not dumb.
In french. What do you have against french?
You cringe at the use of a word that is now common in our vernacular? The fact that it's not listed in most dictionaries does not mean the word is not valid. Dictionaries aren't written by all-knowing gods who dictate our languages to us. They are sourced by people, and we frequently opt to add new words to the English language. The more we see people continuing to make use of the word "performant" to mean what we all mean when we use it, the more likely it becomes that we will simply officially adopt it into our dictionaries.
If you understand the meaning behind a "non-word", then the word is already a part of the language. Language is fluid, and the appearance of new words is normal. You can grip onto your dictionary like a bible, while looking down your nose at others who are flexible when change comes knocking. It doesn't make us look dumb; it just makes you appear pretentious.
Aside: "performant" is borrowed from French, which also has "efficace" as a direct translation for "efficient". So there is room for an additional word. ;)
Seriously it makes me cringe every time someone says it.
SOA == Structure of Arrays
A flag that guarantees side effect freedom to a set of operations and suboperations, for the processor that would be great.