Cold, Hard Cache – Insomniac Games’ Cache Simulator [slides]
deplinenoise.files.wordpress.com
deplinenoise.files.wordpress.com
It then integrates with the IDE to give you mouse-over reporting of a single statement's cache hit performance.
Also, the submission has been edited. I mentioned that the system is open source, http://github.com/insomniacgames/ig-cachesim is on the page I submitted, but is on page 105/105 of the slide PDF.
Edit, disclosure: I once had the pleasure of working with deplinenoise, and tried to sponge up as much knowledge as possible. He's great.
On the other hand: It probably seems particularly impressive because its building blocks don't happen to be in your skill set. A rule of thumb: Something is impressive in proportion to how many times you think "How did they do that?!".
In this case, if you're not already familiar with disassembly, binary instrumentation, how caches work, etc, then every step along this way sounds magical. But to a low-level/systems programmer, this sort of thing is their bread and butter. So their reaction wouldn't be "OMG how did he even...?" - it's probably more like "Niiiiiiice".
(If you are, eg, a web developer, you probably have an equivalently ridiculous level of background knowledge about the internals of 5+ levels of the web stack, and how each goes wrong. I have seen systems guys - people who would jump on this project in a heartbeat - take one look at modern web dev and flee with their head in their hands, asking "how do they know all that stuff?!")
https://deplinenoise.files.wordpress.com/2017/03/webtoolspos...
My experience has been that people who are good at the bottom layers of the stack are good with the top layers of the stack, too. And if they're unfamiliar, they can ramp very up quickly.
My comment addresses the perception that "OMG programmers today know nothing; this is what a Real Man looks like". Someone learning web development has spent a lot of brain-space on learning a level of in-depth knowledge that also looks magical to someone who hadn't.
As it happens, I too think that the things you learn lower down the stack tend to make you a better engineer, whereas what you learn higher up the stack is too often an unedifying schlep. High-level systems spend a lot of complexity solving problems created by the layer underneath. (This goes 10x for web frameworks.) Low-level systems are more tightly constrained by the boundaries of the possible, so they spend their complexity on more fundamental problems. The same amount of time and intelligence spent learning the ins and outs of Angular yields less transferable skills than learning how compilers work. But if you need to build a web startup, compiler expertise on its own won't help you.
It's like learning physics vs biology. Sure, physics is more fundamental, and a physicist learning biology usually has an easier time than the other way round. But a research biologist has also spent their time acquiring an immense amount of expertise, and fundamentally we need medical advances more directly than we need confirmation of the Higgs boson.
Completely serious question/request: I'd be hugely interested in psychology hacks/ideas that short-circuit the brain (safely) to remedy this "thinking anti-pattern".
I think self-worth and confidence might be buried in the answer - I don't know if this is my own bias (depression?), I can't help but think that there's the tiniest edge of "I'm not worth learning this", which snowballs in a big way. Lack of confidence would be an obvious blocker though.
I'm not saying that programmers are getting worse--actually wait, no, that's what I'm saying (myself included).
You work hard when your job is your hobby, passion, and profession.
For our project, we only needed to simulate cache accesses and not the content. We also kept track of the LRU block, simulated sub-blocking, and added a FIFO victim cache to the simulation. It was a fun exercise, and I got to learn a bit more C++ in the process.
The next project involves building a CPU pipeline simulator, which might be more challenging.
Regardless, very nice to see tools like this being open sourced. The games industry has a lot of valuable experience with performance-critical code and I'd love to see more of it shared publicly.
Trying to add in a prefetch handler would be painful, time consuming, and may actually hide some cache misses that you would want to find. Also each CPU family and potentially in between families there could be changes to the prefetch algorithm used.
Certainly an important point, but only relevant if what you are doing had a hope of being prefetched in the first place. Maybe the tool could be extended to give an idea of how 'prefetchable' something is? This should be simpler than a full prefetcher simulation. Then you can concentrate on the cache misses it predicts that are unlikely to prefetched. The tool is more focussed on pointing out areas you can improve, it doesn't matter if it gets it wrong sometimes.
Generally, I think considering the prefetcher is roughly as important (and, indeed, inextricably linked) as considering the cache subsystem itself. Any performance sensitive code will benefit from careful consideration of memory access and traversal patterns.
Fun fact: A certain game-dev kit had another machine sitting along-side on the bus. It enabled you to sample the Program Counter at N=1 for a small amount of samples(I want to say ~4k) off of a breakpoint.
The amount of perf stuff you could trivially track down was staggering. Load-Store-Fetch? Boom, there it is. Cache misses? Clear as day.
Never worked on another platform with such an incredible profiler and still miss it to this day. You could set N to any power of 2 for coarser profiles but it always shined at N=1.
You can use the CALLGRIND_START_INSTRUMENTATION and CALLGRIND_STOP_INSTRUMENTATION to start and stop the stats collecting. Start valgrind with the –instr-atstart=no option to disable collecting on everything outside of those blocks.
It's still really, really, slow but much faster than when it collects everything.
Some bugs are going to cause problems regardless of the cache config (like the "access a cold page unnecessarily every iteration of the loop" examples); some might be more sensitive to exact cache configuration.
https://en.wikipedia.org/wiki/Jaguar_%28microarchitecture%29
I am sure the next steps will be better visualization and simulator accuracy.
There's nothing statistical about it, no "samples" that are hoped to represent some hidden full system, since the entire instruction stream is analyzed when the simulator is enabled.