Now any load from the L3 cache memory or from the main memory takes much more time than any other instruction (not counting exceptions generated by instructions, which include many memory accesses that slow them down, or deprecated instructions that are kept for backwards compatibility and that are executed by long microcode sequences).
Paste in assembly code, check "Trace Table" and run, then "Open Trace". Not sure if it will help with your annoying colleague, but it gives a much more concrete idea about how a processor will execute any given code.
Or, if you want to channel their energy into something slightly more direct, there's also https://quick-bench.com/ which allows easy micro-benchmarking. Still not guaranteed to be relevant to any real-world scenario, but more data-driven than "vibes".
> I have a colleague who has really no idea what he's talking about with respect to machine performance, and who did not have the requisite knowledge of how to peep at the assembly code of a given function with the standard tools like objdump, who now loves to send everyone godbolt links in slack, along with his suppositions about which function will be faster, based entirely on vibes (mostly, instruction count).
This, however, just no.
Instruction counts are only useful if everything is guaranteed to be in registers.
It was a completely unnecessary instruction from a correctness perspective, because it had no effect on the answer. However, it was important for performance; removing the instruction made the calculation slower.
It would be fun to non-destructively randomize the instruction stream and have an ML model learn how to remove hazards.
I do totally get how some people learn just enough to be annoying. Generally I still think that's not a good reason to gatekeep them.
Write program -> compile -> disassemble w/ some mapping -> make notes -> repeat.
Eventually your brains pattern recognition starts to allow you to do neat things with disassembles of programs without source code.