An Interview with Zen Chief Architect Mike Clark
computerenhance.com
computerenhance.com
Is that really true? I was under the impression that ARM core density was better than x86.
The first Google result I found agrees: https://www.bitsnbites.eu/cisc-vs-risc-code-density/
> It may take a little bit more microarchitectural work for us to account for stronger memory order on the x86 side, but of course the software has to account for the weaker memory ordering on the ARM side. So there are tradeoffs.
I don't think this makes sense either since software has to have synchronisation anyway. I don't think much software relies on x86 TSO (intentionally anyway; I'm sure many bugs are covered up by it).
Anyway fantastic interview! Very refreshing when they're actually technical with interesting questions and answers.
This isn't a great metric for 2 reasons. Firstly, it's looking at static rather than dynamic instruction counts, and 2nd x86 tends to get bigger vectorized implimentations, but these are often on bigger vectors so the "instructions per data" which is the more important metric in many cases is lower.
https://news.ycombinator.com/item?id=19329128 has a pretty interesting discussion on this from a few years ago, but 2 places where x86 shines are the lea and mov instructions. They are both incredibly compact for representing potentially very complicated operations.
There's talk about vectors, longer basic blocks and cache utilization and how they wish programmers used more of it, it misses the real world.
Regardless of how "hardcore" programmers feel about it, so much real world executed code is built in JS,etc. The JIT:ed code is super-branchy (to cater for deopt fallbacks) and won't use vectors at all.
This is something Apple has gotten correct with their vertical integration as they seem to have put more focus on making real-world code go fast. Even most games will have huge swaths of non-vectorized code that would benefit from a more scalar focused way of optimizing.
Considering transistor counts in use today, instead of bigger vector units,etc it could be spent on bigger UOP buffers, BTB buffers, bigger caches to eat up branchy-but-linear code flows that are reality for less "optimal" languages).
After some levels of abstraction, it really doesn't matter how fast the processor is, it will run like crap anyway.
> Why isn't there a conditional exception? to replace if (!cond) { __builtin_trap(); }
There's not really any way to make that into something that doesn't branch; the best you can hope for is only one instruction that may branch but hopefully gets predicted accurately. But in the event of an exception, there really does need to be something that can cause the instruction pointer to do something other than advance to the next byte after the current instruction.