Please tell me you forgot the /s.
I have a PhD in computer engineering from a top-20 school in US. Took a bunch of grad level classes, passed the quals (my specialty was ML accelerators).
I do NOT have an “intimate understanding of CPUs”. I probably know a little bit more about CPUs than an average programmer. Which is very little.
Modern CPUs are extremely complex. Almost as much of impenetrable black boxes as modern neural networks.
and 101 other hilarious jokes you can tell yourself!"
In my school to pass the computer architecture course you had to read and present a recent paper on CPU design.
I'm very confident in people having the ability to read by the time they are in college. And considering that summarizing a research paper doesn't even have to be perfect, plus very little need to scrutinize the experiment itself, undergrads should be able to do that.
Otherwise they don't belong in college.
I did an undergrad in CS, where I did well. I don't feel like I understand CPUs very well. Certainly not anywhere in the realm of "intimate."
I think the descriptor “intimate” is overstating the case if you don’t know how to optimize code on different implementations of the same ISA. Most formal CS programs give you a generic understanding of CPUs, more like a survey course, not enough information to do serious optimization.
Compilers are good at finding very local optimizations targeting specific microarchitectures. If you use godbolt and switch out the architecture flags you can see how the codegen changes.
However, as with higher level code, the compiler can’t rewrite your data structures or see patterns spread across too much code. As a simple example, modern cores have 4+ concurrent ALUs with different capabilities. Some of those ALUs are usually idle if there are not enough independent instructions or independent instructions are too far away from the current instruction. You can see large performance improvements simply by reorganizing your C code so that the compiler and CPU can see more opportunities to use more ALUs in parallel. Interestingly, a lot of these code changes are trivial no-ops at the code semantics level, but they bring the ALU concurrency opportunity within view of the CPU.
There is an active niche community on the Internet that studies how various instruction sequences interact with various microarchitectures. This is probably the best resource because a lot of detail is not well documented by the CPU companies themselves. It requires a fair amount of experimentation to develop an intuition for how code will run on a given CPU at this level of detail.
Could elaborate on that? I mean where do I find this community?
If your day to day involves "intimate knowledge with the CPU" and bumping program counters, I can rule out a lot of things that you probably aren't building, like a compelling iOS app or forum HTTP server, for example.
Maybe you're doing impressive work on emulators or something though. But it's nothing to get pompous about just because other people don't share that interest.
We too easily go off careening into a circlejerk.