Some of the things you claim are very context-specific and not true in general, which I find jars a bit with the (perceived by me to be) authoritative tone of your post.
> I love these "rules of thumb", "black magic sorcery", and "best practices" that grey-haired old wizards pick up.
Just quoting this to point out what I mean by addressing your tone. It's fine to disagree, but this feels like you're setting up that everyone who disagrees with you must be condescending and stuck in the past. Yes, performance advice is often tied to the specific situation, but that's also true for the things you state.
> Worse still, these internalised tricks and rules are often related to Moore's law scaling, making them exponentially wrong over time. Not in the figurative sense, but a literal one.
Implying that Moore's Law makes anything exponentially faster is an anachronism (talk about old knowledge!), it only makes things exponentially more parallel nowadays. See next point why that's relevant.
As others have already clarified, -fomit-frame-pointer was sometimes 20% on x86. For some applications, people would kill for a 20% boost in single-threaded performance, even today!
And importantly, that this specific optimization has become useless on x64 says nothing about the value of optimizations in general. Different ones are still useful, to give an example let's say SIMD (where it applies).
> You see this with people claiming that I'm exaggerating when I say all code should be parallel. Meanwhile AMD is about to sell a 128 core / 256 thread processor. You can have two per host for a staggering 512 hardware threads. Single threaded code is capable of utilising just 0.2% of that machine!
I'm not sure which strawman "grey-haired" developer you have in mind that buys a 512-thread server to run single-threaded code on it and then needs you to explain to them that they should parallelize it. Maybe you've met that person. But in general: Doesn't most code live in web servers nowadays anyway? If so, it's either already running in parallel, or can be made to do so with very little effort.
BTW, parallelizing code in a web server can be harmful (!) to total throughput due to communication overheads. This is something that I find people don't always realize when they naively parallelize something, and is also sometimes missed by benchmarks that don't put the server under heavy load.
> A similar "tuning best practice" that I still see applied in the field is database servers with dedicated data, log, temp, and templog drives... on cloud VMs where this is guaranteed to have worse performance than simply pooling the same amount of disk into a single logical drive.
A statement like this depends on many concrete details about the specific cloud infrastructure (you didn't say AWS, you said "cloud"), and cloud VMs are usually not the best performing way to run a database anyway. So again, this is very context-specific.
Also, if database performance is critical enough that you consider tuning these things, then renting bare metal should at least be a consideration as well. People who have never measured it are sometimes surprised by how large the virtualization overhead is (and how cheap it is to rent physical servers). On actual hardware, this advice might still be relevant.
> I can't think of many (any?) other industries where more experienced senior staff can be this wrong about as many things...
The actual point is that you should measure things in concrete cases, not make general claims. In that sense, I unfortunately don't find your "updated" way to state what supposedly is and isn't relevant not much of an improvement to the "outdated" claims you argue against so vehemently.