something as seemingly innocuous as checking and incrementing an atomic counter across a bunch of threads can induce a 10x performance drop
it's totally batshit insane how much faster things run when they're able to work with stuff in L1 cache where they're not waiting on slower memory.
modern CPUs are insanely fast and crazy concurrent, but the rest of the machine can't keep up.
The story is similar for LLVM. The fact that a CPU-bound task can be pessimized into stalling on cache and TLB misses all the time doesn't mean that it isn't CPU-bound.
A commonly deployed compression format which is lauded as revolutionary has a decoder that contains 0 SIMD intrinsics and uses significantly-fewer-than-optimal number of interleaved streams (according to papers published many years before the format was standardized!), which drives up the length of the dependency chains, and if you try to decode many chunks at the same time then you also get into trouble because you can't fit many tables into cache. The decoder does not saturate memory bandwidth and toy decoders of similar codecs from papers published years before are predictably faster.
The more ubiquitous compression formats are worse.
It's common to encounter hash tables that underperform robin hood tables that support bulk operations by a factor of 100, applications that use red-black trees with individually allocated nodes as priority queues (as opposed to ternary heaps or b-trees or whatever), and sorts that underperform by a factor of 10 compared to radix sort of multi-key quick sort because they are written in terms of a comparator function that must run in its entirety, so that the subset of the items whose first-sorted-by field are the same cannot skip this comparison. This comparison is commonly a comparison of data behind a pointer rather than data present in the struct.
Specialized workloads, for sure.