This is a valid point, and this is an overly long response because it distracts me from watching frightening current events.
There are two ways to look at these sorts of numbers, "CPU performance" and "Systems performance". To give an example from my history;
NetApp was dealing with the Pentium P4 being slower than the Pentium 3 and looking at how that could be. All of the performance numbers said it should be faster. They had an excellent OS group that I was supporting who had top notch engineers and a really great performance analysis team as well, the results of their work was illuminating!
Doing a lot of storage (and database btw) code means "chasing pointers." That is where you get a pointer, and then follow it to get the structure it points to and then follow a pointer in that structure to still another structure in memory. That results in a lot of memory access.
The Pentium 4 had been "optimized for video streaming" (that was the thing Intel was highlighting about it and benchmarking it with.) in part because videos are sequential memory access and just integer computation when decoding. So good sequential performance and good integer performance gives you good results on benchmarking video playback.
The other thing they did was they changed the cache line size from 64 bytes to 128 bytes. The reason they did that is interesting too.
We like to think of things a computer does as "operations" and you say "this operation takes 0.x second, I can do 1/x operations per second." And that kind of works except for something I call "Channel semantics" (which may not be the official name for it but it's in queuing theory somewhere :-).
Channel semantics have two performance metrics, one is how much bandwidth (in bytes/second) a channel has, and the other is what is the maximum channel operation rate (COR) in terms of transactions per second. Most engineers before 2005 or so, ran into this with disk drives.
If you look at a serial ATA, aka SATA, drive it was connected to the computer with a "6 Gb" SATA interface. Serial channels encode both data and control bits into the stream so the actual bytes that go through a 6 gigabit line can be < 600 Mb when the encoding is 10 bits in the channel for every 8 bits sent (called 8b/10b encoding for 8 data bits per 10 channel bits (or bauds)). That means that the channel bandwidth of a SATA drive is 600MB per second. But do you get that? It depends.
The other thing about spinning rust is that the data is physically located around the disk, each concentric ring of data is a track, and moving from track to track (seeking) takes time. Further you have to tell the disk what track and sector you want, so you have to send it some context. So, if you take the "average" seek time, say 10mS, then the channel operation rate (COR) 1/.010 or 100 operations per second.
So let's say you're reading 512 byte (1/2K) sectors from random places on the disk, then you can read 100 of them per second, but wait 100? That would mean you are only transferring 50 kB per second from the disk, what happened to 600MB?
Well as it turns out your disk is slow when randomly accessed, it can be faster if you access everything sequentially because 1) the heads don't have to seek as often, and 2) the disk controller can make guesses about what you are going to ask for next. You can also increase the size of your reads (since you have extra bandwidth available) so if you read, say 4 kb sectors, then 100 x 4 kB is 400 kB/second. And 8 fold increase just by changing the sector size. Of course the reverse is also true, if you were reading 10 Mb per read, at a 100 operations per second that would be 1000 Mb per second which is 400 Mb more than your available bandwidth on the channel!
So when your channel request rate is faster than the COR and/or the data size requests are greater than the available bandwidth, you are "channel limited" and you won't get any more out of the disk no matter how much faster the source of requests improves its "performance" in terms of requests/second.
So back to our story.
Cache lines are read in whenever you attempt to access virtual memory that has not been mapped to the computer's cache. Some entry in the cache is "retired" (which means over written, or written out first if it has been modified, and then overwritten) and the new data is read in.
The memory architecture of the P4 has a 64 bit memory bus (in 72 data bits if you have ECC memory) That means every time you fetch a new cache line, the CPU's memory controller would to two memory requests.
Guess what? The memory bus on a modern CPU is a channel (they are even called "memory channels in most documentation") that are bound by channel semantics. And while Intel often publishes it's "memory bandwidth" number, it rarely would publish its channel operation limits.
The memory controller on the P4 was an improvement over the P3, but it didn't have double the operation rate of the P3. (it was like 20% faster as I recall, but don't quote me on that.) But the micro-architecture of the cache doubled the number of memory transactions for the same workload. This was especially painful on code that was pointer chasing because the next pointer in the chain shows up in the first 64 bytes and that means the second 64 bites the cache fetched for you are worthless, you'll never look at them.
As a result, on the same workload, the P3 system was faster than the P4 even though on a spec basis the P4's performance was higher than that of a P3.
After doing the analysis some very careful code rewriting and non-portable C code which packed more data in the structures into the 128 byte "chunks" where both 64 byte halves had useful data in them. Improved the performance enough for that release. It also was that analysis that gave me confidence that recommending Opteron (aka Sledgehammer) from AMD with its four memory controllers and thus 4x memory operations per second rate was going to vastly outperform anything Intel could offer. (spoiler alert: it did :-))
Bottom line, there are performance" numbers and there is system performance* which are related, but not as linearly as certainly Intel would like.