However, It's also clear what gravypod is trying to say. The architecture itself has a greater influence on the "processing power" nowadays because the clock has reached physical limits and can't be increased much further without getting expensive (see IBM Power Sytems). Bottomline is that you can't compare CPUs by frequency alone, but it was never any different, the architectures 15 years ago just sucked so hard that all of them were too similar regarding instructions per cycle.
I think we can blame AMD for that line of thinking, they had to push hard against Intel during a time where everyone used clock frequency to compare CPUs.
*edit: Sorry I remembered it wrong, the commonly used power equation is P = C x V^2 x f (capacity, voltage, frequency). Power is linear dependent on the frequency!
I don't see how this is ultimately true.
The amount of electrons needed to flip a capacitor is constant given that the voltage is constant. The more often you flip, the more current you need per time.
The amount of electrons needed to flip a capacitor is proportional to the voltage. So, lower voltage means a) less power per electron and b) less electrons. Therefore the power is quadratic with the operating voltage.
Third, if you want to increase the frequency, at some point, you have to increase the voltage or you get errors. Then you increase frequency and voltage at the same time and your power consumption shoots up.
In the meantime I can only offer this [1] stackexchange post. Power of a CPU is definitely quadratic in f, or even worse.
[1] http://physics.stackexchange.com/questions/34766/how-does-po...
edit: I probably remembered it wrong and I now think that that the equation was P=C x V^2 x f. However, transition power effectively has an exponent much greater than 2, but that's apparently caused by the increased die temperature (according to the post on Anandtech).
e.g. ARM Cortex A7: 2,850 MIPS at 1.5 GHz
Qualcomm Krait (Cortex A15-like, 2-core): 9,900 MIPS at 1.5 GHz
Both processors have the same clock frequency, but one has over three times the processing speed in MIPS.
[1] https://en.wikipedia.org/wiki/Instructions_per_second#Timeli...
Referring to the list on wikipedia again, compare two different 4-core CPUs:
Intel Core i5-2500K 4-core: 83,000 MIPS at 3.3 GHz
Intel Core i7 875K: 92,100 MIPS at 2.93 GHz
Modern CPUs use pipelining which executes many instructions parallel. This only works well if everything goes as predicted. If you have an algorithm which works contrary to what the branch prediction thinks and a cache which does not hold the data you need, your performance goes down the drain. Those MIPS mean nothing if not put into the right context.
Cache misses make CPU speed irrelevant, but when you look at your memory system it's another instance of the same problem with the same tradeoffs of frequency vs. work per cycle. And when the "right context" is waiting for the outside world, that's not the most important context.
http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-93-6.pdf
It explores the limits of instruction-level parallelism. If you had a processor that could dispatch an unlimited number of independent instructions simultaneously, how much of an improvement in common algorithms would you get?
I highly recommend taking a digital logic and computer architecture class in a CS program to understand how pipelines/instruction units are laid out (microcoding, branch prediction, pipeline stalls, macro ISA, etc.), what makes them fast/slow and challenges to implementation.
As a side note, the Tomasulo algorithm is a huge pain in the ass to do in hardware. Props to the processor engineers who are implementing modern OOO execution hardware in whatever awful proprietary HDL/IDE combo your company makes you use.
A really excellent example here are ARX (add rotate xor) algorithms in cryptography. Take BLAKE2 for example; the reference C code, compiled to plain x86-64 assembly mainly consists of long rows of addq/xorq/rorq instructions. A modern CPU manages to schedule these so efficiently that the performance difference to a hand-coded SSE or AVX version practically never matters.
If you take a peek at Haswell's port map (eg. http://images.anandtech.com/reviews/cpu/intel/Haswell/Archit... ) you'll see that there are four out of eight ports who can do integer arx operations. Since memory and addr ops got their own ports this means that even the straight addq/xorq/rorq code can almost fully utilize the core.
This is no accident, though, the algorithm was designed to take advantage of designs like this :)
GP argues that even though frequency has not increased recently, performance has, a point better made by the article.
You can't.
For instance, CPUs have sped up a lot compared to memory. As a result, if you just calculate as fast as you can, you run out of data to calculate on. Suddenly it becomes more relevant how fast you can access memory. So now we have all sorts of complicated cache mechanisms to help us keep the data to hand. Coding needs to be cache friendly as well.
If you look at CPU prices, you can see the amount of cache increases with the price.
Frequency is far from an accurate measure of performance. It doesn't take in IO, context switching, pipe lining, interrupts, and multicore systems. Even within the same family clock speed may not have much to do with performance. Clock speed is just a measure of how many times a second some increment in an operation may occur. Doesn't mean it will increment it's progress in an operation, doesn't mean the operation can't happen between then, it just means the CPU will find out about the state of that between those updates.
You also need to take into account the BPU and the vector math sections of the die which this sort of metric completely ignores.