I don't understand why we should compare performance in artificial benchmarks to power consumption. Electricity is cheap and I assume that most laptop users have a wall socket nearby. Instead, it makes sense to compare typical user performance (not CPU instruction count) versus price.
For example, let's say CPU A has 32 cores each capable to perform 1B instructions and costs $1000, and CPU B has 4 cores with same speed, but costs $100. Let's say typical task for a machine is browsing social network sites. If this task uses only 4 cores and a user can browse same number of web pages per unit time then effectively both A and B have the same performance, but CPU A has 10 times worse typical performance per dollar.
A metric in this case should be not the count of executed instructions, but for example, a count of web pages viewed by user per unit time, or share of time spent actually viewing pages (and not waiting for them to load).
Of course, you can make similar metrics for other uses - for example, a time taken by illustrator to draw a picture, or count of written lines of code per unit time for a programmer. All divided by hardware cost.
That's the type of comparison I would like to see. Would Apple be winning in such fair and realistic benchmark is an open question. I doubt that.