I.e., LAPACK/BLAS benchmarks are just really big linear algebra matrix problems, so obviously your pre-fetch and branch prediction performance will be significantly better since you aren't dealing with interrupts, locatedb, or Windows DCOM events firing off in the background. You have a huge set of matrices with a very predictable set of branches, fetches, and decodes, so obviously your CPU can optimize for that load, you're just paying for it in latency on the back-end (RAM fetches are the new disk swap ;)).
On all those benchmarks, (i.e. your standard LU matrix decomposition which previously was the basis of the LAPACK benchmarks, though things might have changed in the ~10 years since I've really looked at things) isn't CPU-bound anymore, so of course your instruction-per-cycle load on the CPU isn't where you'll be bottlenecking (and hasn't been since "let's avoid floating-point operations and just use static look-ups instead since we don't want the 10x cost of using the FDIVP instruction!"). Your processor can very easily anticipate from where in that sparse-matrix your next data fetch is going to be. It's the cost of that RAM fetch[1] going along that copper trace which is going to be where you're going to bottleneck on any heavy numerical computation.
The power consumption on your CPU might drop a nominal amount which is great for those marketing white papers, but for a numerically heavy load, you're paying just as much (in total power consumption per 4U in the data center, total heat generation/dissipation within the case, and total processing time) on the back-end for those fetches.
[1] https://i.stack.imgur.com/a7jWu.png (I normally cite academic references, but this is 'good enough' to convey my point, I hope).