For instance there is the kind of simple interpreter that I know how to write off the top of my head which evaluates the nodes of an expression tree which is horribly inefficient.
Then there is something like Python which "compiles" code to bytecode which is then interpreted by a "virtual machine". That kind of system admits very complex optimizations about as far you as can imagine, including compiling some of the bytecode all the way to machine code the way PyPy or the JVM does it.
Note most interpretations of the Intel architectures break complex instructions up into RISC-like "micro-ops", which is a bit like bytecode interpretation.
One of the biggest concerns in a modern CPU is that memory accesses are very slow relative to a CPU cycle so you want to be executing a large number of instructions at a time to hide the latency of memory access. Intel's failed Itanium had the compiler try to explicitly schedule this but it didn't work because the compiler can't know ahead of time (for general purpose code) what is in what level of the cache and what isn't -- and in a worse case scenario this can hang up not only one instruction but many other instructions that depend on the result of that instruction, ether because they really depend on the result or because the CPU gets hung up. The CPU, on the other hand, does know, and can be opportunistic about taking advantage of available parallelism.