Also, for rarely-run pieces of code (micro-)performance may not matter, but size certainly does - the smaller the non-performance-critical code, the more room there is in the cache for the performance-critical code. On a larger scale, the smaller the whole process is, the less chance there is of the non-performance-critical code being swapped out (and suddenly becoming very performance-critical...) when memory is constrained, and of course it also reduces overall memory usage, which is a good thing in a multiprocess environment.
Perhaps we need to fundamentally change how compilers and their optimisers work, since the current model seems to be "generate the most stupidly unoptimised code possible, then apply optimisation passes to it". I think it's a rather counterintuitive and wasteful way of doing it, compared to the "generate only the code you need" method a human would likely follow. Thus I'd say that current compiler optimisations are still useful, if only for cleaning up the mess they would make otherwise. :-)