Consider the problem of optimizing the response time of a set of servers. Load balancing - keeping the hardware busy - will keep the request queues at roughly equal and therefore minimize the maximum queue length. That ensures that the hardware is as busy as possible, but you can improve on that.
When a new request arrives, you can either add all of that request to one queue, or you can split it into pieces and distribute those across all the queues. Distributing the request, assuming perfect load balancing, means the latency for that request will be k + 1/n instead of k + 1, where k is the length of the queue.
In other words: Parallelizing gets you a substantial improvement in response time without additional hardware, and without making the code run any faster. I didn't realize the implications of this until recently.