In scientific computing it's not unusual to get 2x-3x speedups by doing simple optimizations like this one. For desktop applications it's nothing, but when your simulation is taking a week to complete then 2x-speedup is a huge improvement.
This was especially the case on the old Pentium 4's. What we used to do was over-allocate the amount of memory needed, then offset the pointer into the allocated memory to avoid cache collisions.
It's not always nothing, particularly when you're dealing with video. :)