That's not optimal on real machines; it's often faster to recalculate x*y than to cache it.
In general this "optimal" strategy seems to assume infinite memory that's all accessible at the same speed.
In general this "optimal" strategy seems to assume infinite memory that's all accessible at the same speed.
In a similar vein, if typical strategies for unstable sorting of arrays in place use quicksort splitting for long spans and insertion sort for short spans it doesn' mean that either algorithm is "better".
I don't think "massively parallel" is all that good for performance either, because battery life is part of that.