http://fabiensanglard.net/duke3d/duke3d_code_review_unrolled...
http://fabiensanglard.net/duke3d/duke3d_code_review_unrolled...
For example, if you know that the iteration count is divisible by 4, you could do something like:
int unrolledN = n / 4;
for (int unrolledI = 0; unrolledI < unrolledN; unrolledI++) {
int i = unrolledI * 4;
// loop body...
i++;
// loop body...
i++;
// loop body...
i++;
// loop body...
}
This wouldn't really offer any advantage over the plain loop, though. Next you'd need to reorganize the loop body so that e.g. memory reads for all iterations would occur at the start of the unrolled loop. This kind of optimizations can offer significant performance increases because you get more control over what the CPU is doing within the loop, but they also depend greatly on the target platform. Even different x86 processors can be very different in this respect, so unrolling can become a disoptimization easily.In this case, I think he's replaced function calls with the body of those functions. For example, "displayrooms" is immediately followed by { } surrounding the contents of what "displayrooms" actually does: interpolate, animate, and so on. This means you can read just the one source file and know what's being executed, instead of having to read the source of displayrooms.c and various other source files separately.
Another common form of "unrolling" is repeating the body of a loop some number of times. See, for example, http://en.wikipedia.org/wiki/Duff%27s_device . Basically, by reducing the number of branches (and therefore potential pipeline flushes) you can decrease computation time.