A fun optimization trick from rsync
blog.plover.com
blog.plover.com
typedef int (*MethodFunc)();
void doMethod() {
static int methodId = 0;
static const MethodFunc methodFuncs[] = {
#ifdef METHOD_1_CAN_COMPILE
method1,
#endif
method2,
...,
NULL
};
while (methodFuncs[methodId]) {
int success = (methodFuncs[methodId])();
if (success)
return 0;
methodId++;
}
return -1;
}While syncing some million files, I'm sure it'll have an impact tho.
P.S.: Yes, we use it for a million files, regularly.
the syscall is going to be many, many orders of magnitude slower than the difference between a direct and an indirect call
In theory, it's a very cheap piece of assembly after that point.
Now, as a working curiosity, this is pretty damn cool.
Switch on the last return point at the top of the function, and make all variables that need to persist between calls `static`.
(I've written it before, and I swear I didn't invent it... Maybe it's just a short skip from Duff's Device? Still, using it to pick up before the last return instead of after is new trick for me!)
1: https://www.chiark.greenend.org.uk/~sgtatham/coroutines.html
Or the code smell around fall through switch statements.
There are legitimate uses for fall through switch statements and/or mid-file includes, and or static function vars, but combining all of them together is not clever, it's a recipe for bugs.
Can be tricky balancing terseness with readability.
Yes, but.. compared to switch and preprocessor hacks described in the article? Granted it's subjective but I'd find a function pointer that's assigned to once to be way easier to comprehend.
It's like trying to push anything and everything through a little hole.
One concrete issue with this technique: it doesn't "loop". Meaning, for example, if initially method #3 doesn't work but method #4 does, it'll settle on method #4. If later, method #3 starts working, it won't switch it. Worse, if method #4 stops working, it'll just always fail.
I imagine this isn't an issue for rsync's use but could be a concern if applying it more generally.
But it's an "optimization", and avoiding previously attempted methods is the whole point of the trick. Otherwise it would be an easy and uninteresting function.
That‘ll probably try all methods on every single file that is processed, so if method #16 is the one that works, you get 15 unnecessary syscalls for every file.
I think (I haven't actually tested it) you can do what the parent commenter wants by wrapping the whole thing in a do-loop but the original implementation is questionable enough, sticking an extra loop in there isn't doing anything to help make it less horrible.
int set_the_mtime(...) {
static int switch_step = 0;
int loop_again = 1;
do {
switch (switch_step) {
case 0:
if (method_0_works(...))
return 0;
switch_step++;
/* FALLTHROUGH */
case 1:
if (method_1_works(...))
return 0;
switch_step++;
/* FALLTHROUGH */
case 2:
if (method_2_works(...))
return 0;
}
switch_step = 0;
} while (loop_again--);
return -1; /* ultimate failure */
} int set_the_mtime(...) {
static int starting_from = set_mtime_starting_from(0);
set_mtime_starting_from(starting_from);
}
// Try various ways of setting mtime until it works.
int set_mtime_starting_from(int starting_from) {
// horrible switch statement
}Also, if you're going to use static vars in multithreaded contexts doing this, better to write
thing = 5;
(via that macro magic) than `thing++`. Then at least nothing too bad should happen...I don't understand the need for the macro-fu to generate the case values. The use of switch_step++ in each previous case seems to drive that, but couldn't the same be accomplished by setting switch_step to a constant?
static int switch_step = 0;
switch(switch_step) { // Falls through between cases
default:
#if METHOD_0_AVAILABLE
switch_step = 0;
case 0:
if (method_0_works(...))
break;
#endif METHOD_0_AVAILABLE
#if METHOD_1_AVAILABLE
switch_step = 1;
case 1:
...
}One thing it doesn't have natively (unless it's really well hidden) is the ability to look at two directory structures, and generate a delta of the two out to a separate location (say a USB stick).
A very useful feature for backups where you've got good enough (to run an rsync --dry-run) network connectivity, but not good enough to actually do the transfer.
[1] https://jeddi.org/b/2016/05/28/rsync-from-a-to-b-via-usb/
you can feel free to write the batch directly to some portable media(Sadly it's clear it was in fact there long before I wrote a workaround in response to not finding it.)