Faster memory comparison in C
macosxfilerecovery.com
macosxfilerecovery.com
You can see that CBMC proves that (1) a zero result from mycmp guarantees a zero result from memcmp, and (2) a nonzero result from mycmp guarantees a nonzero result fro mycmp. But (3) a negative result from mycmp doesn't guarantee a negative result from memcmp and (4) a positive result doesn't guarantee a positive result.
So, modulo problems with "strict aliasing" in modern C standards, you can use this code to compare blocks of memory for equality, but you cannot use it to order blocks according to a less-than or greater-than predicate.
https://gist.github.com/jepler/760cfcd4b326d6d11241256a6a4c7...
Regardless of alignment penalties, memcmp always needs to test for alignment so that it doesn't read past the end of a page, causing a segfault. This hack isn't appropriate for general purpose use; only where you know for sure an unaligned read won't overflow a page boundary.
And I'm not sure why you think inlining would result in the speed increase? I believe the compiler will inline it when the standard memcmp is used too.
Regarding inlining, using GCC 6.2 with "-O3 -march=native" on macOS I first had to remove "111111" and "222222" as constants, otherwise GCC precomputed both loops.
When I used __attribute__((noinline)) on mycmp, half of the difference went away. When I forced mycmp to check the alignment, mycmp became slower by almost a full second.
That was on a 2011 Mac Mini. Using a newer box with a Haswell chip (Xeon E3-1230 v3) and GCC 5.4.0 performance is about the same when preventing inlining and adding the alignment check, rather being slower. In both cases the assembly confirms that neither mycmp nor memcmp were inlined.
if (* src1_as_int ++ != * src2_as_int ++) return 0; /* compare as ints */