Just because a piece of code incurs hardware related performance issues does not mean it's "no longer close to the real machine". Cache misses? Reorder your data, or start inserting prefetch statements. Mispredicted branches? Issue a hint, or structure your code better.
Both of these are profile-guided optimizations. It's very difficult for compilers to optimize for access patterns that will only become clear in the context of execution, and often depend on your target spec machine.
RAM cache misses are pretty bad, but with the blocks that he's dealing with, he's looking at maybe 1 miss per block, and that's assuming that the processor is letting the blocks get entirely out of the 3 levels of cache to RAM.
The compliment array should stay in memory 100% of the time because it's tiny and referenced constantly, and the individual arrays shouldn't be that bad. He's not branching all that much, so unless the while and for loops are triggering after 1 or 2 characters, it's unlikely that's the issue.
The actual problem he's more likely running into is the OS rescheduling the thread to a different core, which is a gigantic hit.