Heisenbugs[1] can be incredibly frustrating. Computers are supposed to be
deterministic. Even when bugs are incredibly complex[2], it's possible to systematically investigate
iff the behavior of the system is deterministic.
I once had to add 7 NOP instructions at the beginning of the bootstrap/"BIOS" code I wrote for a Z80 clone. I couldn't understand why test programs seemed to crash[3] about a 15% of the time. Later, I discovered the same behavior in code that previously worked. I spent over two weeks trying to investigate, which only produced more confusion as the behavior would sometimes go away or get much worse randomly with each change I made.
I finally found the bug using a (hardware) logic analyzer to watch[4] what the CPU was doing on the memory buss, Something wasn't finished resetting inside the CPU. Any instructions that ran too early would trash the internal state of the CPU, causing later instructions to have problems like asserting multiple chip select pins. Multiple ROM/RAM chips would try to drive the buss, and everything dies. The instruction this happened on depended on which instructions were run while the CPU was still resetting. The NOPs simply delay startup to let the CPU stabilize.
[1] http://www.catb.org/~esr/jargon/html/H/heisenbug.html
[2] http://www.catb.org/~esr/jargon/html/M/mandelbug.html
[3] "crash" == CPU locked hard with no activity on the memory buss until /RESET was grounded by the watchdog timer (or me)
[4] with a 100-pin PQFP clip-on probe that wouldn't stay attached