"This algorithm is much faster now that I've used some lock free programming techniques I just read about!!!"
Code that works correctly on x86 but fails intermittently on another processor. It's like a nightmare unfolding...
"This algorithm is much faster now that I've used some lock free programming techniques I just read about!!!"
Code that works correctly on x86 but fails intermittently on another processor. It's like a nightmare unfolding...
This stuff is terrifying.
What sort of code are you thinking about? Even lower-level / non-abstracted atomics and memory fences that are specific to a particular architecture? I actually did run into code like this for a MIPS-based CPU and we converted the code to using the GCC builtins for portability. Funny thing is that the next hardware architecture toolchain wasn't GCC-based, so there was no benefit... we had to make wrappers for the atomics... and hunt down some absurd bugs related to L2 cache writebacks and mem fences. It was an exercise in how long I could stick with a problem... some serious hunting and staring at debug output / hex dumps..
[1] http://gcc.gnu.org/onlinedocs/gcc-4.1.2/gcc/Atomic-Builtins....