Java Memory Model Pragmatics
shipilev.net
shipilev.net
Given that most developers work at best with embarrassingly parallel problems they wouldn't need to know much of these details.
But if you're up for it: It's a very rewarding feeling when you're both allowed to figure these things out and someone are also paying you while you figure it out.
I think the problem in the example from the article is that the spec doesn't say what happens to i when it appears twice in the same expression. What is the "current" value of i each increment sees?
Maybe they both see the initial value 5, and the result is (5+1) + (5+1) = 12 like you say. But compiling the example with gcc 4.6.3 gives me 14 as a result! It seems as if gcc is seeing the two updates and then replacing each occurrence of i with the final value 7 (but I'm guessing). I guess the point is that, since it isn't in the spec, each implementation of C is free to do whatever the implementors decided.
I'd expand it beyond Java to shared-everything models in general. Any language with a shared-everything model is going to incur all this complexity. At least Java goes through this pain to ensure memory safety in the presence of unsynchronized accesses from multiple threads.
In most cases you can get away with coarser constructs (which are higher-performing anyway, eg task parallelism vs. trying to decompose an individual operation) and only dive into the language lawyering semantics when you, eg, really really need a lockfree concurrent linked list.
As you note you rarely need to concern yourself with the details of the memory model as most of us are users, not designers, of concurrency primitives. But you should you need to design a new primitive you can actually have some guarantees it will work in a cross-platform manner, which is more than you an say about any other platform I know of.
Having it defined is very important. However part of the problem is that definition actually takes away guarantees that many CPUs (eg. x86) give you. For example double checked locking works in x86 but famously not in Java. So the fact is that Java is giving you in some ways a less sequential model than what you started with is what makes it particularly insidious.