The idea looks interesting (I have not yet read the actual paper), but calling it «The first new coherence mechanism in 30 years» is ridiculous.
In fact, you could start by looking at the Related Work section of this paper itself. The version at https://people.csail.mit.edu/devadas/pubs/tardis.pdf is better. It is quite telling that the paper does not make the outrageous claim of the title of MIT's press release.
O(log N) of memory overhead per block is nothing new. There were commercial systems in the 1990s that achieved that (search for SCI coherence). Note that there are other overheads to consider (notably latency and traffic).
This paper is very interesting and looks sound, but MIT's press release makes it look silly.
Excuse me for not even trying to make a summary of the last 30 years of research in this field.
The quote from the article
> MIT researchers unveil the first fundamentally new approach to cache coherence in more than three decades
is more hedged ("fundamentally"). I don't know if anyone in this thread has the expertise to evaluate that judgement.
The technique kind of reminds me of Jefferson's virtual time (and time warp), which rather exists in a distributed simulation context. Virtualizing time to manage coherence of reads and writes is a very good idea.
But my knowledge is a bit out of date, since GPUs are actually including caches now, and what they can do seems to evolve rapidly.
http://www.nvidia.com/content/pdf/fermi_white_papers/nvidia_fermi_compute_architecture_whitepaper.pdf