My first reaction: Some Nvidia cards can synchronize between warps? Nice!
I've been living in OpenCL world which is pretty much just everyone except Nvidia (because nvidia intentionally ignores OpenCL and cripples their support for it) so I have unfortunately missed this development.
On the other hand the particular use case in the article was to circumvent some other limitations of the architecture, such as relatively small register file and the shared instruction counter for a block of 32 threads. Clever and interesting nevertheless.