Intel's ISA-L ended up implementing this method. Their implementation is interesting because they took this further and took advantage of knowledge of the instruction latency to pipeline multiple iterations of this method to achieve some really amazing throughput.
For reference, see https://01.org/intel%C2%AE-storage-acceleration-library-open... and the source code at https://github.com/01org/isa-l (see the erasure code folder for details).
In general, I've found Prof. Plank's other papers and presentations very interesting, innovative, and accessible.