I'm a bit confused as to why they cite https://www.usenix.org/conference/usenixsecurity16/technical..., but then don't evaluate their performance relative.
The selection of JTR and HashCat rule eval is troubling - 'Best64' isn't the way a skilled attacker uses those tools.
It would be cool to see these teams participating in cracking event; http://contest.korelogic.com/ or tasked against un-recovered corpus; https://hashes.org/
Of course, maybe a more accurate view was that the paper isn't actually seeking to advance to the state of the art in password cracking, and has other motivations.