Full threadbmc7505·Fast matrix multiplication would be a more useful benchmark: https://fmm.univ-lille.fr/View on HN