Cutlass is from NVidia and AFAIK gets pretty close to cuBLAS performance. I think the use case for cutlass is when you only need a few kernels and don’t want to pull in a huge cuBLAS dependency and are ok paying a small perf penalty for that. Then there are times when you need custom kernels that not available in cuBLAS and for that cutlass is about as fast as it gets.