That's right, we've tested down to pascal, but this should work on Kepler too since CUDA and the underlying libraries support it.
Contribution guide is here: https://github.com/NVIDIA/MatX/blob/main/CONTRIBUTING.md
Like most benchmarks it really depends on what you want to do, and since it's a general library everyone might care about different things.
Versus the advert in the root Readme, which is impressive but gives no data on the pareto.