Will you publish benchmarks for e.g. K80? Or provide a way for users to contribute? It's really handy to know, e.g. comparable to what is Resnet50 inference on a bunch of architectures.
Contribution guide is here: https://github.com/NVIDIA/MatX/blob/main/CONTRIBUTING.md
Like most benchmarks it really depends on what you want to do, and since it's a general library everyone might care about different things.
Versus the advert in the root Readme, which is impressive but gives no data on the pareto.